- Full Stack Engineer
- Serverless AWS
- Compliance Automation
- AI Tooling
Steven Gorrell
I build software and solutions at scale in AWS
-
4
AI certifications Including Palantir Professional Certification
-
5
AWS certifications Including DevOps Professional
-
60+
Services and automations shipped to production Several running continuously for 5+ years
About
Continuous adaptation, refinement, and improvement.
I started on Windows servers and kept evolving. Systems administration became SharePoint engineering. SharePoint engineering became running a team of six engineers and five consultants with a multi million dollar book of business. Then I transitioned back to delivering custom solutions.
After moving my focus has been on building and delivering reliable, scalable, event-driven, decoupled, serverless applications — API Gateway, Lambda, Step Functions, DynamoDB, SQS, ECS, ECR, Fargate, OpenSearch — provisioned entirely through Terraform and deployed by automated pipelines.
Many of these solutions are required to be PCI DSS compliant. I have multiple years of experience implementing requirements, and creating and presenting evidence to auditors.
Most recently I've been working in the AI space — building Model Context Protocol servers, and delivering full stack solutions for clients on top of AWS Bedrock and Palantir.
What I do
Things I get called in for
-
AI tooling & MCP
Model Context Protocol servers that expose internal APIs to LLM assistants with per-tenant authorization enforced server-side — plus agentic playbook runtimes built on LangGraph and Amazon Bedrock. Custom solutions built on top of Palantir. AWS Certified AI Practitioner. Palantir Data Engineer Certified.
-
Serverless architecture
Event-driven services built from API Gateway, Lambda, Step Functions, DynamoDB, SQS and SNS. Designed for fault tolerance first: idempotent handlers, dead-letter queues, and a failure alerting on every state machine.
-
Platform APIs & front ends
Versioned REST APIs with custom Lambda authorizers, OpenAPI documentation and Athena-queryable request logs — paired with React and TypeScript single-page front ends that engineers actually use during an incident.
-
Data & analytics pipelines
Ingestion and transform pipelines feeding OpenSearch, Kinesis streams and Palantir Foundry — ending in dashboards that executives use for revenue trend and business-review reporting. Palantir certified Foundry Data Engineer.
-
Infrastructure as code
Terraform or CloudFormation from the VPC up. Reusable modules, plan-and-approve pipelines with manual gates for production, and disposable per-branch test environments that tear themselves down.
-
Compliance engineering
PCI DSS subject-matter expert. I build the controls rather than document them after the fact: continuous anomaly detection against an analysis engine, with daily evidence committed to version control for review.
Selected work
Systems I designed, built and operated
Nearly all of my professional code lives in private repositories, so what follows describes each system by its problem, architecture and outcome. Happy to go deeper on any of these in conversation.
-
AI & MCP backed troubleshooting and analysis tool
Engineers needed a way to spend less time reading and analyzing support tickets, and help diagnosing issues. I built a tool to take the ticket data in real time and run automated analysis against existing cloud infrastructure, knowledge databases, and various other supporting API's. Data gathered in this analysis is then sent to AWS Bedrock which has a multiple custom MCP servers. The AI provides a detailed summary of the issue and provides suggested troubleshooting steps with a confidence score, and will also generate remediation scripts or code patches if it deems them appropriate.
- LangGraph pipeline, presents data to AI "personas" tuned for engineering and software development.
- DynamoDB caching layer.
- React frontend.
- Python MCP servers.
- AWS Bedrock, OpenSearch, APIs, and MCP servers give the LLM context.
- End users are presented with actionable solutions.
-
Automated alarm triage & remediation pipeline
Support engineers were spending their shifts running the same diagnostic steps against recurring cloud resource alarms. This service intercepts the monitoring event stream, filters noise, and hands qualifying alarms to Lambda backed Step Function workflows that execute diagnostic and remediation automation against effected resources. After running it provides the engineer with a detailed analysis of findings and fixes performed.
- Around forty per-alarm handlers covering compute, database, load balancer, backup, VPN and DNS conditions.
- Filtering keeps invocation cost proportional to actionable events rather than total event volume.
- Catch-all exception handling on every branch: a failure still produces an event, so nothing is silently dropped.
- Extended to Google Cloud, where a single remediation absorbed 20% of incoming alarm volume; later extended again to the private-cloud ticketing platform, taking a public-cloud tool company-wide.
- Set new records for executions and support hours saved month over month across its first full year.
-
Multi-cloud support tooling automation platform
A platform for deploying standardized configurations and monitoring tooling across a managed fleet spanning multiple cloud vendors. A product catalogue defines what can be deployed; a versioned REST API and worker fleet handle deployment, update and removal; every request lands in an Athena-queryable audit log.
- Lambda-router API behind API Gateway with a custom authorizer and published OpenAPI documentation.
- Athena schemas over structured request logs, so auditing a change is a simple SQL query.
- React for end user interaction.
- Over 1,100 automation's supported.
-
Continuous PCI anomaly detection with audit trail
PCI-regulated environments must prove that every privileged API call was expected. This service evaluates an environment's control-plane event stream against a declarative ruleset of authorized principal, action and resource combinations. Anomalies are alerted on and recorded in real time.
- Declarative ruleset keyed on principal, action and target resource — reviewable as code.
- Deployment pipelines are themselves modelled in the ruleset, so automated infrastructure changes are analyzed.
- Fail-open-to-alert design: any processing exception emits an unexpected-event record rather than losing the event.
- Provides assessors with an immutable, timestamped, reviewed trail.
-
Multi-cloud account assessment & reporting tool
Answering "what is actually running in this customer's cloud, and is it configured the way we recommend?" was a multi-day manual exercise. This tool automated it: a parser expands the request, agents collect inventory and configuration across accounts, results are processed asynchronously with progress tracking, and the output is rendered as a spreadsheet or published straight to the internal wiki. Built first for AWS, then extended to Google Cloud.
- Asynchronous job model with progress reporting, so long-running scans don't hold a request open.
- Agent-based collection via Systems Manager rather than requiring inbound network access.
- Pluggable output writers — Excel for customers, JSON, wiki markup for internal review.
- Parallel AWS and GCP implementations sharing the job and reporting model.
-
MCP servers for LLM-assisted cloud operations
Giving an AI assistant useful access to cloud tooling is mostly an authorization problem, not a prompting one. I built the internal AWS tooling API and the Model Context Protocol server that fronts it, exposing resource lookup, configuration diffing and instance inspection as MCP tools with tenant scoping enforced server-side: when a tenant header is present, the server validates the model-supplied account identifier against it and independently confirms ownership before any call proceeds.
- Server-side tenant verification, so a compromised or confused model cannot widen its own scope.
- Tool surface deliberately narrow and read-oriented — capability granted by design, not by prompt instruction.
- Deployed as infrastructure-as-code alongside the API it wraps.
-
Executive analytics pipelines and dashboards
Data engineering for leadership reporting on the Palantir Foundry platform: Python transforms that shape CRM, ticketing and revenue data, TypeScript server-side functions that compute the metrics, and React front ends built with Vite and Recharts that turn them into revenue-trend and monthly business-review dashboards. Also built the streaming transforms that feed operational ticket data into the same platform.
- Transform layer separates raw ingestion from the metric definitions consumed by dashboards.
- Server-side TypeScript functions keep metric logic in one place across dashboards and applications.
- Front ends built against a shared internal design system for consistency with other tools.
Open source & publications
Work that's public
Most of what I write is closed source, but I have contributed to a number of open source projects.
-
Contributor to the widely used Python library for generating CloudFormation templates. My merged work added support for AWS Auto Scaling Plans and Aurora backtrack windows, corrected the EC2 Spot Fleet load-balancer structure and an over-strict launch-template requirement, filled in missing CloudFront parameters and validators, and fixed a DynamoDB validator error.
-
Author and long-term maintainer of several modules in Rackspace's public Terraform library. I wrote the SQS module outright, brought S3, SNS and SQS to parity with their CloudFormation equivalents, added Route 53 support and queue policies, fixed redrive-policy handling, reworked bucket ACL behaviour to match AWS defaults after the provider's ACL changes, and set up their CI.
-
Contributed Canon EOS R3 support to darktable, the open-source photography workflow and raw development application. Camera support there depends on contributors doing the measurement work: I shot and processed calibration reference frames to mathematically derive values for white-balance presets across the camera's colour-temperature and preset range, plus a denoise profile covering its ISO range. Raised the tracking issue, then delivered both data sets.
-
Contributing author on Wrox's professional administration title. Alongside it, I presented pro bono at SharePoint Saturday community events, wrote and taught internal Rackspace University classes, and delivered customer-facing webinars.
-
One of my solo-built internal services is described in the solution section of Rackspace's published AWS case study.
Credentials
Certifications
Amazon Web Services
-
DevOps EngineerProfessional · 2024
-
Solutions ArchitectAssociate · 2016
-
DeveloperAssociate · 2021
-
SysOps AdministratorAssociate · 2024
-
AI PractitionerFoundational · 2025
Data platform & AI
-
Palantir Foundry Data Engineer2026
-
Palantir Foundry & AIP Builder Foundations2026
-
Palantir Foundry Aware2026
-
Generative AI with Large Language ModelsDeepLearning.AI & AWS · 2023
Microsoft
-
MCTS: SharePoint 2010
-
MCITP: Server AdministratorWindows Server 2008
-
MCSA: Windows Server 2008
-
MCTS: Active DirectoryWindows Server 2008
-
MCTS: Server Network Infrastructure
-
Microsoft Certified Professional
Other
-
FCC Amateur Radio Technician
-
Wilderness First ResponderNOLS Wilderness Medicine
Contact
Get in touch
Email is the reliable way to reach me, and I answer it. I'm happy to talk through any of the work above in more detail, and open to conversations about interesting problems in cloud platform engineering. Generally available for calls 2:00–7:00 PM Central — email first and we'll find a time.