DevOps on AWS, part 5: Pipelines and observability — shipping and watching
Part 5from the DevOps on AWS series · 6 parts in all
Part 5: the two loops of DevOps — the delivery loop (code → production) and the feedback loop (production → you). Teams that automate the first but not the second ship faster and know less.
The delivery loop
The canonical pipeline for a container service (MapleCart's web tier), in GitHub Actions — the same shape CodePipeline gives you natively:
# .github/workflows/deploy.yml (abridged; secrets configured in repo settings)
name: deploy
on: { push: { branches: [main] } }
jobs:
deploy:
runs-on: ubuntu-latest
permissions: { id-token: write, contents: read } # OIDC: no stored AWS keys!
steps:
- uses: actions/checkout@v4
- uses: aws-actions/configure-aws-credentials@v4
with:
role-to-assume: arn:aws:iam::<acct>:role/gha-maplecart-deploy # short-lived creds
aws-region: us-west-2
- name: Build and push image
run: |
docker build -t $ECR/maplecart/web:${{ github.sha }} .
docker push $ECR/maplecart/web:${{ github.sha }}
- name: Run tests against the image
run: docker run --rm $ECR/maplecart/web:${{ github.sha }} npm test
- name: Deploy
run: |
aws ecs update-service --cluster maplecart --service web \
--force-new-deployment \
--task-definition $(sed "s|IMAGE_TAG|${{ github.sha }}|" task-def.json | ...)
Two details worth copying: OIDC login (the pipeline assumes an IAM role per run — no static keys to leak) and the git SHA as the image tag (every running container is traceable to an exact commit; rollback is redeploying the previous SHA). For higher-stakes services, CodeDeploy adds automatic traffic shifting — 10% of users see the new version, CloudWatch alarms watch error rates, and any breach rolls it back before a human notices.
The feedback loop: CloudWatch
Three layers, in order of installation (most teams install them backwards):
- Logs — every container's stdout flows to CloudWatch Logs
(
awslogsdriver in the task definition). Searchable, retention-managed. - Metrics — the golden four per service: request rate, error rate (5xx), latency (p50/p95/p99), saturation (CPU/memory/queue depth).
- Alarms on symptoms — page on user-visible problems ("p95 latency > 800ms for 5 minutes", "5xx rate > 1%"), not causes ("CPU > 80%"):
aws cloudwatch put-metric-alarm --alarm-name maplecart-web-5xx \
--metric-name HTTPCode_Target_5XX_Count --namespace AWS/ApplicationELB \
--statistic Sum --period 60 --evaluation-periods 3 --threshold 10 \
--comparison-operator GreaterThanThreshold \
--dimensions Name=LoadBalancer,Value=$ALB_NAME \
--alarm-actions <sns-topic-arn> <asg-rollback-policy>
And the cheap trick that catches what metrics miss: synthetic canaries — a Lambda that loads your checkout page from outside every minute and alarms on failure. NewsGrid uses exactly this to detect "viral story is down" before Twitter does.
The weekly rhythm
DevOps on AWS settles into a loop: Terraform plan on every PR (reviewed like code), deploy
on merge, alarms watch, post-mortems turn incidents into new alarms and new guardrails.
Next part closes the series the best way possible: one Terraform file that builds MapleCart's
entire stack — network, database, cluster, pipeline plumbing — with a single
apply.