Trust is earned, not given

A different perspective

2020-07-08 · Projects

DevOps on AWS, part 2: Compute — EC2, ECS/Fargate, and Lambda

Part 2from the DevOps on AWS series · 6 parts in all

Part 2 of the AWS series: choosing where your code runs. The three compute shapes — VMs, containers, functions — are not competitors; they are three answers to "who patches the OS and who pays for idle?" Our sample companies choose differently, and that's the lesson.

EC2: the classic VM

You pick an instance type, an Amazon Machine Image (AMI), and you own everything — patching, scaling groups, capacity. Maximum control, maximum responsibility:

# Launch a minimal web server from the CLI (dev/test only - in prod this is Terraform)
aws ec2 run-instances \
  --image-id ami-0abcdef1234567890 \        # a pinned AMI, never "latest"
  --instance-type t3.micro \                # 2 vCPU, 1 GB - the free-tier staple
  --subnet-id subnet-0123456789abcdef \     # private subnet, ALB in front
  --iam-instance-profile Name=webapp-role \ # the instance gets permissions via role
  --tag-specifications 'ResourceType=instance,Tags=[{Key=app,Value=maplecart-web}]'

# An Auto Scaling Group keeps N healthy instances behind the ALB:
aws autoscaling create-auto-scaling-group --auto-scaling-group-name maplecart-asg \
  --min-size 2 --max-size 8 --desired-capacity 2 \
  --health-check-type ELB --health-check-grace-period 60 \
  --vpc-zone-identifier "subnet-a,subnet-b"   # two AZs - always

Use EC2 when you need long-running, specialized, or license-bound workloads, or full kernel access. For the three companies: nobody's web tier starts here anymore.

ECS and Fargate: containers without the babysitting

You package the app as a Docker image; ECS schedules it. With the Fargate launch type, AWS runs the underlying instances — you specify vCPU/RAM per task and pay per second of usage:

# Build and push the app image to ECR (Elastic Container Registry)
aws ecr create-repository --repository-name maplecart/web
docker build -t maplecart/web:v1 . 
aws ecr get-login-password | docker login --username AWS --password-stdin <acct>.dkr.ecr.<region>.amazonaws.com
docker tag maplecart/web:v1 <acct>.dkr.ecr.<region>.amazonaws.com/maplecart/web:v1
docker push <acct>.dkr.ecr.<region>.amazonaws.com/maplecart/web:v1

# Register a task definition (the "pod spec"): image, CPU, memory, port, role
aws ecs register-task-definition --cli-input-json file://task-def.json

# Run it as a service behind the ALB, Fargate, 2 tasks across AZs
aws ecs create-service --cluster maplecart --service-name web \
  --task-definition maplecart-web:1 --launch-type FARGATE --desired-count 2 \
  --network-configuration 'awsvpcConfiguration={subnets=[subnet-a,subnet-b],
    securityGroups=[sg-web],assignPublicIp=DISABLED}' \
  --load-balancers targetGroupArn=<tg-arn>,containerName=web,containerPort=8080

MapleCart and FleetView both live here. Blue/green deploys become "update the service with the new image tag"; CodeDeploy shifts traffic 10% → 50% → 100% with automatic rollback on alarm.

Lambda: functions, billed per millisecond

Upload code, set a trigger (HTTP via API Gateway/ALB, S3 events, queues), and it runs — no servers at any level of abstraction:

# A Lambda handler (Python) that resizes images dropped into S3
def handler(event, context):
    import boto3
    from PIL import Image
    s3 = boto3.client("s3")                    # credentials come from the
    bucket = event["Records"][0]["s3"]["bucket"]["name"]   # function's IAM role
    key    = event["Records"][0]["s3"]["object"]["key"]
    if not key.lower().endswith((".jpg", ".png")):
        return {"status": "skipped"}           # guard clauses first
    local = "/tmp/img"                         # /tmp is the only writable dir
    s3.download_file(bucket, key, local)
    im = Image.open(local); im.thumbnail((800, 800)); im.save(local, optimize=True)
    s3.upload_file(local, bucket, f"thumbs/{key}")
    return {"status": "ok"}

NewsGrid's API and image pipeline are Lambda: idle hours cost zero, and the viral spike just means more concurrent executions — AWS scales it automatically. The constraints to respect: 15-minute max runtime, cold starts (mitigate with provisioned concurrency), and no state between invocations (that's what DynamoDB/Redis are for).

Choosing, in one table

EC2ECS/FargateLambda
Unit you manageOS + appimagefunction
Idle costfull priceper-taskzero
Scale unitinstancetaskrequest
Cold startnone (warm VM)seconds100ms-1s
Fitspecialized/legacylong-running servicesevent-driven, spiky

Next: where the data lives — S3, RDS, DynamoDB — and how MapleCart, FleetView and NewsGrid partition their data across them.