DevOps on AWS, part 3: Data — S3, RDS/Aurora, DynamoDB, and the caching layer
Part 3from the DevOps on AWS series · 6 parts in all
Part 3: the data layer, where architecture mistakes are most expensive to undo. The AWS mental model: S3 for everything byte-shaped, RDS/Aurora when you need SQL and transactions, DynamoDB when you need scale-without-ops on key-shaped access, and a cache in front of all of it.
S3: the infinite disk that isn't a disk
S3 stores objects in buckets, scales to exabytes, and versions everything if you ask. It is MapleCart's product-photo store, FleetView's export dump, NewsGrid's entire static site:
# Static website hosting for NewsGrid (origin for CloudFront)
aws s3 website s3://newsgrid-site/ --index-document index.html --error-document 404.html
aws s3 cp ./dist s3://newsgrid-site/ --recursive --cache-control "public,max-age=300"
# Versioning + lifecycle: keep 90 days of versions, then glacier the old ones
aws s3api put-bucket-versioning --bucket maplecart-media \
--versioning-configuration Status=Enabled
aws s3api put-bucket-lifecycle-configuration --bucket maplecart-media --lifecycle-configuration '{
"Rules": [{"ID": "archive", "Status": "Enabled", "Filter": {},
"Transitions": [{"Days": 90, "StorageClass": "GLACIER"}]}]}'
The lifecycle rule is quiet money: 90-day-old media moves to Glacier at ~1/5 the price, transparently.
RDS/Aurora: SQL without the 3 a.m. pages
RDS runs Postgres/MySQL/SQL Server with automated backups, patching, and failover. MapleCart's orders live here:
# Multi-AZ Postgres: standby in another AZ, automatic failover
aws rds create-db-instance --db-instance-identifier maplecart-db \
--engine postgres --db-instance-class db.t3.medium --multi-az \
--allocated-storage 50 --storage-encrypted \
--backup-retention-period 14 \
--db-subnet-group-name private-subnets \
--master-username admin --master-user-password <from-secrets-manager>
# Read replica for the product-browsing traffic (reports, catalog pages)
aws rds create-db-instance-read-replica --db-instance-identifier maplecart-db-ro \
--source-db-instance-identifier maplecart-db
Two disciplines that matter: credentials never live in code (store them in Secrets Manager; the app's IAM role reads them at boot), and set backup-retention before you need it. Aurora is RDS's high-performance sibling — same API, storage auto-scales, up to 15 read replicas.
DynamoDB: scale without an ops team
FleetView's telemetry — one row per device ping, billions of rows, always accessed by
(tenant_id, device_id, timestamp) — is the textbook DynamoDB shape. You declare
the key, DynamoDB handles partitioning:
# Table: partition key tenant_id, sort key device_ts (device#timestamp)
aws dynamodb create-table --table-name telemetry \
--attribute-definitions AttributeName=tenant_id,AttributeType=S \
AttributeName=device_ts,AttributeType=S \
--key-schema AttributeName=tenant_id,KeyType=HASH \
AttributeName=device_ts,KeyType=RANGE \
--billing-mode PAY_PER_REQUEST # on-demand: pay per request, no capacity planning
# Put an item (from Python with boto3; credentials via the task's IAM role)
table.put_item(Item={
"tenant_id": tenant, "device_ts": f"{device}#{ts.isoformat()}",
"speed_kph": speed, "fuel": fuel, "ttl": int(time.time()) + 90*86400,
})
That ttl attribute auto-deletes rows after 90 days — telemetry ages out
without a single cron job. The trade-off to internalize: DynamoDB is blazingly fast
for key access and useless for ad-hoc SQL joins. Choose by access pattern, not by fashion.
ElastiCache: the speed-of-light layer
MapleCart's product catalog barely changes but is read constantly — classic Redis territory. Cache-aside pattern: check Redis, on miss read Postgres and populate, TTL 5 minutes. One well-placed cache typically absorbs 80-95% of read traffic on hot data, which downsizes the database you have to pay for.
Next part: the network that ties it together and the IAM that locks it down.