Posts

Showing posts with the label AWS

From Logs to Alerts: SLOs for Go APIs on AWS

Image
Logs are useful after something breaks. SLOs are useful before users start sending screenshots. The shift from log-based debugging to service-level objectives is one of the biggest maturity jumps a backend team can make. For Go APIs on AWS, I like starting with a small set of SLOs that match user pain: availability, latency, and freshness. Everything else can grow from there. Define What Good Means A service-level indicator is the measurement. A service-level objective is the target. For an API, the indicators are usually request success rate and latency. For a data pipeline, freshness matters too. 99.9% of API requests should return non-5xx responses over 30 days. 95% of dashboard requests should complete under 500ms. 95% of reporting data should be less than 15 minutes stale. Instrument at the Edge Measure user-visible behavior at the edge of the service. Handler middleware is a good place for request count, status, and duration. Do not build an SLO from internal function timi...

Trace Context Propagation Across Go Workers and AWS Queues

Image
Distributed tracing is straightforward for HTTP calls. A request comes in, middleware starts a span, headers propagate to the next service, and the trace forms a nice chain. Queues break that chain unless you explicitly carry trace context through the message. For systems built with Go workers, SQS, EventBridge, and background jobs, trace context propagation is the difference between seeing a complete workflow and seeing disconnected islands. Put Trace Context in Message Attributes Do not hide trace metadata inside business payloads. Use message attributes when the transport supports them. For SQS, the W3C `traceparent` header can be stored as an attribute and extracted by the consumer. func addTraceAttributes(ctx context.Context, attrs map[string]types.MessageAttributeValue) { carrier := propagation.MapCarrier{} otel.GetTextMapPropagator().Inject(ctx, carrier) for k, v := range carrier { attrs[k] = types.MessageAttributeValue{ DataType: aws.Str...

DynamoDB Single-Table Design: Patterns for High-Throughput Go Services

Image
DynamoDB looks simple until you design your first table wrong and spend a week refactoring. The first time I used DynamoDB, I modelled it like a relational database — one table per entity type, with natural primary keys. It worked fine at low traffic, then became a mess of expensive scans as usage grew. Learning to think in DynamoDB's model — access patterns first, everything else second — was one of the more valuable architectural shifts in my career. The Fundamental Mental Shift In relational databases you normalise first and query later. SQL's query planner can handle most access patterns efficiently as long as you have reasonable indexes. In DynamoDB, there is no query planner. Every query you want to make must be anticipated in the key design. Design for access patterns first; everything else is secondary. Before touching the DynamoDB console, write down every query your application needs to make. For our Amazon Ads management platform, this looked like: Get all c...

ECS Worker Autoscaling with Queue Depth and Lag Metrics

Image
CPU-based autoscaling works well for web services. It works poorly for queue workers. A worker can be at 20% CPU and still be dangerously behind because the queue is receiving messages faster than it can process them. For SQS workers on ECS, the better scaling signal is backlog per task and message age. The Metric That Matters The metric I start with is backlog per running task. If there are 20,000 visible messages and 20 ECS tasks, each task effectively owns 1,000 messages. If the processing rate is known, that number can be translated into expected drain time. backlog_per_task = visible_messages / max(running_tasks, 1) For workloads with variable processing time, combine it with approximate age of oldest message. Queue depth tells you how much work exists. Age tells you whether users are waiting too long. Scaling Policy Shape A simple target tracking policy can work, but I prefer step scaling for important worker pools because it lets you react aggressively when lag is high an...

Event-Driven Architecture on AWS: SQS, EventBridge and Idempotency

Image
Event-driven architecture sounds clean in diagrams: one service publishes an event, another service reacts, and the system becomes nicely decoupled. In production it is messier. Events arrive late, arrive twice, arrive out of order, or fail halfway through a workflow. AWS gives you strong building blocks, but the architecture still depends on how you handle those realities. Use EventBridge for Routing, SQS for Work The pattern I like is EventBridge for routing and SQS for durable work queues. EventBridge is good at publishing domain events and letting consumers subscribe without tight coupling. SQS is good at giving workers a queue they can drain, retry, and monitor. { "source": "ads.campaigns", "detail-type": "CampaignBudgetChanged", "detail": { "companyId": 5000, "profileId": 50000100, "campaignId": 123456789, "oldBudget": 50.00, "newBudget": 75.00 } } ...

Amazon Marketing Stream: Real-Time Ad Data Without Batch Processing

Image
For three years, our reporting pipeline at the platform ran on nightly batch jobs. Every night at 2am, a fleet of Go workers would call the Amazon Advertising API to pull the previous day's campaign performance data for 2,000+ advertisers. It worked. Until it did not scale. The problems accumulated over time: API rate limits tightened, some advertisers grew to hundreds of campaigns making their nightly sync take 20 minutes, and any API downtime meant the entire previous day's data was missing. Most importantly, our bidding algorithms were always working with data that was at least 24 hours stale. Amazon Marketing Stream changed all of this. What Is Amazon Marketing Stream? Amazon Marketing Stream (AMS) is Amazon's event-driven reporting system that delivers advertising performance data as near-real-time events via SNS → SQS. Instead of you pulling data from Amazon, Amazon pushes data to you as events — impressions, clicks, spend, conversions — typically 3–5 hours afte...

ECS Fargate Autoscaling: How We Cut Infrastructure Costs by 35%

Image
When I joined the company the backend ran on a fleet of EC2 instances sized for peak traffic, sitting at 15% CPU utilisation most of the time. Scaling was manual — someone would notice latency going up, SSH into a box to check what was happening, then provision more capacity if needed. Deployments required coordination to drain the load balancer and restart services one by one. Migrating to ECS Fargate with autoscaling was the single biggest infrastructure improvement we made: costs dropped 35%, deployments became zero-downtime, and on-call became less stressful. Why ECS Fargate Over EC2-Backed ECS ECS can run on two launch types: EC2 (you manage the instances) and Fargate (AWS manages the compute). I chose Fargate for three reasons: No instance management : no more AMI updates, no instance type selection, no patching Bin packing is AWS's problem : with EC2-backed ECS you need the right EC2 instance size to fit your tasks efficiently. Fargate handles this transparently. Pe...

SQS Dead-Letter Queues and Visibility Timeout: What I Got Wrong

Image
I spent an embarrassing amount of time debugging an SQS consumer that was processing messages twice and sometimes three times. The root cause turned out to be a fundamental misunderstanding of how visibility timeout, dead-letter queues, and consumer concurrency interact. Here is everything I wish I had known from the start. SQS Delivery Guarantees: At-Least-Once, Not Exactly-Once Before we dive into configuration, the most important thing to internalise: SQS guarantees at-least-once delivery . The same message may be delivered multiple times, even under normal operating conditions — not just during failures. This is not a bug; it is a deliberate design decision that enables SQS to provide high availability and scalability. Your consumer must be idempotent. Visibility Timeout: The Most Misunderstood Setting When you call ReceiveMessage , SQS makes each returned message invisible to other consumers for the visibility timeout duration. If you do not delete the message before the ti...