Posts

Showing posts with the label ECS

ECS Worker Autoscaling with Queue Depth and Lag Metrics

Image
CPU-based autoscaling works well for web services. It works poorly for queue workers. A worker can be at 20% CPU and still be dangerously behind because the queue is receiving messages faster than it can process them. For SQS workers on ECS, the better scaling signal is backlog per task and message age. The Metric That Matters The metric I start with is backlog per running task. If there are 20,000 visible messages and 20 ECS tasks, each task effectively owns 1,000 messages. If the processing rate is known, that number can be translated into expected drain time. backlog_per_task = visible_messages / max(running_tasks, 1) For workloads with variable processing time, combine it with approximate age of oldest message. Queue depth tells you how much work exists. Age tells you whether users are waiting too long. Scaling Policy Shape A simple target tracking policy can work, but I prefer step scaling for important worker pools because it lets you react aggressively when lag is high an...

ECS Fargate Autoscaling: How We Cut Infrastructure Costs by 35%

Image
When I joined the company the backend ran on a fleet of EC2 instances sized for peak traffic, sitting at 15% CPU utilisation most of the time. Scaling was manual — someone would notice latency going up, SSH into a box to check what was happening, then provision more capacity if needed. Deployments required coordination to drain the load balancer and restart services one by one. Migrating to ECS Fargate with autoscaling was the single biggest infrastructure improvement we made: costs dropped 35%, deployments became zero-downtime, and on-call became less stressful. Why ECS Fargate Over EC2-Backed ECS ECS can run on two launch types: EC2 (you manage the instances) and Fargate (AWS manages the compute). I chose Fargate for three reasons: No instance management : no more AMI updates, no instance type selection, no patching Bin packing is AWS's problem : with EC2-backed ECS you need the right EC2 instance size to fit your tasks efficiently. Fargate handles this transparently. Pe...