Posts

Designing Backfill Jobs That Do Not Take Production Down

Image
Backfills are deceptively dangerous. The code is often simple: read old rows, compute a missing value, write it back. The danger is scale. A job that behaves perfectly on ten thousand rows can overload a database, fill a queue, or starve production traffic when it runs across hundreds of millions of records. A production-safe backfill is designed like a service: observable, resumable, throttled, and boring to stop. Make Progress Durable Do not rely on an in-memory cursor for long backfills. Store progress in a table so the job can resume after deploys, crashes, or manual pauses. type BackfillCheckpoint struct { JobName string LastID int64 UpdatedAt time.Time } func (b *Backfill) Run(ctx context.Context) error { checkpoint, err := b.store.LoadCheckpoint(ctx, "campaign-currency") if err != nil { return err } return b.processFrom(ctx, checkpoint.LastID) } Chunk Everything Large transactions are the enemy. Process small batches, co...

Amazon Ads Bulk Operations: Designing for Partial Failure

Image
Bulk operations are where clean API abstractions go to suffer. Updating one campaign budget is simple. Updating ten thousand bids across hundreds of advertiser profiles is a different system. Some updates succeed, some fail validation, some hit rate limits, some time out, and the product still needs to tell the user exactly what happened. The main design principle is to treat partial failure as the normal case. If the code assumes all-or-nothing success, the first real advertiser account will break the workflow. Represent Work Explicitly A bulk operation should become a durable job with child items. Each item has its own status, request payload, response payload, retry count, and error message. This makes the operation resumable and auditable. type BulkItemStatus string const ( ItemPending BulkItemStatus = "PENDING" ItemRunning BulkItemStatus = "RUNNING" ItemSucceeded BulkItemStatus = "SUCCEEDED" ItemFailed BulkItemStatus = ...

Amazon Ads API at Scale: Rate Limiting, Pagination and Bulk Operations in Go

Image
After three years of building and maintaining the platform — a platform that manages Amazon advertising campaigns for thousands of advertisers — I have made every mistake possible with the Amazon Ads API. This post is a practical guide to operating the API at scale: how to stay within rate limits across thousands of advertiser profiles, how to paginate correctly, and how to bulk-process operations without hammering the API into returning 429s. The Scale Problem When you have one advertiser, the Amazon Ads API is straightforward. When you have 2,000 advertisers, each with dozens of campaigns, hundreds of ad groups, and thousands of keywords, the same operations become an engineering challenge. A nightly sync that takes 3 seconds per advertiser profile takes over an hour across the fleet. Any operation that requires multiple API calls per entity — reading, computing, then writing — multiplies that cost. The constraints you need to design around: Rate limits are per profile (per ...

From Logs to Alerts: SLOs for Go APIs on AWS

Image
Logs are useful after something breaks. SLOs are useful before users start sending screenshots. The shift from log-based debugging to service-level objectives is one of the biggest maturity jumps a backend team can make. For Go APIs on AWS, I like starting with a small set of SLOs that match user pain: availability, latency, and freshness. Everything else can grow from there. Define What Good Means A service-level indicator is the measurement. A service-level objective is the target. For an API, the indicators are usually request success rate and latency. For a data pipeline, freshness matters too. 99.9% of API requests should return non-5xx responses over 30 days. 95% of dashboard requests should complete under 500ms. 95% of reporting data should be less than 15 minutes stale. Instrument at the Edge Measure user-visible behavior at the edge of the service. Handler middleware is a good place for request count, status, and duration. Do not build an SLO from internal function timi...

Cost-Aware LLM Routing: Reducing AI API Bills by 60%

Image
Three months after shipping the AI feature, our Anthropic API bill had grown faster than the revenue it was generating. The naive solution was to reduce usage. The right solution was to use the right model for each task. A cost-aware router that directs simple tasks to cheaper models and complex reasoning to powerful ones reduced our monthly AI spend by 60% while maintaining — and in some cases improving — output quality. The Insight: Not All Tasks Are Equal We were using Claude Opus for everything. Extracting a number from a JSON field does not need the same model as synthesising a 500-word campaign performance narrative. Classifying a keyword into one of five categories does not need the same model as generating a multi-step bid adjustment strategy. Using Opus for classification is like using a Ferrari to go grocery shopping. The Anthropic model family maps naturally to task complexity: Claude Haiku : fast, cheap (~50× cheaper than Opus per token), excellent for structured ex...

Trace Context Propagation Across Go Workers and AWS Queues

Image
Distributed tracing is straightforward for HTTP calls. A request comes in, middleware starts a span, headers propagate to the next service, and the trace forms a nice chain. Queues break that chain unless you explicitly carry trace context through the message. For systems built with Go workers, SQS, EventBridge, and background jobs, trace context propagation is the difference between seeing a complete workflow and seeing disconnected islands. Put Trace Context in Message Attributes Do not hide trace metadata inside business payloads. Use message attributes when the transport supports them. For SQS, the W3C `traceparent` header can be stored as an attribute and extracted by the consumer. func addTraceAttributes(ctx context.Context, attrs map[string]types.MessageAttributeValue) { carrier := propagation.MapCarrier{} otel.GetTextMapPropagator().Inject(ctx, carrier) for k, v := range carrier { attrs[k] = types.MessageAttributeValue{ DataType: aws.Str...

Distributed Tracing in Go with OpenTelemetry

Image
When a request takes 800ms instead of the expected 50ms, distributed tracing tells you exactly which service, which database call, and which line of code is responsible. Without it, debugging latency regressions in a microservices system means reading logs across five services, correlating timestamps by hand, and guessing at causality. I implemented OpenTelemetry across our Go services at the platform and it has changed how we debug production issues. Why OpenTelemetry? OpenTelemetry (OTel) is the CNCF standard for observability instrumentation. The key advantage over vendor-specific SDKs (DataDog tracer, X-Ray SDK, etc.) is portability: you write the instrumentation once and can send it to any compatible backend — Jaeger, Zipkin, Honeycomb, Datadog, Grafana Tempo — by changing an exporter configuration. We started with Jaeger and migrated to Grafana Tempo without touching application code. Setting Up the Tracer Provider func InitTracing(ctx context.Context, cfg TracingConfig) (...

LLM Evaluation Harnesses in Go: Shipping AI Features Safely

Image
The first version of an AI feature is usually judged by vibes. You run twenty examples, the output looks good, and everyone gets excited. The problem is that vibes do not survive production. Prompts change, models change, input data changes, and suddenly the feature starts producing recommendations that are plausible but wrong. An evaluation harness turns AI quality into something you can test before every deploy. It will never be perfect, but it is much better than clicking around manually and hoping the model still behaves. Build A Golden Dataset Start with real inputs from the product, anonymized and reduced to the fields the model actually needs. For each input, store the expected properties of a good answer. Not always the exact output, but the constraints that matter. type EvalCase struct { Name string Input RecommendationInput MustInclude []string MustAvoid []string MaxCostCents int } For campaign recommendations, a case might require the...

Structured Outputs with Claude API: Production Patterns in Go

Image
The difference between a demo LLM integration and a production one often comes down to structured outputs. In a demo, free-form text is fine — you are showing a human-readable result. In production, you need to reliably parse the response into typed data structures, validate it, handle failures gracefully, and integrate it into downstream systems that expect specific types. This post covers the patterns that have worked in our Go services at the platform. Why Free-Form Text Fails in Production LLMs are probabilistic. Even with a deterministic system prompt, the same input can produce slightly different output formats across calls. "Return the ACOS as a number" might sometimes produce 23.5 , sometimes 23.5% , sometimes "ACOS: 23.5%" . Any of these can happen, and your production system must handle all of them or crash. Structured outputs — combined with JSON schema validation — eliminate this class of problem. Instead of parsing the LLM response as free text, y...

DynamoDB Single-Table Design: Patterns for High-Throughput Go Services

Image
DynamoDB looks simple until you design your first table wrong and spend a week refactoring. The first time I used DynamoDB, I modelled it like a relational database — one table per entity type, with natural primary keys. It worked fine at low traffic, then became a mess of expensive scans as usage grew. Learning to think in DynamoDB's model — access patterns first, everything else second — was one of the more valuable architectural shifts in my career. The Fundamental Mental Shift In relational databases you normalise first and query later. SQL's query planner can handle most access patterns efficiently as long as you have reasonable indexes. In DynamoDB, there is no query planner. Every query you want to make must be anticipated in the key design. Design for access patterns first; everything else is secondary. Before touching the DynamoDB console, write down every query your application needs to make. For our Amazon Ads management platform, this looked like: Get all c...

Amazon Ads API v1: A New Unified Approach — Notes from a New York Meetup

Image
Last month I was in New York for a series of meetings that included a tech gathering where several Amazon Ads engineers were presenting the direction of their advertising API. I had been aware of the "Amazon Ads API v1" project for a while but had not fully understood its scope. After that evening, I left with a clear picture of what it is, why it matters, and the migration work we need to plan at the platform. Context: The Problem With the Current API Landscape If you have built tools on top of the Amazon Advertising API, you know the pain. Sponsored Products, Sponsored Brands, Sponsored Display, and DSP are all separate product lines — and historically they each have their own API surface, their own endpoint naming conventions, their own request/response shapes, and their own error formats. Want to create a campaign? The Sponsored Products endpoint and the Sponsored Display endpoint have different request schemas. Want to list ad groups? Different pagination implement...

Caching Strategies for Amazon Ads Dashboards

Image
Advertising dashboards are read-heavy, bursty, and expensive to compute. A single page can ask for spend, sales, ACOS, ROAS, campaign status, budget pacing, placement breakdowns, and search-term trends. Without caching, the database becomes the place where every product decision is paid for repeatedly. The hard part is not adding Redis. The hard part is deciding what can be cached, for how long, and how to invalidate it when advertisers expect fresh numbers. Cache Data Products, Not SQL Rows A common mistake is caching low-level query results. That leaks implementation details into the cache and makes invalidation painful. I prefer caching data products: the exact response shape used by the dashboard card or API endpoint. type DashboardCacheKey struct { CompanyID int64 ProfileID int64 DateRange string Marketplace string Card string Version int } The `Version` field is important. When the calculation changes, bump the version and old entrie...

Zero-Downtime Schema Changes in Go Services

Image
Database migrations are easy in small applications because deploys are linear. Change the schema, deploy the code, done. In a real production system with multiple Go services, background workers, rolling deploys, and long-running jobs, schema changes need choreography. The safe pattern is expand, migrate, contract. Add the new shape while the old code still works, move traffic gradually, backfill data, then remove the old shape only after every consumer has moved. Step 1: Expand The expand migration only adds things: a nullable column, a new table, a new index, or a trigger. It should be safe to run while old code is still deployed. ALTER TABLE campaigns ADD COLUMN budget_currency VARCHAR(3) NULL; CREATE INDEX CONCURRENTLY idx_campaigns_company_currency ON campaigns (company_id, budget_currency); Avoid migrations that rewrite huge tables during business hours. Even if the database supports online operations, test the migration with realistic data volume before trusting it. Step...

ECS Worker Autoscaling with Queue Depth and Lag Metrics

Image
CPU-based autoscaling works well for web services. It works poorly for queue workers. A worker can be at 20% CPU and still be dangerously behind because the queue is receiving messages faster than it can process them. For SQS workers on ECS, the better scaling signal is backlog per task and message age. The Metric That Matters The metric I start with is backlog per running task. If there are 20,000 visible messages and 20 ECS tasks, each task effectively owns 1,000 messages. If the processing rate is known, that number can be translated into expected drain time. backlog_per_task = visible_messages / max(running_tasks, 1) For workloads with variable processing time, combine it with approximate age of oldest message. Queue depth tells you how much work exists. Age tells you whether users are waiting too long. Scaling Policy Shape A simple target tracking policy can work, but I prefer step scaling for important worker pools because it lets you react aggressively when lag is high an...

Event-Driven Architecture on AWS: SQS, EventBridge and Idempotency

Image
Event-driven architecture sounds clean in diagrams: one service publishes an event, another service reacts, and the system becomes nicely decoupled. In production it is messier. Events arrive late, arrive twice, arrive out of order, or fail halfway through a workflow. AWS gives you strong building blocks, but the architecture still depends on how you handle those realities. Use EventBridge for Routing, SQS for Work The pattern I like is EventBridge for routing and SQS for durable work queues. EventBridge is good at publishing domain events and letting consumers subscribe without tight coupling. SQS is good at giving workers a queue they can drain, retry, and monitor. { "source": "ads.campaigns", "detail-type": "CampaignBudgetChanged", "detail": { "companyId": 5000, "profileId": 50000100, "campaignId": 123456789, "oldBudget": 50.00, "newBudget": 75.00 } } ...

ClickHouse for Advertising Analytics: Lessons from High-Cardinality Data

Image
Advertising analytics is a perfect stress test for databases. The data looks simple at first: date, campaign, ad group, keyword, spend, clicks, sales. Then you add thousands of advertisers, multiple marketplaces, placement breakdowns, search terms, hourly metrics, attribution windows, and suddenly every dashboard query has high-cardinality dimensions. ClickHouse is very good at this workload, but only when the table design respects how ClickHouse reads data. The first version of a schema can feel fast in development and become painful once the real cardinality arrives. Model Around Query Patterns For analytics tables, I start from the dashboards and API endpoints. Which dimensions are always filtered? Which dimensions are grouped? Which time ranges are common? The answers drive partitioning, ordering, and materialized views. CREATE TABLE campaign_daily_metrics ( event_date Date, company_id UInt64, profile_id UInt64, campaign_id UInt64, marketplace LowCardinalit...

Designing Multi-Tenant Go Services Without Data Leaks

Image
Multi-tenancy is easy to underestimate because the first version is usually just a `company_id` column. Add the column, add an index, filter by it in queries, and move on. That works until the product grows, background jobs are added, exports are introduced, and one missing filter becomes a serious data leak. For a platform that manages advertiser data, tenant isolation is not a nice-to-have. It is a core security boundary. The safest design is the one where the boring default path is also the secure path. Make Tenant Context Explicit I avoid passing raw IDs through twenty function calls. Instead, request-scoped tenant context becomes a first-class value. It contains the company, profile, marketplace, permissions, and any constraints needed by the downstream service. type TenantContext struct { CompanyID int64 ProfileID int64 Country string Roles []string } func TenantFromRequest(r *http.Request) (TenantContext, error) { claims := auth.ClaimsFromContext(...

Retry Budgets in Go Services: Preventing Cascading Failure

Image
Retries are useful until they become the reason your system is down. I have seen this pattern more than once: one dependency gets slower, callers retry aggressively, queue depth grows, CPU jumps, and the service that was already struggling now receives three times the normal traffic. The incident starts as a dependency issue and turns into a self-inflicted denial of service. The fix is not to remove retries. The fix is to give retries a budget. A retry budget makes every caller spend from a limited allowance, so the system can absorb short failures without amplifying long ones. The Failure Mode Imagine a Go API that calls a reporting service. The reporting service usually responds in 80ms, but during a deploy it starts taking 900ms. The API has a one second timeout and retries twice. Every user request can now become three downstream calls, and each one waits almost the full timeout before failing. If the original traffic is 200 requests per second, the downstream service may sudde...

Building a Production LLM Layer in Go

Image
In Q4 2024 we shipped the AI feature — an AI layer that analyses campaign performance data, surfaces anomalies, and generates human-readable recommendations for advertisers. Building it taught me more about production AI systems than any course or blog post. This is the unfiltered account: what worked, what failed, and the architecture we ended up with after several iterations. The Problem We Were Solving Our platform manages advertising campaigns for hundreds of advertisers on Amazon. Each advertiser has dozens of campaigns, hundreds of ad groups, thousands of keywords, and daily performance metrics for all of them. Identifying what needs attention — which keyword bid is too high, which campaign is bleeding budget without converting, which new product launch is outperforming expectations — requires reading a lot of data and making nuanced judgments. We were doing this manually in customer success calls. our AI product was the attempt to automate it. Architecture: Thin LLM Servic...