Posts

Showing posts with the label Go

Designing Backfill Jobs That Do Not Take Production Down

Image
Backfills are deceptively dangerous. The code is often simple: read old rows, compute a missing value, write it back. The danger is scale. A job that behaves perfectly on ten thousand rows can overload a database, fill a queue, or starve production traffic when it runs across hundreds of millions of records. A production-safe backfill is designed like a service: observable, resumable, throttled, and boring to stop. Make Progress Durable Do not rely on an in-memory cursor for long backfills. Store progress in a table so the job can resume after deploys, crashes, or manual pauses. type BackfillCheckpoint struct { JobName string LastID int64 UpdatedAt time.Time } func (b *Backfill) Run(ctx context.Context) error { checkpoint, err := b.store.LoadCheckpoint(ctx, "campaign-currency") if err != nil { return err } return b.processFrom(ctx, checkpoint.LastID) } Chunk Everything Large transactions are the enemy. Process small batches, co...

Amazon Ads Bulk Operations: Designing for Partial Failure

Image
Bulk operations are where clean API abstractions go to suffer. Updating one campaign budget is simple. Updating ten thousand bids across hundreds of advertiser profiles is a different system. Some updates succeed, some fail validation, some hit rate limits, some time out, and the product still needs to tell the user exactly what happened. The main design principle is to treat partial failure as the normal case. If the code assumes all-or-nothing success, the first real advertiser account will break the workflow. Represent Work Explicitly A bulk operation should become a durable job with child items. Each item has its own status, request payload, response payload, retry count, and error message. This makes the operation resumable and auditable. type BulkItemStatus string const ( ItemPending BulkItemStatus = "PENDING" ItemRunning BulkItemStatus = "RUNNING" ItemSucceeded BulkItemStatus = "SUCCEEDED" ItemFailed BulkItemStatus = ...

Amazon Ads API at Scale: Rate Limiting, Pagination and Bulk Operations in Go

Image
After three years of building and maintaining the platform — a platform that manages Amazon advertising campaigns for thousands of advertisers — I have made every mistake possible with the Amazon Ads API. This post is a practical guide to operating the API at scale: how to stay within rate limits across thousands of advertiser profiles, how to paginate correctly, and how to bulk-process operations without hammering the API into returning 429s. The Scale Problem When you have one advertiser, the Amazon Ads API is straightforward. When you have 2,000 advertisers, each with dozens of campaigns, hundreds of ad groups, and thousands of keywords, the same operations become an engineering challenge. A nightly sync that takes 3 seconds per advertiser profile takes over an hour across the fleet. Any operation that requires multiple API calls per entity — reading, computing, then writing — multiplies that cost. The constraints you need to design around: Rate limits are per profile (per ...

From Logs to Alerts: SLOs for Go APIs on AWS

Image
Logs are useful after something breaks. SLOs are useful before users start sending screenshots. The shift from log-based debugging to service-level objectives is one of the biggest maturity jumps a backend team can make. For Go APIs on AWS, I like starting with a small set of SLOs that match user pain: availability, latency, and freshness. Everything else can grow from there. Define What Good Means A service-level indicator is the measurement. A service-level objective is the target. For an API, the indicators are usually request success rate and latency. For a data pipeline, freshness matters too. 99.9% of API requests should return non-5xx responses over 30 days. 95% of dashboard requests should complete under 500ms. 95% of reporting data should be less than 15 minutes stale. Instrument at the Edge Measure user-visible behavior at the edge of the service. Handler middleware is a good place for request count, status, and duration. Do not build an SLO from internal function timi...

Cost-Aware LLM Routing: Reducing AI API Bills by 60%

Image
Three months after shipping the AI feature, our Anthropic API bill had grown faster than the revenue it was generating. The naive solution was to reduce usage. The right solution was to use the right model for each task. A cost-aware router that directs simple tasks to cheaper models and complex reasoning to powerful ones reduced our monthly AI spend by 60% while maintaining — and in some cases improving — output quality. The Insight: Not All Tasks Are Equal We were using Claude Opus for everything. Extracting a number from a JSON field does not need the same model as synthesising a 500-word campaign performance narrative. Classifying a keyword into one of five categories does not need the same model as generating a multi-step bid adjustment strategy. Using Opus for classification is like using a Ferrari to go grocery shopping. The Anthropic model family maps naturally to task complexity: Claude Haiku : fast, cheap (~50× cheaper than Opus per token), excellent for structured ex...

Trace Context Propagation Across Go Workers and AWS Queues

Image
Distributed tracing is straightforward for HTTP calls. A request comes in, middleware starts a span, headers propagate to the next service, and the trace forms a nice chain. Queues break that chain unless you explicitly carry trace context through the message. For systems built with Go workers, SQS, EventBridge, and background jobs, trace context propagation is the difference between seeing a complete workflow and seeing disconnected islands. Put Trace Context in Message Attributes Do not hide trace metadata inside business payloads. Use message attributes when the transport supports them. For SQS, the W3C `traceparent` header can be stored as an attribute and extracted by the consumer. func addTraceAttributes(ctx context.Context, attrs map[string]types.MessageAttributeValue) { carrier := propagation.MapCarrier{} otel.GetTextMapPropagator().Inject(ctx, carrier) for k, v := range carrier { attrs[k] = types.MessageAttributeValue{ DataType: aws.Str...

Distributed Tracing in Go with OpenTelemetry

Image
When a request takes 800ms instead of the expected 50ms, distributed tracing tells you exactly which service, which database call, and which line of code is responsible. Without it, debugging latency regressions in a microservices system means reading logs across five services, correlating timestamps by hand, and guessing at causality. I implemented OpenTelemetry across our Go services at the platform and it has changed how we debug production issues. Why OpenTelemetry? OpenTelemetry (OTel) is the CNCF standard for observability instrumentation. The key advantage over vendor-specific SDKs (DataDog tracer, X-Ray SDK, etc.) is portability: you write the instrumentation once and can send it to any compatible backend — Jaeger, Zipkin, Honeycomb, Datadog, Grafana Tempo — by changing an exporter configuration. We started with Jaeger and migrated to Grafana Tempo without touching application code. Setting Up the Tracer Provider func InitTracing(ctx context.Context, cfg TracingConfig) (...

LLM Evaluation Harnesses in Go: Shipping AI Features Safely

Image
The first version of an AI feature is usually judged by vibes. You run twenty examples, the output looks good, and everyone gets excited. The problem is that vibes do not survive production. Prompts change, models change, input data changes, and suddenly the feature starts producing recommendations that are plausible but wrong. An evaluation harness turns AI quality into something you can test before every deploy. It will never be perfect, but it is much better than clicking around manually and hoping the model still behaves. Build A Golden Dataset Start with real inputs from the product, anonymized and reduced to the fields the model actually needs. For each input, store the expected properties of a good answer. Not always the exact output, but the constraints that matter. type EvalCase struct { Name string Input RecommendationInput MustInclude []string MustAvoid []string MaxCostCents int } For campaign recommendations, a case might require the...

Structured Outputs with Claude API: Production Patterns in Go

Image
The difference between a demo LLM integration and a production one often comes down to structured outputs. In a demo, free-form text is fine — you are showing a human-readable result. In production, you need to reliably parse the response into typed data structures, validate it, handle failures gracefully, and integrate it into downstream systems that expect specific types. This post covers the patterns that have worked in our Go services at the platform. Why Free-Form Text Fails in Production LLMs are probabilistic. Even with a deterministic system prompt, the same input can produce slightly different output formats across calls. "Return the ACOS as a number" might sometimes produce 23.5 , sometimes 23.5% , sometimes "ACOS: 23.5%" . Any of these can happen, and your production system must handle all of them or crash. Structured outputs — combined with JSON schema validation — eliminate this class of problem. Instead of parsing the LLM response as free text, y...

DynamoDB Single-Table Design: Patterns for High-Throughput Go Services

Image
DynamoDB looks simple until you design your first table wrong and spend a week refactoring. The first time I used DynamoDB, I modelled it like a relational database — one table per entity type, with natural primary keys. It worked fine at low traffic, then became a mess of expensive scans as usage grew. Learning to think in DynamoDB's model — access patterns first, everything else second — was one of the more valuable architectural shifts in my career. The Fundamental Mental Shift In relational databases you normalise first and query later. SQL's query planner can handle most access patterns efficiently as long as you have reasonable indexes. In DynamoDB, there is no query planner. Every query you want to make must be anticipated in the key design. Design for access patterns first; everything else is secondary. Before touching the DynamoDB console, write down every query your application needs to make. For our Amazon Ads management platform, this looked like: Get all c...

Zero-Downtime Schema Changes in Go Services

Image
Database migrations are easy in small applications because deploys are linear. Change the schema, deploy the code, done. In a real production system with multiple Go services, background workers, rolling deploys, and long-running jobs, schema changes need choreography. The safe pattern is expand, migrate, contract. Add the new shape while the old code still works, move traffic gradually, backfill data, then remove the old shape only after every consumer has moved. Step 1: Expand The expand migration only adds things: a nullable column, a new table, a new index, or a trigger. It should be safe to run while old code is still deployed. ALTER TABLE campaigns ADD COLUMN budget_currency VARCHAR(3) NULL; CREATE INDEX CONCURRENTLY idx_campaigns_company_currency ON campaigns (company_id, budget_currency); Avoid migrations that rewrite huge tables during business hours. Even if the database supports online operations, test the migration with realistic data volume before trusting it. Step...

Designing Multi-Tenant Go Services Without Data Leaks

Image
Multi-tenancy is easy to underestimate because the first version is usually just a `company_id` column. Add the column, add an index, filter by it in queries, and move on. That works until the product grows, background jobs are added, exports are introduced, and one missing filter becomes a serious data leak. For a platform that manages advertiser data, tenant isolation is not a nice-to-have. It is a core security boundary. The safest design is the one where the boring default path is also the secure path. Make Tenant Context Explicit I avoid passing raw IDs through twenty function calls. Instead, request-scoped tenant context becomes a first-class value. It contains the company, profile, marketplace, permissions, and any constraints needed by the downstream service. type TenantContext struct { CompanyID int64 ProfileID int64 Country string Roles []string } func TenantFromRequest(r *http.Request) (TenantContext, error) { claims := auth.ClaimsFromContext(...

Retry Budgets in Go Services: Preventing Cascading Failure

Image
Retries are useful until they become the reason your system is down. I have seen this pattern more than once: one dependency gets slower, callers retry aggressively, queue depth grows, CPU jumps, and the service that was already struggling now receives three times the normal traffic. The incident starts as a dependency issue and turns into a self-inflicted denial of service. The fix is not to remove retries. The fix is to give retries a budget. A retry budget makes every caller spend from a limited allowance, so the system can absorb short failures without amplifying long ones. The Failure Mode Imagine a Go API that calls a reporting service. The reporting service usually responds in 80ms, but during a deploy it starts taking 900ms. The API has a one second timeout and retries twice. Every user request can now become three downstream calls, and each one waits almost the full timeout before failing. If the original traffic is 200 requests per second, the downstream service may sudde...

Building a Production LLM Layer in Go

Image
In Q4 2024 we shipped the AI feature — an AI layer that analyses campaign performance data, surfaces anomalies, and generates human-readable recommendations for advertisers. Building it taught me more about production AI systems than any course or blog post. This is the unfiltered account: what worked, what failed, and the architecture we ended up with after several iterations. The Problem We Were Solving Our platform manages advertising campaigns for hundreds of advertisers on Amazon. Each advertiser has dozens of campaigns, hundreds of ad groups, thousands of keywords, and daily performance metrics for all of them. Identifying what needs attention — which keyword bid is too high, which campaign is bleeding budget without converting, which new product launch is outperforming expectations — requires reading a lot of data and making nuanced judgments. We were doing this manually in customer success calls. our AI product was the attempt to automate it. Architecture: Thin LLM Servic...

Idempotency Keys in Distributed Go Services

Image
Distributed systems are fundamentally unreliable. Networks drop packets, services restart mid-request, clients retry on timeout, and load balancers reroute connections. The standard response to this reality is to design every state-modifying operation to be idempotent — safe to call multiple times with the same result as calling it once. Idempotency keys are the primary tool for achieving this at the API level. The Core Problem Consider a client that sends a POST request to create a campaign. The server receives the request, creates the campaign, and then a network failure prevents the response from reaching the client. The client, having received no response, retries the request. Without idempotency, the server creates a second campaign. Now you have two identical campaigns — a data integrity problem that is difficult to detect and painful to clean up. We manage campaigns for thousands of advertisers. A double-creation bug is not just a data integrity issue — it is a budget issu...

Amazon Marketing Stream: Real-Time Ad Data Without Batch Processing

Image
For three years, our reporting pipeline at the platform ran on nightly batch jobs. Every night at 2am, a fleet of Go workers would call the Amazon Advertising API to pull the previous day's campaign performance data for 2,000+ advertisers. It worked. Until it did not scale. The problems accumulated over time: API rate limits tightened, some advertisers grew to hundreds of campaigns making their nightly sync take 20 minutes, and any API downtime meant the entire previous day's data was missing. Most importantly, our bidding algorithms were always working with data that was at least 24 hours stale. Amazon Marketing Stream changed all of this. What Is Amazon Marketing Stream? Amazon Marketing Stream (AMS) is Amazon's event-driven reporting system that delivers advertising performance data as near-real-time events via SNS → SQS. Instead of you pulling data from Amazon, Amazon pushes data to you as events — impressions, clicks, spend, conversions — typically 3–5 hours afte...