Posts

Trace Context Propagation Across Go Workers and AWS Queues

Image
Distributed tracing is straightforward for HTTP calls. A request comes in, middleware starts a span, headers propagate to the next service, and the trace forms a nice chain. Queues break that chain unless you explicitly carry trace context through the message. For systems built with Go workers, SQS, EventBridge, and background jobs, trace context propagation is the difference between seeing a complete workflow and seeing disconnected islands. Put Trace Context in Message Attributes Do not hide trace metadata inside business payloads. Use message attributes when the transport supports them. For SQS, the W3C `traceparent` header can be stored as an attribute and extracted by the consumer. func addTraceAttributes(ctx context.Context, attrs map[string]types.MessageAttributeValue) { carrier := propagation.MapCarrier{} otel.GetTextMapPropagator().Inject(ctx, carrier) for k, v := range carrier { attrs[k] = types.MessageAttributeValue{ DataType: aws.Str...

Distributed Tracing in Go with OpenTelemetry

Image
When a request takes 800ms instead of the expected 50ms, distributed tracing tells you exactly which service, which database call, and which line of code is responsible. Without it, debugging latency regressions in a microservices system means reading logs across five services, correlating timestamps by hand, and guessing at causality. I implemented OpenTelemetry across our Go services at the platform and it has changed how we debug production issues. Why OpenTelemetry? OpenTelemetry (OTel) is the CNCF standard for observability instrumentation. The key advantage over vendor-specific SDKs (DataDog tracer, X-Ray SDK, etc.) is portability: you write the instrumentation once and can send it to any compatible backend — Jaeger, Zipkin, Honeycomb, Datadog, Grafana Tempo — by changing an exporter configuration. We started with Jaeger and migrated to Grafana Tempo without touching application code. Setting Up the Tracer Provider func InitTracing(ctx context.Context, cfg TracingConfig) (...

LLM Evaluation Harnesses in Go: Shipping AI Features Safely

Image
The first version of an AI feature is usually judged by vibes. You run twenty examples, the output looks good, and everyone gets excited. The problem is that vibes do not survive production. Prompts change, models change, input data changes, and suddenly the feature starts producing recommendations that are plausible but wrong. An evaluation harness turns AI quality into something you can test before every deploy. It will never be perfect, but it is much better than clicking around manually and hoping the model still behaves. Build A Golden Dataset Start with real inputs from the product, anonymized and reduced to the fields the model actually needs. For each input, store the expected properties of a good answer. Not always the exact output, but the constraints that matter. type EvalCase struct { Name string Input RecommendationInput MustInclude []string MustAvoid []string MaxCostCents int } For campaign recommendations, a case might require the...

Structured Outputs with Claude API: Production Patterns in Go

Image
The difference between a demo LLM integration and a production one often comes down to structured outputs. In a demo, free-form text is fine — you are showing a human-readable result. In production, you need to reliably parse the response into typed data structures, validate it, handle failures gracefully, and integrate it into downstream systems that expect specific types. This post covers the patterns that have worked in our Go services at the platform. Why Free-Form Text Fails in Production LLMs are probabilistic. Even with a deterministic system prompt, the same input can produce slightly different output formats across calls. "Return the ACOS as a number" might sometimes produce 23.5 , sometimes 23.5% , sometimes "ACOS: 23.5%" . Any of these can happen, and your production system must handle all of them or crash. Structured outputs — combined with JSON schema validation — eliminate this class of problem. Instead of parsing the LLM response as free text, y...

DynamoDB Single-Table Design: Patterns for High-Throughput Go Services

Image
DynamoDB looks simple until you design your first table wrong and spend a week refactoring. The first time I used DynamoDB, I modelled it like a relational database — one table per entity type, with natural primary keys. It worked fine at low traffic, then became a mess of expensive scans as usage grew. Learning to think in DynamoDB's model — access patterns first, everything else second — was one of the more valuable architectural shifts in my career. The Fundamental Mental Shift In relational databases you normalise first and query later. SQL's query planner can handle most access patterns efficiently as long as you have reasonable indexes. In DynamoDB, there is no query planner. Every query you want to make must be anticipated in the key design. Design for access patterns first; everything else is secondary. Before touching the DynamoDB console, write down every query your application needs to make. For our Amazon Ads management platform, this looked like: Get all c...

Amazon Ads API v1: A New Unified Approach — Notes from a New York Meetup

Image
Last month I was in New York for a series of meetings that included a tech gathering where several Amazon Ads engineers were presenting the direction of their advertising API. I had been aware of the "Amazon Ads API v1" project for a while but had not fully understood its scope. After that evening, I left with a clear picture of what it is, why it matters, and the migration work we need to plan at the platform. Context: The Problem With the Current API Landscape If you have built tools on top of the Amazon Advertising API, you know the pain. Sponsored Products, Sponsored Brands, Sponsored Display, and DSP are all separate product lines — and historically they each have their own API surface, their own endpoint naming conventions, their own request/response shapes, and their own error formats. Want to create a campaign? The Sponsored Products endpoint and the Sponsored Display endpoint have different request schemas. Want to list ad groups? Different pagination implement...

Caching Strategies for Amazon Ads Dashboards

Image
Advertising dashboards are read-heavy, bursty, and expensive to compute. A single page can ask for spend, sales, ACOS, ROAS, campaign status, budget pacing, placement breakdowns, and search-term trends. Without caching, the database becomes the place where every product decision is paid for repeatedly. The hard part is not adding Redis. The hard part is deciding what can be cached, for how long, and how to invalidate it when advertisers expect fresh numbers. Cache Data Products, Not SQL Rows A common mistake is caching low-level query results. That leaks implementation details into the cache and makes invalidation painful. I prefer caching data products: the exact response shape used by the dashboard card or API endpoint. type DashboardCacheKey struct { CompanyID int64 ProfileID int64 DateRange string Marketplace string Card string Version int } The `Version` field is important. When the calculation changes, bump the version and old entrie...