Posts

Amazon Ads API v1: A New Unified Approach — Notes from a New York Meetup

Image
Last month I was in New York for a series of meetings that included a tech gathering where several Amazon Ads engineers were presenting the direction of their advertising API. I had been aware of the "Amazon Ads API v1" project for a while but had not fully understood its scope. After that evening, I left with a clear picture of what it is, why it matters, and the migration work we need to plan at the platform. Context: The Problem With the Current API Landscape If you have built tools on top of the Amazon Advertising API, you know the pain. Sponsored Products, Sponsored Brands, Sponsored Display, and DSP are all separate product lines — and historically they each have their own API surface, their own endpoint naming conventions, their own request/response shapes, and their own error formats. Want to create a campaign? The Sponsored Products endpoint and the Sponsored Display endpoint have different request schemas. Want to list ad groups? Different pagination implement...

Caching Strategies for Amazon Ads Dashboards

Image
Advertising dashboards are read-heavy, bursty, and expensive to compute. A single page can ask for spend, sales, ACOS, ROAS, campaign status, budget pacing, placement breakdowns, and search-term trends. Without caching, the database becomes the place where every product decision is paid for repeatedly. The hard part is not adding Redis. The hard part is deciding what can be cached, for how long, and how to invalidate it when advertisers expect fresh numbers. Cache Data Products, Not SQL Rows A common mistake is caching low-level query results. That leaks implementation details into the cache and makes invalidation painful. I prefer caching data products: the exact response shape used by the dashboard card or API endpoint. type DashboardCacheKey struct { CompanyID int64 ProfileID int64 DateRange string Marketplace string Card string Version int } The `Version` field is important. When the calculation changes, bump the version and old entrie...

Zero-Downtime Schema Changes in Go Services

Image
Database migrations are easy in small applications because deploys are linear. Change the schema, deploy the code, done. In a real production system with multiple Go services, background workers, rolling deploys, and long-running jobs, schema changes need choreography. The safe pattern is expand, migrate, contract. Add the new shape while the old code still works, move traffic gradually, backfill data, then remove the old shape only after every consumer has moved. Step 1: Expand The expand migration only adds things: a nullable column, a new table, a new index, or a trigger. It should be safe to run while old code is still deployed. ALTER TABLE campaigns ADD COLUMN budget_currency VARCHAR(3) NULL; CREATE INDEX CONCURRENTLY idx_campaigns_company_currency ON campaigns (company_id, budget_currency); Avoid migrations that rewrite huge tables during business hours. Even if the database supports online operations, test the migration with realistic data volume before trusting it. Step...

ECS Worker Autoscaling with Queue Depth and Lag Metrics

Image
CPU-based autoscaling works well for web services. It works poorly for queue workers. A worker can be at 20% CPU and still be dangerously behind because the queue is receiving messages faster than it can process them. For SQS workers on ECS, the better scaling signal is backlog per task and message age. The Metric That Matters The metric I start with is backlog per running task. If there are 20,000 visible messages and 20 ECS tasks, each task effectively owns 1,000 messages. If the processing rate is known, that number can be translated into expected drain time. backlog_per_task = visible_messages / max(running_tasks, 1) For workloads with variable processing time, combine it with approximate age of oldest message. Queue depth tells you how much work exists. Age tells you whether users are waiting too long. Scaling Policy Shape A simple target tracking policy can work, but I prefer step scaling for important worker pools because it lets you react aggressively when lag is high an...

Event-Driven Architecture on AWS: SQS, EventBridge and Idempotency

Image
Event-driven architecture sounds clean in diagrams: one service publishes an event, another service reacts, and the system becomes nicely decoupled. In production it is messier. Events arrive late, arrive twice, arrive out of order, or fail halfway through a workflow. AWS gives you strong building blocks, but the architecture still depends on how you handle those realities. Use EventBridge for Routing, SQS for Work The pattern I like is EventBridge for routing and SQS for durable work queues. EventBridge is good at publishing domain events and letting consumers subscribe without tight coupling. SQS is good at giving workers a queue they can drain, retry, and monitor. { "source": "ads.campaigns", "detail-type": "CampaignBudgetChanged", "detail": { "companyId": 5000, "profileId": 50000100, "campaignId": 123456789, "oldBudget": 50.00, "newBudget": 75.00 } } ...

ClickHouse for Advertising Analytics: Lessons from High-Cardinality Data

Image
Advertising analytics is a perfect stress test for databases. The data looks simple at first: date, campaign, ad group, keyword, spend, clicks, sales. Then you add thousands of advertisers, multiple marketplaces, placement breakdowns, search terms, hourly metrics, attribution windows, and suddenly every dashboard query has high-cardinality dimensions. ClickHouse is very good at this workload, but only when the table design respects how ClickHouse reads data. The first version of a schema can feel fast in development and become painful once the real cardinality arrives. Model Around Query Patterns For analytics tables, I start from the dashboards and API endpoints. Which dimensions are always filtered? Which dimensions are grouped? Which time ranges are common? The answers drive partitioning, ordering, and materialized views. CREATE TABLE campaign_daily_metrics ( event_date Date, company_id UInt64, profile_id UInt64, campaign_id UInt64, marketplace LowCardinalit...

Designing Multi-Tenant Go Services Without Data Leaks

Image
Multi-tenancy is easy to underestimate because the first version is usually just a `company_id` column. Add the column, add an index, filter by it in queries, and move on. That works until the product grows, background jobs are added, exports are introduced, and one missing filter becomes a serious data leak. For a platform that manages advertiser data, tenant isolation is not a nice-to-have. It is a core security boundary. The safest design is the one where the boring default path is also the secure path. Make Tenant Context Explicit I avoid passing raw IDs through twenty function calls. Instead, request-scoped tenant context becomes a first-class value. It contains the company, profile, marketplace, permissions, and any constraints needed by the downstream service. type TenantContext struct { CompanyID int64 ProfileID int64 Country string Roles []string } func TenantFromRequest(r *http.Request) (TenantContext, error) { claims := auth.ClaimsFromContext(...