OpenTelemetry Usage Metering Engine

go services that turn opentelemetry traces, logs and metrics into billable usage, published exactly once and reconciled against the invoice.

OpenTelemetry Usage Metering Engine
role
sole author
stack
Go · PostgreSQL 17 · ClickHouse · Kubernetes · Helm · React 19 · TanStack · OpenTelemetry

The problem

Usage-based billing sounds simple: count what customers use, multiply by a price. In practice the counting has to stay correct through retries, late data and partial failures, because every bug is either lost revenue or an angry customer. This engine is a personal build that works through those problems properly, end to end.

How data flows

  1. An ingest gateway receives OTLP data. It forwards to storage first, then records usage, then acknowledges. The worst failure is "stored but not billed", never "billed but not stored".
  2. Every batch is deduplicated over a 6-minute window, which covers the OpenTelemetry Collector's retry span.
  3. A rollup worker closes each hour and writes the hourly totals and their outgoing billing events in one transaction.
  4. A publisher sends those events to a Stripe-compatible billing provider.
  5. A reconciler compares the ledger, the outbox, the provider and the invoices, and can hold an invoice from finalising until the numbers agree.

Correctness, not just counting

  • Append-only ledger. The app's database role can only read and insert usage rows. A correction is a new, reconciler-approved event, never an edit.
  • Tenant isolation in the database. Row-level security scopes every transaction to one tenant and fails closed, and a lint rule bans opening transactions any other way.
  • Exactly-once publication. Each billing event has a deterministic ID built from tenant, meter, signal, hour and sequence, plus a lease so only one writer publishes per tenant.
  • Late data never reopens the past. It's counted in the hour it arrived, so a closed hour never changes. A written proof covers the clock-skew and timeout bounds, and a sentinel checks them in production.
  • Drift is classified, not just detected. Differences are sorted into timing, structural or systemic. Three or more tenants drifting at once freezes publishing everywhere.
  • Testable time. Reading the clock directly is banned outside one package, so a simulated clock can backfill three months of history in a test.

Running it

It runs on Kubernetes from day one, via a Helm chart with CloudNativePG, a ClickHouse operator and Keycloak. Observability is OpenTelemetry, Prometheus and Grafana, with an SLO dashboard, 20 unit-tested alert rules and 23 runbooks. A React 19 dashboard (TanStack Router and Query, types generated from OpenAPI) covers usage, invoices, budgets and alerts, and ingest keys.

By the numbers

  • 11 services and tools, 81 Go packages
  • 584 tests and fuzz targets across 185 test files
  • 24 Postgres migrations, 14 design docs and 74 recorded decisions
  • load test: 2,250 events per second sustained on a 10-CPU local cluster, limited by ClickHouse inserts