OpenTelemetry solved the schema. It never shipped a backend.
Over the past year, OpenTelemetry's GenAI Semantic Conventions have quietly done something useful: they standardized what a model call looks like on the wire. gen_ai.system, gen_ai.request.model, gen_ai.usage.input_tokens, the same attribute names whether you're calling OpenAI, Anthropic, or anything else. That's real progress. Telemetry that used to be a different shape per vendor now has one shape.
But a schema isn't a product. Read the spec closely and it says so itself: the working group is standardizing attributes, not shipping a backend. To see any of this data today, you still need a collector, an exporter, a storage layer, and a UI. That's Datadog, or Grafana and Tempo, or Jaeger, or some other platform you didn't want to evaluate before your first trace. And none of them compute cost. OTel captures token counts and stops there, because pricing isn't a telemetry concern. Every team ends up hand-rolling that part anyway.
So the actual first-run experience, for a solo developer who wants to know why their API bill spiked, is: stand up a collector, configure an OTLP exporter, connect a backend, and only then see a number. That's 30 to 60 minutes of infrastructure decisions before any insight. For a side project, that's most people's stopping point.
What we built
TraceMeter is the missing local-first consumer of that data. Not a competing schema: every span it emits follows the same gen_ai.* attribute names, so it's a legitimate citizen of the OTel ecosystem, not a silo. What it adds is the part the spec deliberately leaves out: a zero-infra way to actually look at the data, and a cost layer on top of it. Auto-instrumentation covers direct OpenAI, Anthropic, and LiteLLM clients, plus LangChain and LlamaIndex through their own callback systems, so it's not limited to one SDK.
pip install "tracemeter[server]"
tracemeter demo
That command seeds a couple weeks of realistic synthetic pipeline data and opens the dashboard on it, no API keys or real LLM calls involved, so you can see what it does before wiring anything up. Instrumenting a real pipeline looks like this:
from openai import OpenAI
import tracemeter
client = tracemeter.instrument_openai(OpenAI())
with tracemeter.span("summarize_doc"):
client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "..."}],
)
tracemeter serve
That's the whole setup. No collector, no exporter config, no account. Traces land in a local SQLite file, and tracemeter serve opens a dashboard on localhost with a waterfall view per run, cost and latency broken down by model or by day, run-vs-run comparison, and CSV/JSON export.
Interoperable in both directions
Because the wire format is standard, it works both ways. Point any OTel-instrumented app's OTLP exporter at a running tracemeter serve instance and it becomes a lightweight local backend, cost computed automatically from the same gen_ai.usage.* attributes, with zero TraceMeter-specific instrumentation. And a pricing table is just a JSON file: we'd rather have the community keep it current than have any one team pretend to own the truth about what every model costs.
Built fast, checked twice
We shipped this in one focused pass, but "fast" only stays honest if you verify as you go instead of at the end. Two things from that process are worth sharing, because both were assumptions that looked reasonable and turned out to be wrong.
The first was about the OTLP ingest endpoint itself. The plan was JSON-only for v1: simpler to parse, no extra dependency, and the spec supports it. It passed every unit test built from hand-written payloads. Then we pointed a real opentelemetry-sdk exporter at it, and every request failed. It turned out the mainstream Python OTLP/HTTP exporter always sends protobuf, regardless of what you set OTEL_EXPORTER_OTLP_PROTOCOL to, there's no way to make it emit JSON. A JSON-only endpoint would have looked finished in tests while not actually interoperating with the primary tool it needed to interoperate with. We added protobuf support instead, via opentelemetry-proto (just the compiled message classes, not the full SDK), and re-verified against the real exporter: nested spans, trace hierarchy, and cost computed automatically, all correct.
The second was the pricing table. It's easy to ship a snapshot of provider pricing and call it done. We checked it against the live OpenAI and Anthropic pricing pages instead, and found real drift: one model was priced five times too high after a cut we'd missed, two others were still on pre-cut introductory pricing, and one entry had quietly been removed from the provider's own pricing page, so we pulled it rather than keep an unconfirmed number. Removing it surfaced an actual bug: a fallback matcher meant to handle dated model snapshots (gpt-4o-2024-08-06 falling back to gpt-4o) was also matching unrelated models that happened to share a prefix, like o1-mini silently pricing itself as o1. Fixed by requiring the matched suffix to actually look like a date. That's exactly the failure mode the pricing engine is supposed to prevent: an unknown model should show up as "unknown," never as a confident, wrong number.
Neither of those would have shown up without testing against the real thing, a real SDK, real live pricing pages, instead of trusting what looked complete on paper.
Get started
pip install "tracemeter[all]"
It's early. The pricing table will drift again the moment providers change prices, which is exactly why it's a plain JSON file with an open PR path rather than something we gatekeep. If a model's missing or a number's off, that's the easiest and most useful contribution to make.