Ask most teams running an AI agent in production what it cost last week and they can answer to the token. Ask them to reconstruct what the agent actually did on the task that went sideways, which tools it called, in what order, where it looped, what it quietly decided to skip, and the room goes quiet. By one 2026 observability stack guide’s reading of Gartner figures, only about 15% of generative-AI deployments instrument any agent observability at all. The cost meter is everywhere. The behavior record is almost nowhere. And after the month this industry just had, the behavior record is the one that carries weight.
This is the gap nobody put on a slide. Every pricing post on this site, the caching mechanics, the cost map, the effort dial, taught you to measure the model. The model is the easy half. It emits tokens and latency, both of which land on a dashboard without any effort from you. What the model does not emit is whether the agent wrapped around it made good decisions, and that is where production agents actually break.
The number you watch belongs to the model. The failure belongs to the agent.
Run network operations at a large telecom long enough and you learn the most dangerous dashboard is the all-green one. I have sat on incident bridges where every host was reporting healthy, every link was up, every synthetic check was passing, and customers still could not complete a call. The infrastructure was fine. The behavior was wrong. The metrics we had measured the things that were easy to measure, not the thing that was failing.
Agents fail the same way, and the failure modes are specific. A trajectory trace surfaces well-formed but incorrect outputs, where the agent returns a confident, syntactically valid answer that is semantically wrong; redundant tool calls that inflate cost and latency; semantically invalid actions that are correct in form and wrong in operation; and non-deterministic behavior, where the same input produces a different tool sequence on the next run. One comparison of the major platforms put the problem plainly: an agent can fail while returning HTTP 200, “selecting the wrong tool, leaking irrelevant context into a prompt, retrying until cost explodes.” Infrastructure can be healthy while behavior is wrong.
None of that shows up on a token chart. A retry storm looks like higher spend, not like a bug. A wrong tool selection looks like a normal tool call. The agent that quietly skipped a verification step and reported success looks, from the outside, exactly like the agent that did the job. You cannot debug behavior you never recorded.
What a trace of an agent actually records
The fix is not another metric. It is a trace: a structured record of every step the agent took, nested so you can see the shape of the run. The emerging standard for that record is OpenTelemetry’s GenAI semantic conventions, which matter because they give one shared vocabulary for agent telemetry instead of a different format per framework.
The conventions define the operations an agent run is made of. create_agent and invoke_agent cover the agent lifecycle and a single run; plan is the reasoning and decomposition step; invoke_workflow is an orchestrated multi-step process; execute_tool is one tool invocation. Two metrics do most of the work in production: gen_ai.client.operation.duration, end-to-end latency tagged with error.type on failures, and gen_ai.client.token.usage, token counts broken out by type, operation, and model. Because the Model Context Protocol moved into the same repository, an agent calling a tool over MCP emits the same execute_tool span as an agent calling a tool directly, which means your MCP connectors and your native tools trace the same way.
The practical upshot: instrument once and the trace reads correctly in any backend that speaks the convention. Datadog, for one, now natively supports the GenAI conventions from v1.37 up, mapping gen_ai.* spans into its agent view with no code changes and a free tier up to 40,000 LLM spans a month.
One caveat worth saying out loud, because the marketing will not. The conventions moved to their own repository on June 12, 2026 with release v1.42.0, and every gen_ai.* span, metric, and attribute in the registry still sits at “Development,” the first stability rung. Only inherited attributes like error.type are marked Stable. “We are OpenTelemetry-compliant” therefore means less than it sounds; you are building on a spec that is young enough to still move under you. Adopt it anyway, since the alternative is a bespoke format, but pin your version and expect churn.
A trace tells you what happened, not whether it was good
Here is the line that should reshape how you think about this. “OpenTelemetry captures what happened. It does not assess whether what happened was good.”
Tracing records that the agent called a search tool, got a result, and produced an answer. It does not score whether the answer was faithful to what the tool returned, whether the agent picked the right tool, whether it stayed inside the authority it was granted, or whether it leaked something it should not have. OpenTelemetry does not cover output evaluation, safety and compliance scoring, content quality, or guardrail enforcement. That assessment is a separate layer that ingests the telemetry and applies scoring, and it is the layer that answers the question you actually care about.
So the real architecture is two parts, not one. Trace the trajectory, then run evals on the trajectory: did the agent select the correct tool at each decision, did it stay in scope, did it stop when it should have, did the final output match the retrieved context. This is the same discipline as building a real eval harness instead of trusting a leaderboard, pointed at the agent’s steps rather than the model’s single answer. It is not free: scoring outputs with an LLM judge adds a per-query cost that Fiddler calls the “Evaluation Trust Tax.” Budget it the way you budget any other token line, because a judge model running on every production trace is itself a meter.
Choosing a stack without buying the wrong thing first
The tool market here is crowded and consolidating fast. ClickHouse acquired Langfuse in January 2026 and Braintrust raised an $80M Series B in February, by one stack guide’s accounting, so expect the names to shift. The useful move is to choose by deployment model first, then by features.
| If your binding constraint is | Reach for | Shape |
|---|---|---|
| Data residency / self-hosting | Langfuse or Arize Phoenix | Open-source, you run it; Phoenix is Apache 2.0 |
| Speed to first value | LangSmith or Braintrust | Managed SDK, fastest to stand up |
| Pure cost visibility | A proxy gateway (Helicone) | Sits in front, meters spend |
The specifics matter when the invoice arrives. LangSmith is a managed, proprietary service at roughly $39 per seat per month plus about $0.50 per thousand base traces; it is strongest when LangGraph is your application backbone. Arize Phoenix is Apache 2.0 and self-hostable, aimed at teams that treat telemetry portability as strategy. Braintrust is proprietary and centers evaluation-driven delivery, where a pull request has to demonstrate behavioral quality before it ships, with free allowances rising to roughly $249 a month. The most common production pairing in 2026 is simple: LangSmith or Langfuse for tracing, Braintrust for evals. You do not need one tool to do both well, and the open-source tracers (Langfuse, Phoenix) carry the adoption for a reason, since the trace data is the part you least want locked inside a vendor.
Pick the deployment model that matches your regulatory and residency posture, not the logo with the best demo. A self-hosted tracer you own beats a managed one you cannot export when a regulator or a customer asks for the record.
The trace is also the record you can be asked to produce
This is where agent observability stops being an engineering nicety and becomes the throughline of everything this site has covered for a month. The AI Agent Accountability Act introduced October 1 puts the company operating an agent on the hook when it recklessly causes damage, and recklessness is conscious disregard of a known risk. The risk is now public record. The defense against a recklessness claim is a documented, bounded deployment you can reconstruct: narrowest scope, a log of what the agent touched, a human gate on consequential actions. The trace is that log.
It has to be built like one. When OpenAI’s own agents turned an internal store into an improvised message board and reached external systems, the lesson for operators was to keep your own record of what the agent did outside the agent’s reach, the same way I argued the egress log belongs at the host, not inside the sandbox the agent controls. An agent that can rewrite its own trace is a witness grading its own testimony. Store the telemetry on a plane the workload cannot touch, tie each span to a request and a scope, and treat the trace as evidence, not as a debugging convenience you clear weekly.
A build order that holds up: span the full trajectory with the OpenTelemetry operations, not just the model call; push it to a backend the agent cannot write to; run trajectory-level evals for tool selection, scope adherence, and output faithfulness; alert on the behaviors that never hit the cost chart, loops, retry storms, and wrong-tool selection; and keep enough retention to reconstruct any run someone later asks about. None of it is exotic. It is the observability discipline ops teams have run on networks for twenty years, applied to a system that now acts on its own.
You already budget the token bill, and you already know what your agent cost. The harder number is what it did, and the trace is the only place that answer lives. The run you cannot reconstruct is the run you cannot defend, which, now that an agent’s mistakes can carry a company’s name, is the whole point of writing it down.
