What LLM Tracing Captures and What It Deliberately Omits
Standard tracing frameworks miss the causal failures that matter most in LLM agents.
An agent completes a run. No exceptions fired, latency sat within normal range, the response came back with a clean 200, and the answer was still wrong. That gap between a healthy-looking trace and a broken outcome is not a bug in the tooling, and it will not get patched by the next release.
Tracing, at its core, captures tool names, arguments, return values, latency, retry counts, and error states. It captures reasoning steps and the plan-act-observe transitions that string them together. It captures state before and after each step, and it captures memory reads and writes along with retrieval scores. Together, these form what one widely used agent observability guide calls a minimum viable schema, the floor below which a trace stops being useful at all. That's a genuinely useful floor. It just was not built to answer causal questions, and it never claimed to be.
This pattern is inherited from distributed observability frameworks, not invented for agents. Distributed observability frameworks such as OpenTelemetry model executions as traces and spans, a design built for deterministic service architecture: the same input produces the same output, and a 200 response means the request succeeded. That model works fine for a payments API or a database call. It breaks against LLM agents on all three counts. The same prompt can produce different tool calls from one run to the next. The execution path itself branches depending on what the model outputs at each step, so there is no fixed path to compare against. And a 200 can wrap an answer that is confidently, completely wrong.
So the trace is faithful to what happened at the mechanical level: which tool got called, what came back, how long it took. It has no vocabulary for why that tool got selected over another, whether the arguments passed were justified by the task at hand, or whether the final output was actually correct. Those are the questions engineers actually need answered when something goes wrong in production, and they sit outside the schema by design, not by oversight.
4|Recording execution without causality
Each of the four span types captures one execution surface with real fidelity, and each one goes quiet at a specific, predictable boundary. That boundary is exactly where the most expensive agent failures tend to start.
Tool-call spans record the tool name, the arguments passed, the raw output, the duration, the retry count, and whether an error fired. What they don't record is whether those arguments were semantically right for the task, whether the tool's schema has quietly drifted since the prompt was written against it, or whether a retry loop is masking a failure that never actually resolves. A tool call can retry three times, succeed on the fourth attempt, and look completely unremarkable in the trace, even though something upstream is clearly broken.
Reasoning spans capture the model's plan, the action it chose, the observation it made, and the decision that followed. They're genuinely good at surfacing plan drift and wrong-branch selection in a way that a single, isolated LLM call span cannot. The span shows that the drift happened, but only after the fact, in the reasoning span that records the plan, action, and observation. The span shows that the agent went left instead of right; it says nothing about why the branch point produced the wrong choice.
State transition spans record context before and after a step, context edits, and the payloads passed during handoffs between agents. They will show that a piece of context existed at one point and was gone at the next. They don't explain why it was dropped, whether a summarization step faithfully preserved the constraints it was supposed to carry forward, or where exactly in a multi-agent handoff the original intent got lost.
Memory operation spans capture reads and writes to long-term stores: the query issued, the entries returned, their relevance scores, their freshness. What they don't capture is whether the entry retrieved actually served the current sub-goal, whether memory from an earlier session bled into the current one, or whether a stale read quietly shaped a decision three steps downstream.
Line all four up and a pattern emerges. Each span type is an honest record of the mechanics at its own layer. None of them carries attribution upward, toward the layer that actually caused a failure, and none of them carries causality downward, toward the effect that failure produced in the next span. The four spans, stacked together, form a complete execution record with no connective tissue between them. You get the full chain of events, but not the chain of causes.
The failure categories that hide most reliably inside well-formed traces
The failures that do the most damage in production are, more often than not, the ones that generate a perfectly clean trace. A well-formed trace is what makes these failures invisible.
Tool schema drift is a straightforward example. An external API adds a required field, or deprecates an endpoint, while the agent keeps calling the old signature it was built against. The tool-call span records a valid call and a returned value, because from the trace's point of view, nothing failed. Nothing in that span signals that the schema the model was trained on no longer matches the live API. External schema changes of this kind are a recurring, largely invisible source of drift precisely because the trace has no field for "this used to be true and now it isn't".
Prompt ambiguity behaves the same way. The model misreads an instruction or misuses a tool because the prompt left the intended behavior under-specified. The reasoning span records the model's plan with total accuracy. The plan is just wrong relative to what the task actually needed, and the trace has no mechanism for flagging that the input, not the output, was the problem. Diagnosing that kind of failure requires decomposing the rollout into its component modules, memory, reflection, planning, action, and system-level behavior, and tracing the failure back to whichever module actually produced it. That kind of root-module attribution is precisely the analysis span-level tracing does not perform on its own.
Workflow loops are subtler still. An agent retries a failed tool call, repeats an action, or stalls without making any real progress toward the goal. Each iteration of that loop produces a structurally normal reasoning-action-observation triple. Aggregate metrics stay green while the agent burns tokens and time going nowhere, because nothing about any individual span looks wrong in isolation.
Multi-agent cascades compound the problem. A flaw in one agent's output gets passed downstream through a handoff, and each receiving agent's trace looks clean, because each one faithfully executed on the input it was given. Nobody's individual trace is lying. The failure only becomes visible once traces from every agent involved get assembled into a single sequence, and even then it takes deliberate work to see it.
The hardest category to catch is the seam failure, a break that happens in the transition between spans rather than inside any one of them. A procurement agent goes to sleep for four days waiting on a vendor reply, wakes up when the reply lands, and finishes the task with a task-completion score of 0.93 and zero flagged errors. A week later, finance catches that the vendor list is missing a supplier the original prompt explicitly named. The signal handler dropped that exclusion when it merged the late reply into state, mid-run. Nothing in the per-turn rubric ever moved, because the break didn't happen in any turn. It happened in the seam between the original prompt, the sleep interval, and the mid-run merge, a gap no span was ever built to watch.
Attributing a failure to its origin layer is harder than it looks
Knowing that a run failed tells you almost nothing about where the failure came from. It could be the prompt, the tool interface, the workflow logic wrapped around the model, the memory layer, or the model itself. Without pinning that down, any fix applied is a guess dressed up as a diagnosis.
Scale makes this worse before it makes it better. Agents running long-horizon tasks generate execution logs large enough that diagnosing failures by hand is not a realistic option, even though reliability depends on doing exactly that. That pushes teams toward automated root-cause attribution using LLMs to read the trace and point at the failure. Those automated methods have a real weakness: diagnostic accuracy drops as traces get longer, because the evidence that actually explains the failure is often sparse, scattered across actions far apart in time, and disconnected from the point where the failure became visible.
A 2026 paper, Root-Cause Attribution Is a Search Problem, names the underlying issue directly. One-shot LLM diagnosis methods tend to seize on a plausible-sounding explanation early and stop looking, leaving evidence buried deeper in the trace unexamined. The paper's answer, an iterative method called Continual Search, keeps prompting for diagnosis until the unresolved evidence actually gets found, rather than accepting the first answer that sounds reasonable. Premature closure, settling for the first plausible story, is treated as a failure mode in its own right, not a shortcut.
Even once you find the right evidence, teams need a shared vocabulary for naming which layer it points to. A 2026 taxonomy, Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures, was built for exactly this. It localizes failures to the specific interaction edge responsible, model, harness, user, tools, memory, or environment, across cases running from minimal single-model setups to full coding agents and long-running personal assistants. Without a taxonomy like that, two engineers looking at the same trace can reasonably disagree about which layer broke, and ship two different, possibly conflicting fixes for what was actually one failure.
The deeper fix that both of these point toward is a provenance layer sitting on top of the trace: typed relations, support, derive, depend-on, contradict, invalidate, trigger, update, that connect retrieved evidence, tool outputs, memory items, observations, intermediate claims, actions, and the final answer into an actual graph. A trace without that layer is a sequence of things that happened, and attribution has to be inferred by whoever reads it. A trace with that layer is a graph where causality is stated outright rather than guessed at. That distinction is the whole difference between reading a trace and being able to act on one. And it's also the pivot point for the rest of the argument: once you accept that attribution has to point somewhere, the evidence keeps pointing to the same layer.
Why most agent failures are harness failures
The harness, the runtime layer wrapped around a fixed model, covering prompts, tool interfaces, context construction, state management, middleware, and recovery logic, turns out to be the main thing determining how an agent performs. That makes it the layer where most failures start, and, more usefully, the layer where they're actually fixable.
The harness is often treated as plumbing, a detail beneath the interesting part of the system. But hold the task and the model constant, and change only how the harness presents auxiliary information to that model, and success rates can shift drastically. That sensitivity is a factor commonly left out of how teams reason about agent reliability. Most of the incidents that observability tools actually surface in production trace back to tool-call failures, context getting truncated, and loops that never terminate, not to errors in the model's underlying weights, and every one of those three failure types lives in the harness.
Research backs this up directly. The Life-Harness project asks a narrow, pointed question: can a frozen model that keeps failing the same deterministic task get better if the runtime harness around it changes, without retraining the model or touching the environment? The answer came back yes, and the margin was not small. Harness-level fixes produced substantially higher pass rates on the first attempt than prompt-only tuning managed, with an average relative gain described in the research as roughly 120%, a difference far beyond a marginal tuning knob. It is evidence that the runtime wrapper around a model can matter more than the model itself for a given task.
A second project, AutoSaddler, treats harness optimization as an offline learning problem in its own right, one with outcomes that can actually be measured and compared, rather than an art practiced by intuition. Putting both results together, reliability engineering for agents looks less like a model-selection question and more like an infrastructure-design question, closer in spirit to configuring middleware than to picking a checkpoint.
That has a direct consequence for how trace analysis should actually run. If most failures live in the harness, then reading a trace should ask which harness component, the prompt template, the tool schema, the context-merging logic, the retry policy, produced the conditions the model was reacting to. The four span types described earlier already contain the evidence for that question. They were just never going to volunteer the answer on their own.



