AI Agent Trace Analysis With Custom Span Facets
Custom span facets add the semantic context standard tracing misses for debugging AI agents.
A trace viewer that reports "tool call failed" and stops there tells an engineer almost nothing worth acting on. Standard OpenTelemetry gen_ai.* attributes record what an agent did at each step of its run, but they carry no information about why a step went wrong or which layer of the system is responsible for the error. The OpenTelemetry GenAI semantic conventions define fields like model name, token counts, and finish reason, which work well for tracking cost and latency, but they say nothing about a hallucinated function argument or a retrieval result that was simply stale. Those are semantic failures, and the schema was never built to describe them.
The consequence appears in infrastructure monitoring itself. Traditional APM tooling treats a 200 response as a successful outcome, but an agent can return 200 after looping twice, calling the wrong tool, or inventing a billing policy that doesn't exist, and none of that registers anywhere in the infrastructure signal. The request succeeded by every metric an ops dashboard tracks, while the task itself failed in a way only semantic inspection would catch. This is the diagnostic gap that custom span facets are built to close, and understanding it requires first understanding what a span can and cannot carry on its own.
What a span actually is in an agent context, and where its schema runs out
An agent trace is a tree of spans, and each tool call, each LLM invocation, and each retrieval step becomes a child span in that tree, together producing a full record of the reasoning chain the agent followed. Prompt text and completion content typically live in span events rather than span attributes, a design choice driven by privacy and payload-size concerns rather than by any limit on what a span could technically hold.
A reasonably complete agent trace schema tends to rest on four pillars: tool-call spans that capture name, arguments, return value, retries, and errors; reasoning spans that capture the plan, the action taken, the observation returned, and the next decision; state-transition spans that capture working memory before and after each step; and memory-operation spans that capture the query issued, the entries returned, their relevance scores, and their freshness. That's a solid scaffold for reconstructing what happened during a run.
Span typing systems extend that scaffold further. Arize Phoenix classifies spans as CHAIN, RETRIEVER, RERANKER, LLM, EMBEDDING, TOOL, and AGENT, while OpenAI's Agents SDK uses its own taxonomy, including AgentSpan, GenerationSpan, FunctionSpan, GuardrailSpan, HandoffSpan, ResponseSpan, CustomSpan, MCPListToolsSpan, TranscriptionSpan, TurnSpan, TaskSpan, SpeechSpan, and SpeechGroupSpan, among others. Both taxonomies make large traces faster to scan, letting an engineer jump straight to the tool spans or the retrieval spans without reading the whole tree. But both still describe execution mechanics rather than business-layer intent.
That's where the schema runs out. None of these standard fields, however well typed, record the rationale behind a tool selection, an assessment of retrieval quality, the memory state an agent was working from when it made a decision, the workflow stage it occupied, or the policy constraint that was supposed to govern that step. Those are the attributes that make a span diagnosable rather than merely observable, and the standard schema leaves the field for them empty by design.
What custom span facets are and how they attach domain meaning to individual spans
Custom span facets are attributes that engineers define and attach to individual spans to encode the domain-specific context a standard schema omits. The mechanism behind them is manual instrumentation through the OpenTelemetry SDK, which allows arbitrary key-value attributes to be added to any span at the moment it's created. Manual instrumentation earns its keep specifically where auto-instrumentation can't reach: custom retrieval steps, evaluation scores, business logic embedded in an agent's decision process, or orchestration across multiple models.
In practice, adding a facet looks like the following:
with tracer.start_as_current_span("tool_call") as span:
span.set_attribute("harness.tool.rationale", "user asked for refund status; policy requires lookup before reply")
span.set_attribute("harness.retrieval.relevance_score", 0.42)
span.set_attribute("harness.workflow.stage", "act")
result = run_tool(args)
Nothing about that call replaces the standard gen_ai.* attributes already present on the span. It sits alongside them, adding the layer of meaning the base schema was never designed to carry.
A few facet categories tend to do most of the diagnostic work. A tool-call rationale facet records why the model selected a given tool at a given step, not merely which tool it picked, surfacing mismatches between a prompt's stated policy and the routing decision the model actually made. A retrieval-quality facet, recording a relevance score, a freshness timestamp, a source identifier, and a count of chunks returned, lets an engineer distinguish a retrieval failure from a model failure occurring on the very same span. A memory-state facet, recording which entries were read, how old they were, and whether they matched the task's current context, exposes stale reads and retrieval against the wrong entity. A workflow-stage facet, a structured label such as plan, act, observe, or verify, locates a span within the agent's reasoning loop and allows failures to be grouped by phase rather than by span type alone. A policy or constraint tag records which behavioral rule governed a given step, enabling a team to filter every span where a specific guardrail fired.
Portability matters here as much as expressiveness. OpenInference semantic conventions, used by Arize Phoenix, hold up well when the same spans are pushed to multiple backends, and custom facets added under a consistent naming convention inherit that same portability. A facet scheme built carelessly becomes backend-specific glue code; one built with naming discipline travels with the trace wherever it's sent.
How facets enable layer-level failure attribution, not just failure detection
The deeper argument here is that facets convert root-cause attribution from an open-ended search across a trajectory into a structured query against a labeled graph. Root-cause attribution in agent systems assigns fault to one of four layers, the model, the harness, the environment, or the grader, and that assignment is what directs the correct intervention: a harness failure calls for harness redesign, not model retraining. Facets are what make that assignment computable directly from trace data, rather than something inferred from an engineer's intuition after staring at logs for an afternoon.
The difficulty attribution has to overcome is structural. The three primary failure categories each become detectable at their origin span with the right facets. Treating attribution as a local inspection problem, one span at a time, when it's inherently a trajectory-level problem, is the source of most debugging dead ends encountered in production. Facets solve this by making the trajectory queryable as a whole: if every retrieval span carries a relevance-score facet and every tool-call span carries a rationale facet, a team can filter every failed run by low retrieval scores and check whether that predicts downstream wrong-tool selections, a correlation that stays completely invisible in unlabeled traces.
Three failure categories illustrate how concretely this works at the origin span. Tool schema drift becomes detectable when a facet records the tool schema version active at the moment of invocation, letting a team correlate argument-validation failures with a specific schema revision instead of discovering the drift only after users start complaining. Prompt ambiguity becomes detectable when a workflow-stage facet is combined with a policy-tag facet, letting engineers filter for spans where the agent was in the "act" phase under a specific behavioral rule and still produced an unexpected tool selection, surfacing the ambiguity in the policy itself before it compounds into a chain of bad decisions. Workflow loops become structurally visible, rather than merely inferred from a spike in token counts, when a step-counter or visited-state facet is attached to reasoning spans; loops burn both tokens and latency and frequently end in outright task failure once the context window fills.
AgentDebugX's DeepDebug component performs multi-turn root-cause diagnosis using global trajectory understanding, structure-guided investigation, and cross-examination. That approach validates the trajectory-level framing directly: diagnosis depends on reading the full trajectory, and facets are simply the mechanism that makes that kind of trajectory-wide reading queryable as structured data rather than something a tool, or an engineer, has to reconstruct by hand from raw spans.
Designing a facet schema for the five harness artifact classes
A facet schema earns its keep when it mirrors the artifact classes an agent is actually built from, so every dimension an engineer can query maps onto a layer they can actually change. Five harness artifact classes, prompts, tools, memory, skills, and code, give that schema its organizing structure, and each class carries its own characteristic failure modes that facets should expose at the span level.
Prompt spans benefit from facets recording the prompt version or hash in effect, the policy rule governing the step, and a confidence or ambiguity signal drawn from the model's own completion, enabling a team to filter for every span where a specific prompt version produced behavior nobody expected. Tool spans benefit from facets recording schema version, the result of argument validation, the reason behind a retry, and whether the next reasoning step actually used the result or discarded it, exposing drift between the contract a model learned during training and the schema actually live in production. Memory spans benefit from facets recording query type, whether it was semantic search or a direct key lookup, the identifier of the retrieved entity, a relevance score, and the entry's age, which make stale reads, retrieval against the wrong entity, and memory leakage across sessions visible instead of unnoticed.
Skill spans benefit from facets recording the skill identifier, the routing decision that selected it, and whether negative examples in the skill manifest were matched against the current request. Adding negative examples to a versioned skill manifest raised routing accuracy from under three-quarters to over four-fifths, per OpenAI's engineering guide on versioned Skill bundles with a SKILL.md manifest, and a routing-decision facet makes that kind of improvement measurable in production. Code and orchestration spans benefit from facets recording workflow stage, a loop-iteration counter, and the target of any handoff, which together locate orchestration failures such as dead ends and missed escalation paths.
Tagging spans with user and workflow identifiers alongside these harness-layer facets turns isolated failures into detectable patterns once trace volume is large enough. A single bad memory read looks like noise, but the same facet combination appearing across hundreds of traces is a pattern worth fixing at the harness level. Naming discipline is what keeps this schema usable as it grows: facets organized under a consistent prefix, something like harness.prompt.* or harness.tool.*, inherit OpenInference portability and move across backends without requiring custom parsing rules for each one.
Using faceted traces to power replay-based validation before shipping harness changes
Facets don't only serve diagnosis after the fact; they make the validation that precedes a harness change far more targeted. Re-running an agent without recording the model outputs, retrieval results, tool responses, and environment state that shaped the original run produces a new trial, not a faithful replay of the old one, and facets on each span are part of what makes a recorded trace replayable in the first place rather than merely rerunnable.
Recovery replay is used for crash recovery, debug replay is used to reproduce one specific failure, forensic replay is used to audit an incident after the fact, and evaluation replay is used to validate a proposed change against historical behavior, and these four modes are worth keeping separate rather than treating "replay" as one undifferentiated activity. Facets are most directly useful in debug and evaluation replay, where querying a trace set by harness layer is the central operation rather than a side benefit.
This is where the golden-set workflow becomes practical: a curated set of production traces paired with expected outputs or rubric scores, replayed through an updated pipeline before that pipeline ships. Facets let a team build that golden set by querying directly for traces where a specific tool schema version was active, where a specific skill routing decision fired, or where a specific memory query returned a low relevance score, rather than sampling traces blindly and hoping the sample happens to cover the failure mode in question.
Running LLM-as-judge scorers against a sampled slice of production traces, bucketed by time window, supports alerting on score movement over time. A drop in production scores with no corresponding drop against the golden set usually signals that the input distribution has shifted, not that the pipeline itself has regressed, and facets on retrieval and memory spans are what make that distinction visible rather than a matter of guesswork. When a production run fails, stalls, loops, overspends, or simply produces a weak answer, the faceted trace identifies where the breakdown occurred, whether in retrieval, tool schema, tool result, prompt policy, model call, or state transition, and a sanitized version of that case can then be promoted straight into an eval dataset, with facets accelerating the promotion by making the failing layer queryable rather than requiring manual inspection of the raw trace.
Validating a harness change this way, against traces filtered on the exact layer being modified, is a categorically different form of evidence than validating it against a synthetic benchmark built before the change was conceived.
Integrating custom facets into an existing observability stack without a full rewrite
Custom facets are additive to any OTel-compatible stack, and a team can instrument one harness layer at a time without touching existing spans or switching backends in the process.
Framework adapters normalize spans into a common schema across the OpenAI Agents SDK, LangGraph, Mastra, Pydantic AI, and custom or unsupported stacks, falling back to plain OpenTelemetry where no dedicated adapter exists. Custom facets added under a consistent naming convention survive that normalization step and show up correctly in any OTel-compatible backend afterward.
A practical sequence for teams starting from a standard OTel setup runs roughly as follows:
- Add tool schema version and retry-reason facets to tool-call spans first. Schema drift is the failure category most likely to stay invisible in existing logs, which makes this the highest-signal place to start.
- Add relevance score and entry-age facets to memory and retrieval spans next, surfacing the stale-read failures responsible for silent quality drift over time.
- Add workflow-stage and loop-counter facets to reasoning and orchestration spans, making loops and dead ends structurally visible in the trace rather than something inferred from anomalies in token counts.
- Add prompt version and policy-tag facets last, enabling filtering by prompt change so regressions introduced by a specific edit can be isolated from everything else happening in the system at the same time.
Platforms that accept OTLP with automatic span conversion, including those built with OpenTelemetry integration from the start, carry custom facets through without any additional configuration. That preserves portability as a team's backend choice evolves. The facet schema a team builds today keeps working regardless of where the traces end up being stored and queried tomorrow.
Sources
- OpenTelemetry for AI Systems: LLM and Agent Observability (2026)
- AgentDebugX: An Open-Source Toolkit for Failure Observability, Attribution, and Recovery in LLM Agents
- GitHub - ai-boost/awesome-harness-engineering: Awesome list for AI agent harness engineering: tools, patterns, evals, memory, MCP, permissions, observability, and orchestration. · GitHub



