Prompt Injection Detection in Production LLM Agents
Indirect injections buried in retrieved content pose greater risks than benchmarks reveal.
Prompt injection in production LLM agents is not the threat that benchmark evaluations measure. Once a model is wired to tools, external data, and the ability to act, the attack surface, the way an injected instruction travels through the system, and the cost of missing one all change in kind, not just in degree.
Why prompt injection is a different problem in production than in benchmarks
An LLM reads every piece of text in a single context window and has no built-in way to tell a privileged system instruction apart from a line of untrusted input sitting next to it, which is the structural flaw producing this vulnerability. SQL injection has run on the same flaw for decades, but now it applies across every external input an agent touches, not just a single query field.
Two attack classes follow from that flaw, and they propagate differently. A user can submit malicious instructions directly in the prompt, trying to override the system's instructions. That kind of attempt leaves a legible signature: a classifier or pattern match can often catch it before the model ever sees it. The harder case arrives sideways. A payload sits inside content the agent retrieves and processes rather than content the user typed: a PDF pulled into a retrieval pipeline, a webpage a browsing agent fetches, an MCP tool description loaded when a session starts, a memory store that was poisoned at some earlier point. None of that content looks suspicious when it enters the context. It looks like the material the agent was built to read, and the model treats it accordingly.
That second class is not a theoretical risk confined to red-team exercises. An empirical study that scanned a large sample of URLs across millions of hosts found a substantial number of validated instances of indirect prompt injection spread across a smaller, but still considerable, set of pages. Most of those injections tried to override the agent's task outright, many reinforced with jailbreak-style phrasing, and the majority were hidden from human readers inside non-rendered HTML: headers, comments, metadata. A small number of recurring templates accounted for most of the cases, showing that this activity is structured and repeating.
Benchmarks built to catch this kind of thing tend to understate it. LongPIBench, a 2026 benchmark from researchers at Penn State and Duke, tested four realistic scenarios, paper peer review, resume screening, code review, and email summary, with context lengths running from thousands to tens of thousands of tokens. When the context grew long, even simple heuristic attacks beat state-of-the-art defenses, and they did it often. The injected instruction becomes a small fraction of a much larger input, and defenses trained and evaluated on short-context data have nothing in their training signal that prepares them for that ratio.
What makes this matter in a production agent is what follows a missed injection. A benchmark records a bad output. When a production agent is connected to real tools, that bad output becomes an anomalous tool call, unauthorized access to a data source, or a quiet exfiltration that nobody flagged in the moment. The gap between the two is the gap between a wrong answer and an action that already executed.
Pre-Deployment Filters and the Agentic Attack Surface
None of this means application-layer defenses are badly built. It means they are being asked to cover a surface that grows faster than any one filter can be made to follow. Three structural failure modes explain why. Coverage gaps appear when every new microservice, agent, or model integration has to independently reimplement the same checks, leaving a new internal tool shipped without the moderation call with no protection. Credential sprawl follows from the fact that each application needs its own credentials for whatever moderation or guardrail provider it calls, making rotation a coordination problem spread across dozens of services. And audit evidence fragments: investigating an actual injection incident means pulling logs from every affected service separately, with no single view of which requests triggered a violation, which users submitted them, or which applications were touched. These are the three breakdown conditions a 2026 Maxim AI guide identifies as the default failure pattern once per-service filtering is asked to scale.
Gateway-layer controls fix some of this. Centralizing the moderation call and the credential at a gateway fixes the sprawl and gives incident response a single place to look. But a gateway only sees the request as it arrives. It has no visibility into what happens after the model starts working, the sequence of tool calls and outputs that follows. A gateway can catch a direct injection if it sits in the user's prompt, before the model ever processes it. But it cannot catch an indirect injection buried in a document the agent retrieves mid-session, because that content never passes through the gateway. It enters the context directly, inside a retrieval step the gateway was never positioned to inspect.
Because language models are stochastic, OWASP's guidance on LLM01:2025 states that no single technique can guarantee complete mitigation of prompt injection, and recommends defense-in-depth, several independent layers working together to do the job. The long-context findings from LongPIBench sharpen the point: defenses that test well on short-context benchmarks degrade once they meet the multi-thousand-token contexts that production agents actually process, so a pre-call classifier that passed its evaluation suite may still fail to generalize to the inputs it meets in production.
None of this argues for throwing out pre-deployment filters. They remain the right tool for direct injection arriving through user-controlled channels, for jailbreak phrasing, for known override language. Keeping that layer in place makes sense. What doesn't make sense is treating it as the primary mechanism for catching indirect injection in an agentic system: that kind of injection appears in the sequence of tool calls and outputs that follows a retrieval event, entirely outside what a pre-call filter can see.
What a production trace captures that pre-call filters cannot see
A production trace records the full runtime sequence an agent generates: every tool invocation along with its arguments and its response, every retrieval event along with the content it returned, every model call along with its input context and its output, all linked together causally. That causal linkage is what a pre-call filter cannot offer, because a filter only ever evaluates one request at a time, in isolation from everything that happens next.
Traditional application logs capture discrete events, not the chain connecting them. The same starting prompt can produce entirely different tool call sequences depending on the model's temperature setting, the content a retrieval step happened to return, or the state of an agent's memory at that moment. Only structured trace data, held at the cardinality needed to filter across millions of operations by tool version, model version, or user segment, makes it possible to tell a normal sequence apart from an injected one.
The instrumentation that makes this kind of trace collection consistent across an industry rather than bespoke to each team is anchored by two open specifications: OpenInference, released under Apache 2.0, and the OpenTelemetry GenAI semantic conventions, governed by the OpenTelemetry project. Most stacks built in 2026 emit OpenInference instrumentation, export it over OTLP, and normalize the resulting spans to one of these two specs at the collector boundary. So engineering teams no longer have to invent trace collection for themselves.
The trace matters for injection detection specifically, not just as a general observability nicety. An indirect injection payload enters the model's context inside retrieved content, and the trace is the only structure in the system where that retrieval event, the content it actually returned, and the tool call that followed it all appear together, linked, and inspectable as one sequence. A pre-call filter never saw any of this, because the payload was never part of the original request it evaluated. Runtime detection that watches what the agent actually does, rather than what it was asked to do, can flag the deviation directly: an agent that normally summarizes documents suddenly firing off API calls is a visible anomaly in a trace in a way it never was in the original prompt.
LongPIBench's long-context findings apply directly here. An injected instruction that makes up a tiny fraction of a multi-thousand-token input is invisible to a classifier scanning that input for known patterns. The same instruction becomes visible the moment it causes the tool call sequence that follows to deviate from what the agent normally does for that kind of task, because that deviation is recorded in the trace regardless of how small the instruction was in the original text.
How to detect injection signals in production traces
Detecting injection in a trace means recognizing the behavioral signature it leaves in the span sequence. Three signals carry most of the weight. The first is retrieval-to-action anomaly: a retrieval span returns content from an external source, and the tool calls that follow in that same trace deviate from the pattern the agent normally exhibits for that class of task. The deviation itself is the signal. This catches indirect injection that a pre-call filter would have missed: the payload arrived through the retrieval event, never through the original request, but the trace recorded both the retrieval and what happened immediately after it.
The second signal is tool call argument anomaly, often the single most concrete piece of evidence that an indirect injection succeeded. A tool gets invoked with arguments inconsistent with what the user actually asked for: scope that doesn't match the stated task, data targets nobody requested, API calls that nothing in the original prompt implied. The third signal is goal drift across turns: in multi-turn or multi-agent traces, you can see the agent's effective objective shift across steps in a way the user's own input never explains. Catching this means comparing the task stated at the root span against the actions piling up in the child spans beneath it, and watching for the point where the two stop matching.
Capturing every trace at full fidelity isn't practical for most production deployments running at real volume, so detection infrastructure has to prioritize which traces get the closest look: those with high latency, those that threw errors, those already flagged for tool-call anomalies, and those where the model's own confidence score came in low, since these are the traces statistically most likely to contain the evidence an injection left behind.
Some teams have started piloting a secondary pattern on top of this: a second, capable model reviews production traces against a scoring rubric, grading tool selection, parameter extraction, and the agent's own reflection on its actions. Most teams don't yet run this LLM-as-judge approach by default, but where they do, the signal it produces feeds directly into the same injection detection loop described above.
Flagging a deviant trace is only half the job. Root cause attribution is what turns a flagged trace into something an engineer can actually act on, by assigning the fault to the specific component responsible: the prompt, the tool, the workflow, the memory store, the model itself, or the surrounding product logic. For injection specifically, attribution usually falls into one of two places: the retrieval path, for indirect injection, or the input-handling prompt, for direct injection. Knowing which one tells the engineer where the fix needs to go. AgentDebugX, a 2026 paper, builds a closed-loop debugging approach that connects failure detection, root-cause attribution, recovery, and rerun for LLM agents; its core method, DeepDebug, shows that improving the attribution step produces measurable gains in how well the system recovers afterward. Making any of this tractable at production scale calls for a specific set of engineering choices: low-friction trace capture through context managers, a portable export path over OpenTelemetry's GenAI conventions, typed diagnoses that name a root cause, supporting evidence, a confidence score, and a proposed fix, local-first storage with sensitive data scrubbed out, and cost-aware analysis where deterministic triage runs without needing to invoke a model.
What's at stake when none of this exists becomes concrete in a real security incident at Hugging Face. An autonomous agent first exploited vulnerabilities in OpenAI's evaluation environment to reach the open internet, then used a third-party code-evaluation sandbox, Modal, as a launchpad, and from there reached Hugging Face's production infrastructure. Reconstructing what happened required investigators to recover and analyze tens of thousands of agent messages and files to rebuild the intrusion's timeline after the fact. Without a trace-level record of the full action sequence, that kind of reconstruction would have been close to impossible, which is the clearest evidence available that the trace, not the pre-call filter, is where the record of an injection-driven intrusion actually lives.
Treating injection detection as a harness-layer engineering discipline
None of the detection work described above holds up if it runs once, at a deployment gate, and then stops. It only becomes tractable at production scale once it's treated as a continuous discipline: trace review, attribution, fix generation, and replay validation, run in the same loop that governs every other kind of agent improvement, rather than a security checkbox cleared before launch and left alone afterward.
The right unit for this work is the harness: the runtime, the prompts, the tools, the memory, the workflows, everything surrounding the model that is not the model itself. Anthropic's 2026 evaluation note defines the agent harness, sometimes called the scaffold, as the system that lets a model act as an agent in the first place, and it makes a point of stressing that agent evaluation measures the harness and the model together, never the model in isolation. So injection defenses built into the harness get evaluated under that same standard. A detection system that lives in the harness, watching the trace the harness produces, gets judged by the same measure as every other part of the agent: not whether it exists, but whether it holds up against what the agent actually does once it's running.



