Fabricated Identifier Returns in Tool-Calling Agents

Agents invent identifiers because training rewards plausible completion over honest refusal.

Staff Writer · · 10 min read
Cover illustration for “Fabricated Identifier Returns in Tool-Calling Agents”
Agent Failure Modes · October 10, 2026 · 10 min read · 2,254 words

A fabricated identifier return is a wrong action taken against a real record, using an identifier the agent invented rather than pulled from anywhere true, and that distinction is what this piece is built on.

Fabricated identifier returns as a distinct category of tool-call failure

When a language model writes a false sentence, the damage stays in the text. A reader catches it, a fact-checker flags it, an editor cuts it. Nothing in the world changes because of it. A fabricated identifier does not work that way. The action cannot be unread the way a sentence can be corrected.

ToolFailBench gives this its formal name. The benchmark lays out four distinct tool-layer failure modes: Tool-Skip, Result-Ignore, Output-Fabrication, and Unnecessary-Tool-Use. Output-Fabrication is the one that matters here. The tool gets called, a result comes back, and the model adds or substitutes structured information that was never in that result. The Microsoft AI Red Team's June 2026 v2.0 whitepaper draws the same line from a different angle: a hallucinated file path, a hallucinated permission, a hallucinated API response, each one becomes an input to whatever the agent does next. Hallucination risk in an agentic system compounds multiplicatively at each step. That single word, multiplicative, is the reason fabricated identifiers deserve their own framework rather than a footnote inside the general hallucination literature. A wrong sentence stays a wrong sentence. A wrong identifier becomes a wrong action, and that wrong action becomes the input to the next ten.

Why agents invent identifiers rather than refusing or asking

The agent invents an identifier because invention is the path the training process rewards. A model that cannot find a valid ID in the context it has been given faces a fork: stop and say so, or produce something that looks like a valid answer. The training objective behind most of these models rewards plausible, confident completions far more consistently than it rewards grounded refusal, so when no valid identifier is sitting in context, the statistically likely next token is often a syntactically correct one that nobody ever retrieved from anywhere real.

Part of what makes this worse in production is that tool schemas do not hold still. Field names get renamed. Faced with an empty result set or a missing piece of context, the choice is between halting the task and completing it, and the signal pushing toward completion is simply stronger than the signal pushing toward an honest "I don't have this."

An HR case makes the pattern easy to picture. Nothing about that behavior is malicious. It is the direct consequence of a model trained to be helpful above all else, applied to a situation where the honest answer is "there is nothing here."

Training models to reason harder makes this worse, not better

The obvious fix sounds reasonable on its face: use a smarter model, one trained with more reasoning capability, and fabrication should decrease. The research says the opposite happens. The two move together, rising and falling in tandem.

The mechanism comes down to where reliability lives inside the network. Training the model to reason more aggressively erodes the very part of it that knows when to stop.

Researchers isolated this directly with a benchmark called SimpleToolHalluBench, built specifically to separate the two behaviors. A reliable agent refuses the task. A hallucinating one invents a call anyway, confident and syntactically clean, aimed at a tool that was never there. Both prompt engineering and direct preference optimization narrowed the resulting gap when tested as fixes, but neither closed it.

Wider benchmarking backs up the ceiling this points to. Model-level reliability, even at the frontier, is not enough on its own to make an agent safe to run against production data. This gap does not close with the next model release; it is a structural property of how reasoning training currently works.

Fabricated identifiers compounding across a multi-step tool chain

A fabricated identifier rarely stays contained to the step where it was invented. Once it is produced, it becomes a valid-looking input to whatever tool call comes next, and nothing in a typical agent chain stops to ask where that input came from. Every downstream action builds on it as the ground truth for whatever comes next.

The math is unforgiving. A single bad step rarely stays a single bad step.

A documented case involving a fabricated station identifier, referred to as BOND-1, shows how this plays out concretely. The compounding here is a chain of decisions, each one reasonable given what came before it, built entirely on a foundation that was invented at step one.

A framework for localizing exactly where a failure originates, distinguishing a model decision from a tool response from harness logic, matters directly for this kind of compounding. Without a way to trace back to the true origin point, teams end up debugging the symptom at step five, leaving the cause at step one untouched.

Why a 200 status code hides this failure mode

A fabricated identifier that happens to match a valid schema returns a 200. The database accepts the query. The tool call completes. The agent's run finishes. Every dashboard a team is watching stays green, because nothing about the request looked malformed to the system that received it.

Traditional application monitoring tools are built on an assumption that does not hold here, that failures produce errors. Fabricated identifier returns produce none of that. The failure is semantic rather than structural, and tools built to catch structural failures have no mechanism for catching a semantic one.

The BOND-1 case shows the pattern concretely: the SQL ran cleanly against a real database and returned a real dataset, just the wrong one. The analysis went out to the end user as if it were authoritative, because by every measure the monitoring stack had access to, it was. A fabrication that returns a clean 200 leaves no such trace unless something was built specifically to look for it.

What traces need to capture to make fabricated identifiers visible

Catching this failure mode starts with recording not just what an agent did, but what it was supposed to be able to do. A trace needs the tool name, the full arguments passed to it, the result of validating those arguments against the schema, and a summary of what came back, enough information to later ask whether the identifier used in a call was ever actually present somewhere upstream, or whether it simply appeared.

The OpenTelemetry GenAI specification gives a structure for this, spanning six layers: client spans, agent and workflow spans, MCP conventions, semantic events, metrics, and provider-specific attributes. The current Model Context Protocol specification documents W3C Trace Context propagation inside MCP request metadata, using fixed key names, traceparent, tracestate, and baggage, which gives every tool call in a multi-step trajectory a traceable identity it did not have before. OpenTelemetry v1.39 added MCP-specific semantic conventions, including mcp.method.name, mcp.session.id, and mcp.protocol.version. Together, these attributes are the minimum surface needed to tie a given tool call back to the session context that should have constrained what arguments it was allowed to use.

Production systems need to tell apart two very different kinds of failure. One is a tool call that failed at the transport or protocol layer, which ordinary logging captures without much effort. The other is a model that produced an invalid or fabricated request that still executed cleanly, which is invisible unless arguments are logged at the field level. ToolFailBench's Output-Fabrication category shows this concretely: detecting it requires comparing the content of a result against what should have been there, not just confirming that a call happened and returned a status code. Logging only the call and the status code misses the entire failure class.

Evaluating whether an identifier was retrieved or invented

Good traces are only half the job. Turning them into a detection signal means evaluating individual spans, not just the final answer a trajectory produces. A trajectory can end with the right output and still contain a fabricated intermediate identifier that happened, by chance, to resolve to something usable. Scoring only the end result misses it. Only span-level evaluation catches it.

The question to ask at each tool-call span is whether the identifier was derivable from something that came before it, a prior tool response, the user's own input, or an instruction given to the agent, or whether it showed up for the first time as an outgoing argument with no upstream source at all, a question simple to state and hard to automate well, and one that AgentHallu's finding on step localization as the hardest part of automated hallucination detection shows is hard to answer at scale. Getting it right requires ground-truth labels attached to individual spans, not just labels attached to whole trajectories.

Fault injection offers a useful complement to this kind of scoring. AgentCheck, a July 2026 method from Mazumder and Lia, replays an agent's tool responses from cache except for one deliberately perturbed fault, letting later tool calls go live once the agent's behavior diverges from the clean run. Comparing the clean trajectory against the faulted one reveals something a passive trace never would: whether the agent, faced with a timeout or an error, fabricates a plausible-looking answer rather than surfacing the uncertainty honestly. That willingness to invent under pressure is what AgentCheck's fault taxonomy is designed to expose.

The loop that makes this scale is straightforward. Production traces that fail an argument-correctness scorer become eval cases. The eval suite grows out of real user behavior instead of being hand-built from imagined scenarios, and a regression, say a schema change or a model upgrade that reintroduces an old fabrication pattern, gets caught automatically the next time that eval runs. Online scorers watching live traffic catch fabrication as it happens. CI/CD gates running argument-correctness evals on every pull request catch it before a change ever reaches production. Production failures feed the regression suite, and the regression suite gates the next release.

Pre-defining every failure mode is not a viable detection strategy

The first instinct most engineering teams reach for is schema validation at the call site: reject anything that does not match a known identifier format before the call goes out. This works fine for catching malformed input. It does nothing for a plausible fabrication, because a fabricated UUID is, by construction, structurally identical to a real one. No validator checking shape alone can tell the difference.

An allowlist of every valid identifier sounds like a sturdier answer, but it breaks down at production scale. The BOND-1 case shows exactly where this fails: BOND-1 is a plausible station identifier that would sail through any format-based validator without triggering a single flag. Only a live check against the actual equipment hierarchy, a runtime lookup rather than a fixed rule, would have caught that it referred to nothing real.

A framework for localizing where failures originate, across model decisions, tool responses, and harness logic, underscores why static rules can never fully cover this space: the combinations shift every time a schema changes or a model gets updated, so a fixed rule set is always chasing a target that has already moved. Prompt engineering runs into the same wall from a different direction. The ICLR 2026 "Reasoning Trap" paper tested prompt-based mitigation for tool hallucination directly and found it narrowed the gap without closing it. A model trained under reasoning RL to fabricate will, under enough completion pressure, eventually override an instruction telling it to stay grounded. Detection has to be continuous and built on what actually happened in a trace, evaluated against what the agent was supposed to have access to, rather than checked against a fixed list or left to a prompt's good behavior.

A practical detection stack for engineering teams shipping agents today

A workable detection stack for this failure class needs three layers working together: argument-level tracing that records every identifier passed into every tool call, component-level evaluation that scores each span's arguments against the context available upstream, and alerting that fires before a fabrication has the chance to propagate into the steps that follow it.

Instrumentation comes first. Without that, there is nothing for an evaluator to compare against later, no matter how sophisticated the scoring logic is.

The evaluation layer then asks, for each span, whether the identifier was present somewhere in the prior context, derivable from an earlier tool response, or sourced from the user's own input. A span where the identifier appears for the first time as an outgoing argument, with no traceable origin behind it, counts as a fabrication candidate regardless of whether the call it was used in returned a 200. AgentCheck's fault injection approach, holding every tool response fixed except one deliberately perturbed fault, adds a harness-level check on top of this. Running agents against injected faults before release shows whether they fabricate under error conditions, catching the pattern before it ever reaches a real user.

The loop closes in CI/CD. A schema change or a model update that brings the fabrication pattern back gets caught the moment it reintroduces the behavior the eval suite already knows to watch for. Alerting belongs at the span level rather than the trajectory level: by the time a trajectory-level eval flags a problem, several downstream steps may already have executed against an identifier that was never real to begin with. Catching it at the span where it was invented, before it has a chance to become an input anywhere else, is what keeps a single bad call from becoming a chain of them.

Sources

  1. AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
  2. Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent Failures

More in Agent Failure Modes