Distinguishing Model Errors From Tool Errors in Agent Traces

Trace spans reveal which layer failed, separating model bugs from tool configuration problems.

Contributing Editor · · 11 min read
Cover illustration for “Distinguishing Model Errors From Tool Errors in Agent Traces”
Trace Analysis · October 8, 2026 · 11 min read · 2,519 words

A failing agent trace shows the same thing every time: a wrong answer, a dropped task, a confident response that doesn't match reality. What it doesn't show is which layer caused it, and that omission is the whole problem. An engineer who fixes the model when the fault sits in the tool layer, or rewrites a tool schema when the fault sits in the model's reasoning, has spent real engineering time on a repair that will not hold the next time the trace fails the same way.

Why the Same Trace Symptom Points to Different Fixes

The same final failure state, a wrong answer returned to the user, can come from a model that hallucinated tool arguments, a tool that returned malformed data, or a model that received correct tool output and then ignored it. All three produce an identical symptom at the system boundary. All three demand a different repair, post-training or prompt revision for model-side failures, schema guards, retry policy, or tool description repair for harness-side failures, and applying the wrong one leaves the trace to fail the same way again on the next run because the actual defect was never touched.

This is the exact gap the "Model or Harness?" taxonomy from Scale AI was built to close. Its authors observe that existing evaluations reduce agent failures to system-level outcomes, obscuring where the fault originated and which intervention would actually improve the next iteration of the system. You need to assign repair, not just measure, and that sits upstream of every fix an engineering team might attempt.

A production illustration makes the stakes concrete. In one deployment, a tool call started returning malformed JSON, and the agent silently continued operating on the bad data. At the same time, a prompt that had worked reliably on one model behaved differently on another. Without a way to attribute fault, the team had no way to tell whether the problem was retrieval, the model itself, or an external API returning degraded responses. Three plausible root causes, one visible symptom, and no trace discipline in place to separate them.

The Model Layer and Tool Layer in a Trace

A production agent trace is a span hierarchy. The model layer (the planner, reasoner, and generator) produces one class of spans; the tool layer (external function calls, APIs, and MCP servers) produces another. Fault origin always sits on one side of the edge connecting them, and reading a trace well starts with knowing which spans belong to which side.

Each span carries its own attributes: input tokens, output tokens, latency, model name, status code, error type. That's the raw material fault attribution works from, and none of it is useful until it's organized around the right unit of analysis. When Agent A hands off to Agent B, the child span links back to the parent, preserving the full execution tree. This span hierarchy is what makes root-cause analysis possible instead of guesswork, because it lets an engineer walk backward from a visible failure to the exact point in the sequence where things went wrong.

The "Model or Harness?" taxonomy treats interactions between components, not components in isolation, as the unit of analysis. Each failure gets assigned to an edge between two components and a fault side indicating where the repair belongs: model-side failures target post-training, harness-side failures point to scaffolding and tool-integration fixes. This structure holds across architectures. Coding assistants, long-horizon personal assistants, and multi-agent systems look nothing alike on the surface, but they all share the same edge structure between model and tool.

Instrumentation has to match this structure to be useful. A trace without tool spans hides the most common failure modes, so every tool needs instrumentation, not just the entry agent calling it. Sampling strategy matters just as much: tail-based sampling should keep full failing runs along with a small fraction of successful runs, so the failure-side spans needed for attribution are actually available when an engineer goes looking for them.

The four tool-side failure modes and their trace signatures

Diagram: Four Tool-Use Failure Modes and Their Trace Signatures. Visualizes: Show the four ToolFailBench failure modes as a ranked or stepped visual, each paired with its defining trace signal.

ToolFailBench, a 2026 paper covering 1,000 single-turn tasks across five professional domains, defines the sharpest available sub-taxonomy for tool-use failures. It separates three failure modes for tasks that require a tool, Tool-Skip, Result-Ignore, and Output-Fabrication, plus one control-task label, Unnecessary-Tool-Use. All four look similar under aggregate accuracy scores, so reading them at the span level is necessary to distinguish them.

Tool-Skip occurs when the model never produces a valid executed tool call. The trace signal is a missing tool span where the task required one. Result-Ignore occurs when the model calls the tool correctly but doesn't use what comes back. Here the trace shows a tool span with a valid return, followed by a model response that contradicts or simply disregards it. Output-Fabrication occurs when the model calls the tool, gets a valid return, and then extends its response with invented structured information that was never in that return. The tool span looks clean, but the model span downstream does not match it. Unnecessary-Tool-Use is the control-task failure: the model calls a tool when the task should have been answered directly, leaving a tool span attached to a task that never needed one.

None of these four labels are safe to treat as automatically model-side, and that qualification matters enough to repeat before moving further: Tool-Skip, Result-Ignore, and Output-Fabrication can all originate on either side of the model-tool edge. They surface inside the tool span, but the root cause can be a defective tool description rather than model misbehavior, and the next section works through that distinction in detail.

Two further patterns round out the tool-side picture. Wrong-args failures, where the argument shape fails the tool's schema and the agent retries with the same malformed shape, appear in the trace as repeated tool spans carrying identical malformed arguments and consistent error status codes. No-error-handling failures, where the tool returns a 500 and the agent fabricates a response on top of it, appear as a tool span with an error status followed by a downstream model span that makes no reference to that failure. The agent declared success while the tool layer had already failed, and nothing in the final output hints at the gap.

Why bad tool descriptions cause tool-layer failures that traces misread as model errors

A model that calls a tool with the wrong arguments because the tool's description was ambiguous or incomplete is not making a reasoning error. The defect sits in the specification the harness handed the model, and the fix belongs in the tool description, not in the model.

MCP has become the dominant protocol for tool integration, and its tool descriptions are the only specification the model receives about what a tool does, what arguments it expects, and where its limitations lie. Research into MCP tool description quality found that 97.1% of tool descriptions have at least one quality problem, with Unstated Limitations, Missing Usage Guidelines, and Opaque Parameters as the dominant failure types. The same research found that 56% of the 856 tool descriptions studied exhibit an "Unclear Purpose" smell, failing to state the tool's intended functionality clearly enough for the model to act on it correctly.

The trace fingerprint of a description-origin failure looks deceptively like a model error. The model span shows a coherent reasoning chain. The tool call arguments are wrong, or the wrong tool gets selected. Replaying the same trace with a repaired description resolves the failure without touching the model at all, which is the clearest possible evidence that the original fault never lived in the model's weights or prompt. This is the core misattribution risk running through the whole exercise of trace reading: a wrong tool selection or a malformed argument set looks exactly like a planning failure, but the actual repair surface is the harness, specifically the tool descriptor handed to the model at call time.

The exposure extends past description quality. The MCP protocol has no native defenses against tool poisoning, rug pull attacks, or cross-server tool shadowing, and each of these attack vectors manifests in a trace as a wrong-tool-selected or wrong-args failure that reads, on the surface, like a model mistake. Separating these cases from genuine model error requires a way to reproduce the fault, change one variable, and check whether the failure disappears. AgentCheck, a reproduce-intervene-mitigate workbench for LLM agents running over MCP, does exactly that: it replays a fault, applies a mitigation such as a repaired description, and verifies whether the failure clears, isolating description-origin failures from model-origin ones directly.

The model-side failure modes and their trace signatures

Model-layer failures carry their own trace signature, and it's the inverse of the tool-side pattern: the tool spans are healthy, with a correct return and a valid status, and the fault lives in what the model decided to do before the call or what it concluded after receiving the tool's valid output.

The "Model or Harness?" framework organizes 41 distinct failure modes this way, assigning each one to an edge between two components and a fault side. If a failure is model-side, post-training or prompt revision is the relevant fix; if it's harness-side, scaffolding changes are. Planning errors are model-side: a wrong tool sequence emitted by the planner span, an infinite loop where the same tool fires repeatedly with near-identical arguments and no advance in goal state, premature termination, or plan-vs-execute divergence, where the plan the model emits doesn't match the tool calls it actually makes. In the infinite-loop fingerprint, the step count climbs without the goal state ever moving forward, and the tool span repeats itself almost verbatim. Plan-vs-execute divergence occurs when the plan span's tool sequence and the executed tool span sequence diverge, a pattern that also serves as the canonical fingerprint of indirect prompt injection operating at the plan layer.

Reasoning errors are the other model-side category: the model receives correct, grounded context from the tool layer and still produces a response that's factually wrong, mathematically incorrect, or simply ignores an instruction it was given. The Air Canada chatbot case is a documented instance of this failure class, now cited in "Towards a Science of AI Agent Reliability"; as a real-world example of the gap between average performance and reliable operation. The agent gave incorrect information about bereavement fare eligibility; the tool or retrieval layer may well have returned valid policy data, with the model misreading it or fabricating content on top of it.

Result-Ignore and Output-Fabrication, introduced earlier as tool-use failure labels, reappear here because they can originate on the model side just as easily as the tool side. When the tool's return is valid and complete but the model ignores it or extends it with invented content anyway, the fault sits in how the model handled the output it was given. The diagnostic rule that separates this from a tool-side cause is simple to state and consistent across both sections: if the tool span shows a clean return with valid content and the model span downstream contradicts or extends that content without grounding, the fault is model-side, regardless of which failure-mode label technically applies to the trace.

Silent Success Declarations and Fault Attribution

Diagram: False Success: How Often Agents Declare Done When They Aren't. Visualizes: Show two stark magnitude callouts side by side: in single-control tau-bench settings (airline and retail domains), roughly 50% of all failures were false successes…

The most dangerous failures on either side of the model-tool edge produce no exception. No error span fires. The agent declares success while the actual state of the world contradicts that claim, and the trace carries no visible signal that attribution is even needed.

A study of tau2-bench trajectories across several model families and task domains found that false success, when the agent declares a task complete that it didn't correctly execute, accounted for roughly half of all failures in single-control settings covering the airline and retail domains. On AppWorld, false success accounted for 75.8% of failures among agent architectures that make explicit completion claims, measured against actual database state. That gap between self-report and ground truth confirms that an agent's own claim of completion cannot be trusted as a signal on its own.

Detecting these cases turns out to be harder than it looks for a specific reason: LLM-as-judge approaches failed to reliably separate false successes from genuine ones in these studies, while a simpler retrieval-based detector substantially outperformed LLM judges on this exact failure type. That has a direct implication for how teams build attribution pipelines, since the obvious instinct, asking a model to judge another model's success claim, performs worse than a narrower, more mechanical check against ground truth.

Silent success is a model-side failure at the point where it surfaces, since the model is the one asserting completion, but its root cause can still sit on the tool side: a tool that silently swallowed an error and returned an empty success response the model had no reason to distrust. One bad tool call can poison several steps that follow it, and an agent that keeps building confidently on a wrong answer produces a span tree where the visible failure sits far downstream from where it actually started, making the tree harder to walk correctly the later the failure is caught.

The multi-agent version of this problem is structural. MAST's analysis of 150 annotated traces found that specification and design issues make up the largest share of multi-agent failures, and inter-agent misalignment makes up the next largest share. Most multi-agent failures are architecture-level problems, built into how agents hand work to each other, and they can produce silent bad outcomes without ever tripping an error condition.

A practical walk-the-trace workflow for assigning fault to the correct layer

Reliable fault attribution follows a fixed order through the span tree, and that order matters: tool spans get checked before model spans are judged, because tool-layer evidence is more deterministic and clears away the largest class of misattribution before the harder model-side judgment calls are needed.

Walk the span tree in execution order, following the sequence input, plan, retrieval, tool args, tool output, LLM rewrite, final answer, and score each span before moving to the next one. Check the tool spans first: a missing span where one was required points to Tool-Skip; a clean return followed by a contradicting or ignoring model response points to Result-Ignore; a clean return followed by invented content points to Output-Fabrication; a tool span on a task that needed no tool at all points to Unnecessary-Tool-Use. Where the tool span itself shows an error status, malformed output, or a repeated identical argument structure, the fault sits with the tool or its description before any model judgment is required.

Where tool spans look clean, the harder question follows: is the fault still tool-side, through a defective description that misled a reasoning model that otherwise behaved coherently, or is it genuinely model-side, a planning error, a reasoning error, or a case of the model discarding or embellishing valid tool output on its own. Replaying the trace with a repaired tool description, in the manner AgentCheck enables, answers that question directly. Because silent success breaks the assumption that a trace failure announces itself, every completion claim in the tree needs to be checked against actual state, not against the agent's own report of what happened, before the trace can be marked resolved.

Sources

  1. ToolFailBench: Diagnosing Tool-Use Failures in LLM Agents
  2. Model or Harness? An Interaction-Centric Taxonomy forLocalizing Agent Failures
Filed underTrace Analysis