Back to Knowledge
Framework
13 min read August 21, 2026

Your Agent Didn't Fail at the Tool Call. It Failed at the Plan.

End-to-end agent benchmarks hide the real failure. APB's 4,209-case diagnostic shows planning quality is not one capability — and extra tools make strong models worse.

Alex Cheeseman

Editor, LLM Wisdom

Overview

When an LLM agent misses a task, the post-mortem is almost always written at the tool layer: wrong API, bad JSON, flaky environment, the model “didn’t call search.” That story is convenient because it is visible in the trace. It is often wrong.

Sun, Wang, Song, He, Zhang, Liu, Yang and Cheng (arXiv:2606.04874, June 2026) introduce Agent Planning Benchmark (APB): 4,209 multimodal cases, 22 domains, five settings, 12 models. The point of the benchmark is diagnostic rather than leaderboard-shaped. It asks whether the plan was valid before anyone debates whether the tool executed.

The headline findings for practitioners:

  • Planning is not a single skill. Long-horizon holistic planning, step-wise planning with feedback, and robustness under tool noise come apart.
  • Newer proprietary models (GPT-5, Gemini 3 Pro, Claude Sonnet 4.5) dominate holistic planning; GPT-4o is not in the same class (19.5% holistic correctness vs GPT-5’s 74.5%).
  • Extra tools are not free optionality. On GPT-5, holistic correctness falls from 74.5% to 67.0% when extraneous tools are present, and tool-use errors (E5) jump from 2.3% to 15.8%.
  • Constraint violation (E3) is the dominant error even for frontier models — 16.4% on GPT-5 holistic plans.
  • Inference-time refinement helps long-horizon plans and can over-correct short-horizon ones.

If you are shipping agents and only tracking task success, you cannot tell a planner problem from an executor problem. APB is the framework for splitting them.

✦

The Study

Existing agent evaluations (ToolSandbox, τ²-bench, SWE-bench-class tasks) report end-to-end success. That number entangles plan quality, tool invocation, environment instability and recovery. A failed run is uninformative; a passed run can still have been a lucky execution of a bad plan.

APB isolates planning. Models must decompose goals, select tools, respect constraints, and decide when a task is infeasible — without the alibi of a broken sandbox. Five settings:

  1. Holistic planning — produce the full trajectory up front
  2. Step-wise planning — plan the next step given feedback
  3. Extraneous tools — the catalogue contains tools the task does not need
  4. Broken tools — some tools are unavailable or fail
  5. Unsolvable tasks — contradictory constraints, missing information, inaccessible evidence, or required tools removed

Scoring is hierarchical:

  • Plan Correctness (CR) — did the plan solve the task logically?
  • Plan Grade — severity, not just binary pass
  • E1–E6 taxonomy — why it failed

The authors validate that this is not academic. On 200 ToolSandbox and 200 τ²-bench tasks, APB-guided refinement improved plan correctness, plan grade, and downstream execution for GPT-4o, Qwen3-VL-235B-A22B and Gemini 2.5 Flash.

✦

The Error Taxonomy

This is the part worth adopting even if you never run APB.

| Code | Name | What it looks like in production | |---|---|---| | E1 | Goal understanding error | The agent solves a neighbouring task | | E2 | Premature conclusion / incompleteness | Stops early, or omits a required sub-task | | E3 | Constraint violation | Ignores budget, policy, format, or system-prompt rules | | E4 | Logic error | Wrong order, missing prerequisite, broken causal chain | | E5 | Tool-use error | Calls the wrong tool, or the right tool with the wrong semantics | | E6 | Hallucination | Invents state, evidence, or intermediate facts |

E5 is what most teams already log. E1–E4 are planning. E3 is the silent killer: the trace looks fluent, the tools are valid, and the policy is already broken.

✦

What They Found

1. The capability gap is a generation gap

Holistic planning correctness (CR):

| Model | Holistic CR | Grade | |---|---|---| | GPT-5 | 74.5% | 0.91 | | Gemini 3 Pro | 71.3% | 0.89 | | Claude Sonnet 4.5 | 64.1% | 0.86 | | Gemini 2.5 Pro | 55.0% | 0.81 | | Gemini 2.5 Flash | 36.5% | 0.69 | | GPT-4o | 19.5% | 0.62 |

GPT-4o is not “a bit worse.” It is a different reliability class. Shipping a 4o-class planner on long-horizon work and expecting GPT-5-class behaviour is a category error. Open-source MLLMs in the study cluster with GPT-4o or below on holistic CR (InternVL and Qwen3-VL variants: 8–23%).

2. Extra tools punish planners that have to choose

Holistic CR with a noisy tool catalogue:

| Model | Clean holistic CR | Extraneous-tool CR | E5 (tool error) clean → noisy | |---|---|---|---| | GPT-5 | 74.5% | 67.0% | 2.3% → 15.8% | | Gemini 3 Pro | 71.3% | 76.4% | 4.7% → 8.8% | | Claude Sonnet 4.5 | 64.1% | 76.4% | 3.2% → 6.7% | | GPT-4o | 19.5% | 45.5% | 15.2% → 18.4% |

Read this carefully. Frontier models that already plan well can get worse when you give them a bigger toolbox — GPT-5 is the clean example. Weaker models sometimes improve on the noisy split because the task distribution and scoring interact differently; they do not become reliable. The practical rule is not “more tools = more capable.” It is “every unused tool is a planning distractor.”

Step-wise planning is more robust to this noise than holistic planning, which is why “just make it ReAct” is an incomplete mitigation: you have traded one failure mode for a shorter-horizon one.

3. Constraint violation is not a tail error

On clean holistic planning, E3 (constraint violation) is 16.4% for GPT-5, 16.7% for Gemini 3 Pro, 24.4% for Claude Sonnet 4.5, and 43.7% for GPT-4o. These are not jailbreaks. They are ignored budgets, skipped prerequisites, and system-prompt rules that did not survive contact with a long plan.

If your safety story is “the model is aligned, so the agent is aligned,” E3 is the counterexample. The model can be polite and still plan an action the policy forbade.

4. Unsolvable tasks expose calibration, not knowledge

APB includes 400 unsolvable cases built from contradictory constraints, missing information, inaccessible visual evidence, and tool removal. Models detect explicit contradictions more readily than implicit information gaps. That is the production incident you already have: the agent proceeds because nothing in the prompt said “this is impossible,” even though a required input was never supplied.

Calibrated refusal is a planning skill. It is not the same as sycophancy-avoidance, and it is not the same as safety refusal.

5. Think longer on long plans. Do not think longer on the next click.

Inference-time refinement — extra reflection before committing the plan — helps holistic planning. On short-horizon step-wise decisions, extended reflection can over-correct: the model talks itself out of a valid next step. This is the agent analogue of the 2026 finding that more reasoning tokens are not monotonically better (see follow-on reading). Test-time compute is a lever, not a virtue.

✦

Why This Happens

Three structural reasons, none of them mysterious:

  1. End-to-end RL and SFT teach traces, not plans. A successful demonstration contains a plan, but the loss does not isolate it. Models learn to look like they are acting. APB’s split is what you get when you finally grade the plan.

  2. Tool catalogues are an untrained decision problem. Pre-training does not contain “here are 40 tools, 36 of which are irrelevant.” Extra tools increase the branching factor of the plan. GPT-5’s E5 spike under extraneous tools is the model spending planning capacity on selection instead of decomposition.

  3. Soft constraints lose to fluent next-token planning. E3 is high because constraints live in the system prompt and the plan is generated as narrative. Narrative generation is good at local coherence and bad at global constraint satisfaction — the same reason long system prompts fail in lost-in-the-middle settings.

✦

Concrete Recommendations

1. Log E1–E6 on every failed (and sampled successful) run

Do not start with a new benchmark. Start with the taxonomy. Tag 50 recent failures. If they cluster in E3/E4, stop tuning tools. If they cluster in E5, then tune tools. If they cluster in E2, you have a stopping-rule problem — related to, but distinct from, context-rot premature termination.

2. Shrink the tool catalogue per task class

Give the agent the tools the task needs, not the tools the platform has. GPT-5 losing 7.5 points of holistic CR when extras are present is the cost of a “universal agent.” Route to specialised catalogues.

3. Put hard constraints in code, not in the system prompt

E3 will not be prompt-engineered away at 16%+ on frontier models. Pre-execution gates — schema checks, policy checks, budget checks — belong in the runtime. The plan can propose; the gate decides.

4. Separate planner and executor

APB-guided refinement improved execution on ToolSandbox and τ²-bench. The cheap version: generate a plan, critique it against E1–E6, then execute. The expensive version: a stronger model plans, a cheaper model acts. Do not use extra reflection on every micro-step.

5. Test unsolvable and broken-tool cases on purpose

Your eval set is almost certainly all solvable. Add: missing fields, contradictory user requests, disabled tools. Score refusal quality, not just success. An agent that cannot stop is not autonomous. It is a loop.

6. Do not buy a model on SWE-bench and deploy it as a planner

Holistic CR in this paper is the number that should gate “can this model own a long-horizon workflow.” GPT-4o-class scores are an executor, not a planner.

✦

How This Extends Prior LLM Wisdom

Sycophancy explains why models agree with you. Lost-in-the-middle and context rot explain why they ignore the middle of a long trace. APB explains why a well-retrieved, non-sycophantic agent still does the wrong sequence of things. Planning is the missing layer between “the model can answer” and “the agent can be trusted.”

✦

Follow-On Research

  • Sun et al. (2026). Agent Planning Benchmark: A Diagnostic Framework for Planning Capabilities in LLM Agents. arXiv:2606.04874. Code: github.com/Mikivishy/AgentPlanningBenchmark
  • Yao et al. (2022). ReAct: Synergizing Reasoning and Acting in Language Models.
  • Shinn et al. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning.
  • Related: The relationship between reasoning and performance in large language models (Scientific Reports, 2026) — more reasoning tokens are not monotonically better.
✦

Recommended Reading

  • arXiv:2606.04874
  • LLM Wisdom: Sycophancy in LLMs
  • LLM Wisdom: Context Rot (companion piece)
  • LLM Wisdom: The Lost in the Middle Problem
agents
planning
tool use
APB
GPT-5
evaluation
constraint violation

Recommended Readings

More from Framework