Home Knowledge Base Expert Profiles About Contact
LLM Wisdom
HomeKnowledgeExpertsAboutContact
Library

Knowledge Base

Browse our curated collection of insights, research, and practical frameworks.

framework
13 min

Your Agent Didn't Fail at the Tool Call. It Failed at the Plan.

Sun et al. (2026) introduce Agent Planning Benchmark: 4,209 multimodal cases across 12 models. GPT-5 holistic plan correctness is 74.5% vs GPT-4o's 19.5%. Extra tools drop GPT-5 to 67.0% and spike tool-use errors from 2.3% to 15.8%. Constraint violation remains 16%+ even at the frontier.

The Lost in the Middle Problem: Why LLMs Ignore What You Put in the Centre of Your Prompt
framework
10 min

The Lost in the Middle Problem: Why LLMs Ignore What You Put in the Centre of Your Prompt

Liu et al. (2023) demonstrated that language model performance on multi-document QA degrades significantly when relevant information appears in the middle of a long context, even when models nominally support that context length. This 'lost in the middle' effect has direct, actionable consequences for how you structure RAG outputs, system prompts, and few-shot examples.

LLM Wisdom

Curating the most valuable knowledge, research, and expert insights in the world of large language models.

Explore

Knowledge BaseExpert ProfilesAboutContact

Categories

StudiesFrameworksInterviewsKnowledge

© 2026 LLM Wisdom. All rights reserved.