Home Knowledge Base Expert Profiles About Contact
LLM Wisdom
HomeKnowledgeExpertsAboutContact
Library

Knowledge Base

Browse our curated collection of insights, research, and practical frameworks.

Attention Sink Phenomenon: Why LLMs Systematically Fail on Long Contexts
knowledge
12 min

Attention Sink Phenomenon: Why LLMs Systematically Fail on Long Contexts

LLMs trained on short sequences exhibit a pathological attention pattern on longer inputs: a small number of 'sink' tokens absorb virtually all attention mass, starving the rest of the context. This is not a bug to be patched — it's a structural consequence of softmax normalization. Here's what the research shows, and how to work around it.

LLM Wisdom

Curating the most valuable knowledge, research, and expert insights in the world of large language models.

Explore

Knowledge BaseExpert ProfilesAboutContact

Categories

StudiesFrameworksInterviewsKnowledge

© 2026 LLM Wisdom. All rights reserved.