Govern Agent Access. Don’t Guess. (Sponsored)AI agents are moving from demos into production, and they need more than a model. They need real-time data, fast answers, and strict limits on what they can read, write, and act on. At the Agentic Data Summit on December 9, the session Guardrails, Not Guesswork shows how to lock down agent access with the Agentic Data Plane. You’ll also see how streaming, SQL, and agents run on one platform, and get a look at where the Agentic Data Plane is headed next. It’s free, virtual, and built for engineers taking AI agents to production. We can provide an LLM with exhaustive information through our prompts, but it can still fail to use it properly. When the required information sits somewhere in the middle of a long prompt, this type of failure becomes even more likely. We call it the “lost in the middle” effect or an LLM blind spot. For example, imagine that we give an AI-based coding assistant a long collection of project documents. One specific paragraph explains that audit logs must be retained for 37 days. However, the coding assistant writes a cleanup function that deletes them after 30 days. The relevant rule was mentioned clearly. It was also part of the prompt and fit the model’s limits. Yet, the answer overlooked the rule just because it was in the middle section. In this article, we’ll look at why LLMs have this bias against middle information. Here’s what we will cover:
What Information is Available to the ModelWhen an application sends a request to an LLM, the input usually contains more than the user’s latest question. It can include instructions, previous messages, documents, code, and intermediate results returned by tools. Together, these form the model’s input context. The model processes all of this information as tokens, which are basically small units of text. A token might represent a word, part of a word, punctuation, or another text fragment. The context window sets a limit on how many tokens the model can handle. This window must also be able to accommodate the generated response. Let’s say a model supports a context window of 128,000 tokens. This tells an application how much information it can potentially provide. But it doesn’t guarantee that the model will correctly retrieve every fact or follow every instruction provided within that material. This is similar to the difference between a system’s capacity and reliability. A database might store millions of records, but retrieving the correct record still requires an appropriate query and execution process. Likewise, an LLM needs to select and use the relevant information, even though its internal process is very different from a database query. Does the LLM Have a BlindspotWe can test this by keeping a question and its supporting information unchanged while moving the supporting information to different positions in the input. For example, imagine a collection of project notes containing this fact: “The owner of Project Cedar is Adam.” The question is always, “Who owns Project Cedar?” In the first test, the relevant note appears first within the input prompt. In another test, it appears halfway through the collection. In a third, the note appears at the very end. The other notes remain the same. If accuracy changes substantially, the model is sensitive to where the evidence appears. A study named “Lost in the Middle”, released in 2023 and published in 2024, studied this behaviour and found that performance was often strongest near the beginning and end. However, performance in the middle was weaker. This produces a U-shaped accuracy curve. Here’s how it stacks up:
Of course, these are more like tendencies rather than strict guarantees. We can get wrong answers even from beginning and ending positions. Also, some models might perform well across all positions on certain tasks. Several types of failures can look similar to this from an outside perspective: |