Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

Admin

Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

Google researchers have introduced knowledge profiling, a behavioral framework designed to separate two different reasons why large language models get factual questions wrong: the fact was never encoded, or the fact is encoded but not accessible. The work appears in “Empty Shelves or Lost Keys? Recall Is the Bottleneck for Parametric Factuality” and examines frontier LLMs including Gemini3 and GPT-5.

The researchers say standard accuracy metrics combine encoding failures and recall failures, even though they point to different limits and different fixes. Encoding failures may suggest a need for larger models or broader data coverage, while recall failures may be addressed with post-training or inference-time methods that help models use what they already store.

In the paper, encoding refers to the parametric representation of facts, recall refers to retrieving encoded facts without external cues, and recognition refers to identifying the correct fact when it is shown among alternatives.

Using the framework, the researchers say many factual errors in frontier LLMs are better understood as lost keys rather than empty shelves — in other words, as recall failures rather than encoding failures.

To support the analysis, the team also introduced WikiProfile, a benchmark of 2,150 Wikipedia-derived facts, each paired with ten questions that test encoding, recall, and recognition.

Source: research.google.

Companies can share verified announcements through Newz9’s international press release submission page.