Encoding vs recall as the bottleneck for parametric factuality
Vintage: 2026-08 / 2026-02. Issuer blog Empty shelves or lost keys? is August 12, 2026; paper is arXiv:2602.14080 (February 2026 id). Finding is real. 24h framing in the Grok recap is recap-only and false. Hydrated in 2026-09-02-grok-com-ai-news-digest-2026-09-02-fable-5-1-astra-atlas-g20.
Encoding vs recall as the bottleneck for parametric factuality
One-line summary: Google Research + Technion (Calderon / Yona): for Gemini-3-Pro and GPT-5, 95–98% of facts are encoded, yet models still fail to directly recall 26–34%; thinking recovers roughly 40–65% of encoded-but-not-directly-known facts, not "~65%" of all misses. Real paper; not last-24h.
The insight
The recap presented this as last-24h research via a 7min chip ("recovered up to 65% of the facts the models could not directly recall"). The issuer blog is Aug 12; the arXiv id is 2602.* (February 2026). File the finding with its real date. Do not flatten the recap's 24h window or the single-point "~65%."
Evidence
- From 2026-09-02-grok-com-ai-news-digest-2026-09-02-fable-5-1-astra-atlas-g20 (Google blog, August 12, 2026 / arXiv:2602.14080): authors Nitay Calderon and Gal Yona (Google Research; Calderon's arXiv affiliation is Technion).
- From the same source: for Gemini-3-Pro and GPT-5, 95–98% of facts are encoded, yet models still fail to directly recall 26–34%; even with thinking they still fail on 11–12%.
- From the same source: in thinking-optimized models, thinking recovers roughly 40–65% of encoded-but-not-directly-known facts (not 65% of all missed facts, and not a single-point "~65%").
What this source does not establish
- Not last-24h. Recap 24h framing is false. Do not file this as September 1–2 research.
- Recap "~65% recovered" is the top of a 40–65% range on encoded-but-not-directly-known facts, not a single-point recovery of all misses.
- 7min chip is aggregator copy, not the issuer blog.
- No other recap-named 24h papers ("LLM scientific law discovery"; "construct-validity issues in LLM agent safety evaluations") were fetched under those titles. Do not invent abs URLs.
Open questions
- Does the 40–65% thinking-recovery range hold on models shipped after Gemini-3-Pro / GPT-5?
- How this interacts with agi-definitions-and-benchmark-saturation (benchmarks saturate while recall still fails) is not argued in this clipping.