AI's real-world-data gap — "LLMs are blind"; grounding via Earth-observation and robots
AI's real-world-data gap — "LLMs are blind"; grounding via Earth-observation and robots
Vintage: 2026-06. Primary source recorded 2026-06-06 (All-In IPO panel, Will Marshall); a multimodality/data-frontier capability argument — treat as a snapshot.
One-line summary: Today's frontier models are trained on the text of the internet and are "blind" to the physical world; the argument here is that the next capability frontier is grounding models in real-world data — daily-global Earth observation ("large earth models"), and robot/sensor streams — because "AI is only as good as the data it's trained upon." It's the data-supply face of the multimodality / world-models question.
The insight
The data wall for text is in view (internet text is largely consumed); the proposed next reservoir is physical-world data that LLMs currently lack. will-marshall frames it sharply: text-trained models "don't know shit about the real world" — they can't see the flooded farm field or the security situation around the corner. Two grounding channels recur in the sources: (1) Earth observation — a daily-refresh global imaging archive as a training/inference substrate ("large earth models" rather than large language models); (2) embodiment — humanoid robots and sensors as real-world data collectors. The capability claim is that grounding unlocks "real-world problems" current models can't touch; the open question is whether a grounded-data layer actually moves frontier capability or just adds a vertical application.
Evidence
- will-marshall in 2026-06-06-podcast-all-in-podcast-the-ipo-comeback-why-tech-giants-are-finally (June 2026): "all the cool stuff that we're doing with LLMs now is really based on just the text of the Internet … But they don't know shit about the real world. I call them blind … If you give them real world data, then they can answer real world problems … instead of having large language models, large earth models." And the load-bearing premise: "AI is only as good as the data it's trained upon."
- will-marshall in 2026-06-06-podcast-all-in-podcast-the-ipo-comeback-why-tech-giants-are-finally (June 2026): the data asset is a daily-refresh, full-history global archive ("the Google satellite layer … except it's today's date rather than three years old … every day going back") — a time-series of the whole Earth as a grounding corpus.
- The embodiment channel (source-attributed, 2026-06-01-podcast-moonshots-opus-4-8-beats-gpt-5-5-the-220b-openai-foundation, Diamandis, May 2026): "humanoid robots interacting in the real world are going to be an important source of data. Everybody's talking about where do I get new data? Well, this is definitely one of them." The OpenAI-robotics pivot (SORA video team → robotics under Aditya Ramesh) is read partly as a data-acquisition play.
Why it matters to this thread
- It's the data-supply complement to the benchmark-saturation story: if known-answer text benchmarks saturate, grounded real-world data is one route to the "unsolved-problem" frontier.
- It connects multimodality/world-models (in scope) to a concrete supply question — who owns the physical-world datasets, and do frontier labs need them.
Contradictions / tensions
- Single, self-interested primary source. Marshall runs the Earth-imaging company that benefits if "large earth models" become real; the claim that frontier capability is data-bound on physical-world data (vs architecture/compute) is asserted, not demonstrated.
- It's unclear whether grounded data lifts general capability or only powers vertical applications (agriculture, defense, climate) — the markets version of the same question is ai-real-world-data-gap-to-planet-moat (a stock-market mechanism, PL).
- Competes with the synthetic-data view (frontier gains increasingly from synthetic/self-play data, not new human/real-world corpora — cf. Wissner-Gross on pretraining-tokens being a smaller share of frontier gains).
Related
- will-marshall — primary source
- ai-real-world-data-gap-to-planet-moat — the stock-market (tradeable, PL) version of the same argument
- agi-definitions-and-benchmark-saturation — grounded data as a route past saturated benchmarks
- terrestrial-power-flat-to-orbital-dc-arbitrage — adjacent AI-infra (where the compute to use this data may eventually live)