What If AI Needs to See Before It Can Think?

For years, the AI story has been about words. ChatGPT writes essays. Claude summarizes meetings. Language models pass exams and debug code. It is easy to think language is the key.
But a new white paper asks a simple question: what if we have it backwards?
What if the real path to general intelligence runs through vision, not text?
Eyes Are Older Than Words
Complex eyes first appeared roughly half a billion years ago. Life learned to navigate, hunt, and survive based almost entirely on sight. Human language is brand new by comparison. Writing is only a few thousand years old.
This is the core argument of Visual General Intelligence: A White Paper, from researchers at OpenAI, Google DeepMind, Oxford, Stanford, Harvard, CMU, Cambridge, and more.
In modern AI, we treat vision as an add-on. A camera feeding data into a system that does the real thinking with language. But what if vision is not just an input? What if it is the foundation of how intelligence understands the world?
What Is Visual General Intelligence?
The paper introduces Visual General Intelligence (VGI).
We have seen what scaling language models can do. GPT was trained to predict the next word on massive text, and somehow learned to reason, code, and summarize. This raised a natural question: if scaling text led to general capabilities, what could emerge from scaling visual inputs like images, videos, and geometry?
That is what VGI is about. Can intelligence emerge from seeing the world, not just reading about it?
They are defining what computer vision should pursue in the AGI era.
Why Vision Might Matter More
Think about how a child learns. Before reading, they see.
They learn physics by watching, not reading. A toddler understands gravity before they know the word.
Language describes the world. Vision is direct contact with it. The paper suggests this difference matters more than we realize.
Language models know a lot about the world because they have read a lot about it. But they have never seen it.
The Bitter Lesson for Vision
The paper references Sutton's "Bitter Lesson": methods that scale with computation and data beat hand-engineered rules.
We saw this in language. Grammar rules gave way to neural networks that scaled. The authors suggest the same could happen for vision.
Imagine a model trained on massive video, images, and 3D geometry, learning to predict what comes next in a visual sequence the way GPT predicts the next word. Could it develop spatial reasoning? Physical intuition? The ability to plan in a real environment?
Nobody knows yet. But the paper says we should find out.
What This Means
If the VGI idea is even partly right, the implications are big.
The next generation of AI might need to spend more time interacting with the physical world and less time reading Wikipedia. Robotics, simulation, and video understanding would move from niche fields to the center of the AGI path.
Language models are impressive, but they might be one piece of a larger puzzle. The full picture may need systems that can see and experience the world the way biological intelligence has for half a billion years.
The Bottom Line
The authors are opening a conversation, not offering a final answer.
But if they are right, the most important frontier in AI might not be making language models bigger. It might be teaching machines to see.
And that would change everything.
Read the full white paper: Visual General Intelligence: A White Paper



