Thought
Published:
I believe pretraining for large language models is far from finished. A lot of work today focuses on building agents and on post-training. Those steps can help on specific tasks, but the gains are often narrow. The deeper abilities we still need—making sense of the world and dealing with messy real situations—depend on large-scale pretraining with careful data and real changes to how the model is built. That includes new versions of the basic Transformer design, memory inside the model that ties ideas together more cleanly, and flexible ways to blend representations when the setting shifts.
When we work with models that see images and text together, or with systems that handle vision, language, and action, the lessons from text-focused LLMs carry over in a natural way to multimodal LLMs and to world models. That path should give tighter alignment across modalities and stronger skills for agents that act in the physical world.