Unlike LLMs, which function as "wordsmiths in the dark," world models are designed to understand spatial intelligence, physical consistency, and the consequences of actions within an environment.
This shift has sparked a billion-dollar funding flurry for startups like World Labs and Advanced Machine Intelligence (AMI), as well as new projects from established players like Google DeepMind and Runway.
The technical development of world models currently centers on two primary approaches: generative video models and explicit 3D representations.
Some researchers follow the "bitter lesson," arguing that models should learn physics and geometry implicitly from massive amounts of video data through scale.
Others advocate for hybrid systems that output 3D assets like Gaussian splats, which can be integrated into existing creative workflows for gaming and film.
Beyond entertainment, a major driver for this technology is the "robot training data problem." By creating accurate physical simulations, developers hope to provide the vast amounts of synthetic data needed to train humanoid robots for complex, real-world tasks that are currently difficult to teach via manual operation.