World Labs has recently launched Atlas, a new "frontier model" that has received an "amazing reception." Feifei, Justin, and Ben discussed its capabilities, significance, and future implications. Atlas is described as a "next-generation world model" that can generate, reconstruct, and simulate the world. Its core innovation is "new view prediction," differentiating it from LLMs' "next token prediction" and video models' "next frame prediction."
Atlas takes a given number of views of a scene or a description and creates a "spatial context." From this context, it can predict what the world should look like from any arbitrary virtual camera position in space and time. This enables capabilities such as "camera condition generation," where users can input an image and a camera trajectory to generate video frames from any desired perspective. It also excels at "sparse 3D reconstruction," allowing the reconstruction of the real world from as few as one to up to a hundred frames, which can then be output as a novel video or an explicit 3D model. For simulation, Atlas can create "bullet time videos," reminiscent of the Matrix, and robotic simulations.
A significant breakthrough Atlas achieves in sparse 3D reconstruction is reducing the number of cameras needed for a complex shot from hundreds (as in the original Matrix "bullet time" scene) to just three. This eliminates the need for studio capture, green screens, or expensive calibration, representing a "50-100x reduction" in data input compared to traditional "dense reconstruction." Ben, creator of NeRF, explained that dense reconstruction traditionally required capturing every single point from multiple angles, leading to tedious, hours-long efforts to avoid holes in the 3D model. Atlas, by contrast, uses generation to "imagine" and fill in gaps where data is sparse, making 3D reconstruction much more accessible and efficient.
Feifei highlighted that Atlas represents the first time there's been a "unification of pixel generation and pixel reconstruction" in a single model, a problem that historically has been tackled by separate subfields in computer vision. Atlas is natively multimodal, designed from the start to work with text, images, videos, camera poses, and 3D data (like depth maps) as integrated inputs. The team expressed strong conviction in the scaling hypothesis, noting that making the model bigger and training it longer consistently improved performance, suggesting they are "at the beginning" of its potential, currently limited only by compute resources.
Regarding applications, Atlas extends World Labs' previous product, Marble, by offering a more robust and unified approach. For creatives, it provides a 3D-consistent grounding for generations, enabling precise control for storyboarding, movie production, and game environment design. For designers and architects, it streamlines the virtual design phase, allowing quicker iterations and feedback by translating verbal or sketch-based ideas into concrete 3D models.
A crucial application area is robotics. Feifei explained that Atlas is key to the "real-to-sim" pipeline. Robotics training is data-intensive, often requiring vast amounts of real-world data and randomization in simulated environments. Atlas's sparse reconstruction capabilities dramatically accelerate the process of creating simulated environments from real-world scans, replacing the "excruciatingly painful" dense reconstruction methods. This helps overcome the data bottleneck in robotics, allowing for the development of "neural simulators" that can understand how the world responds to actions, thereby improving robotic policy training.
Looking ahead, the team acknowledged that dynamics are a key frontier. Atlas already incorporates "baby dynamics," learned from its training data, allowing for elements like moving cars or water waves in its outputs, which was not possible with the static Marble model. Further development will focus on enhancing dynamic capabilities, exploring "4D video" (3D + time), and expanding "multimodal control" and editability. The ultimate goal is to allow users to interact with and control various aspects of the generated scenes, from layout to object identity and time. The speakers strongly believe that "new view prediction" is "AI complete," analogous to next token prediction, fundamentally aligning with how nature gave animals eyes to experience new viewpoints.