❯ Fei-Fei Li’s World Labs releases Atlas, rebuilding 3D worlds from one photo and simulating robot views
one model, four jobsWorld Labs, founded by Fei-Fei Li, released the world model Atlas on September 1, calling it the first multimodal world model of its kind. It is a multimodal autoregressive diffusion transformer pretrained from scratch, operating natively on text, images, video and 3D, and treating world generation, 3D reconstruction and simulation as facets of one problem. Output runs up to 1440p and one minute per generation.
camera as native inputPer its technical blog, prior approaches largely treated camera pose as a post-hoc constraint; Atlas’s decisive choice is to take camera geometry as a native input rather than coax it out of a text prompt. That yields pixel-perfect camera control: from a single image, or a few ground-level photos, it reconstructs large scenes and outputs 3D natively, exportable as point clouds and Gaussian splats. On 3D reconstruction the company reports an error of 25.3‰, ahead of specialist models including Pi3X and Depth Anything 3 — a generalist beating purpose-built tools at their own task is the hardest claim here.
the robotics lineThe Real-to-Sim demonstration says more than the benchmarks: World Labs reconstructed two large environments from cell phone video, 24 frames each, then had simulated robots move through them while Atlas rendered the RGB and depth data an onboard camera would see. That collapses the cost of a robot training environment from building a physical set or hand-modeling down to a phone video. Interaction with rigid, articulated and deformable objects is inside the reconstruction as well. Nvidia’s Isaac GR00T simulation stack is the industry default today, and that is the ground Atlas is cutting into.
what was withheldAtlas is in early access with select partners, with a request form on the site, and will power future versions of Marble and other products. The launch came with no paper, no pricing, no latency or compute-footprint figures, and no partner names. Teams building robot data pipelines need exactly those numbers before scoping it in: the cost and latency of rendering a frame decide whether this is a stage in a pipeline or a demo.
▪ SIGNALA generalist world model beat the specialists at 3D reconstruction; the shopping list for robot training data just gained an option.