<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom"><title>Ju Lin's AI Weblog: World-Models</title><subtitle>An independent research notebook on AI engineering, agents, models and the systems around them.</subtitle><id>https://julin.ai/atom/tags/world-models/index.xml</id><link rel="self" type="application/atom+xml" href="https://julin.ai/atom/tags/world-models/index.xml"/><link rel="alternate" type="text/html" href="https://julin.ai/tags/world-models/"/><author><name>Ju Lin</name></author><updated>2026-09-03T00:00:00+12:00</updated><entry><title>Atlas: World Labs' Omni World Model</title><id>https://julin.ai/2026/09/03/atlas-world-model/</id><link rel="alternate" type="text/html" href="https://julin.ai/2026/09/03/atlas-world-model/"/><published>2026-09-03T00:00:00+12:00</published><updated>2026-09-03T00:00:00+12:00</updated><category term="field-notes"/><category term="ai-engineering"/><category term="world-models"/><content type="html">&lt;blockquote&gt;
&lt;p&gt;Introducing Atlas: The world&amp;rsquo;s first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.&lt;/p&gt;
&lt;p&gt;Model the world, move the camera, and simulate space &amp;amp; time.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;a href="https://x.com/theworldlabs/status/2094839756329041984"&gt;World Labs&lt;/a&gt;, announcing Atlas on X.&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;World Labs, the startup Fei-Fei Li co-founded, &lt;a href="https://www.worldlabs.ai/blog/atlas"&gt;pretrained Atlas from scratch&lt;/a&gt; on text, images, video, and 3D together, not a video model with 3D bolted on after. It&amp;rsquo;s a multimodal autoregressive diffusion transformer: it generates the next frame by denoising it, one step at a time, conditioned on that 3D-positioned context.&lt;/p&gt;
&lt;p&gt;From 1 to 6 reference images, it generates up to a minute of video at 1440p with pixel-accurate camera control. Feed it more images, up to dozens, and it reconstructs the scene as an explicit 3D point cloud or Gaussian splat instead of just more video.&lt;/p&gt;
&lt;p&gt;World Labs ran human preference tests against camera-controlled generators: 75% of raters preferred Atlas over MiniMax H, 81% over Gemini Omni Flash, 93% over FLUX. On 3D reconstruction accuracy across seven datasets, Atlas averaged 25.3 mean absolute-relative pointmap error, ahead of the next-best specialist model, Pi3X, at 28.7.&lt;/p&gt;
&lt;p&gt;The target use case is Real-to-Sim: point a phone at a room, and Atlas turns that into a 3D space a robot can train in. Atlas is in early access with select partners now, and will power future versions of World Labs&amp;rsquo; existing product, Marble.&lt;/p&gt;</content></entry></feed>