網誌
WM論文日誌
本文由簡體中文內容確定性轉換,並受版本化術語表保護。
1.使用説明
本文用於長期記錄 World Model(WM,世界模型)在機械人導航、路徑規劃、具身智能方向的論文閲讀、復現、實驗與創新點分析。
當前重點不是先“湊創新點”,而是建立 CV / 3D Vision → World Model → Simulator → Planner / Policy → Robot 的論文譜系,優先關注正式頂會和高水平期刊工作,並從中選擇代碼完整、算力可承受、能夠做閉環實驗的 baseline。
當前重點方向:Navigation World Model、Latent World Model、3D Structured World Model、World Model + Planning、World Model + Policy、Model-Based RL、動態障礙預測、實時化、Sim2Real、ROS2 / Nav2 實機部署。
2.先統一幾個概念
Sensor / Observation
传感器 / 当前观测
↓
Perception / Representation
感知 / 世界表示
↓
World Model
世界模型
↓
Future Prediction
未来预测
↓
Planner / Policy
规划器 / 策略
↓
Controller
控制器
↓
Robot Command
机器人执行命令
- Computer Vision(CV,計算機視覺):現在世界是什麼樣?
- World Model(WM,世界模型):如果執行動作 A,未來世界會怎樣?
- Planner(規劃器) / Policy(策略):應該選擇什麼動作?
- Controller(控制器):把動作真正變成速度、關節或力矩指令。
因此 World Model 的輸出不一定是 RGB,也通常不等於 cmd_vel。它可以輸出 Future RGB、Future Latent、Future Occupancy、Future 3DGS、Future Physical State、Reward / Risk / Value 等。
3.Renderer / Simulator / Planner 框架
3.1.Renderer(渲染器)
回答:未來世界“看起來”是什麼樣?
常見輸出:RGB、Video、Novel View、Depth、3D rendering。
3.2.Simulator(模擬器)
回答:世界在動作作用下實際上會怎樣變化?
常見輸出:Future State、Future Latent、Future Geometry、Future Occupancy、Future 3DGS、Object Pose、Reward / Risk。
3.3.Planner(規劃器)
回答:為了達到目標應該做什麼?
常見輸出:Action、Action Sequence、Trajectory、Velocity Command。
常見方法:CEM(Cross-Entropy Method,交叉熵方法)、MPC(Model Predictive Control,模型預測控制)、MPPI(Model Predictive Path Integral,模型預測路徑積分)、Learned Policy(學習策略)、Actor-Critic(演員-評論家)。
4.核心論文目錄
| 論文 | Venue / 狀態 | 核心路線 | 與移動導航關係 | 當前定位 | 優先級 |
|---|---|---|---|---|---|
| X-MOBILITY | ICRA 2025 | Latent WM + Policy | 直接相關 | 第一主 baseline 候選 | ⭐⭐⭐⭐⭐ |
| Navigation World Models(NWM) | CVPR 2025 Oral / Best Paper HM | Diffusion Video WM + Planning | 直接相關 | Navigation WM 標誌性工作 | ⭐⭐⭐⭐⭐ |
| DINO-WM | ICML 2025 | Pretrained Visual Latent WM + Planning | 高度相關 | Latent WM + Planning 核心參考 | ⭐⭐⭐⭐⭐ |
| World4Drive | ICCV 2025 | Physical Latent WM + Trajectory Selection | 自動駕駛,方法高度相關 | WM → Planning 橋樑重點參考 | ⭐⭐⭐⭐⭐ |
| WMNav | IROS 2025 | VLM + WM Memory + Navigation Policy | 直接相關 | Object Goal Navigation | ⭐⭐⭐⭐ |
| Drive-WM | CVPR 2024 | Multiview Video WM + Reward Planning | 自動駕駛,方法高度相關 | Visual WM → Planning 經典工作 | ⭐⭐⭐⭐ |
| GWM | ICCV 2025 | 3D Gaussian WM | 當前偏 Manipulation | 3D / Geometry-aware WM | ⭐⭐⭐⭐ |
| DreMa | ICLR 2025 | Gaussian Digital Twin + Physics | 當前偏 Manipulation | Compositional WM / Imagination | ⭐⭐⭐⭐ |
| MoSim | CVPR 2025 | Physical-State WM + Model-Based RL | 間接相關 | 非視覺 Dynamics WM | ⭐⭐⭐⭐ |
| MaskGWM | CVPR 2025 | Diffusion Video WM + Mask Reconstruction | 自動駕駛,偏 CV | 長時視頻預測 / 泛化 | ⭐⭐⭐ |
| V-JEPA 2 / 2-AC | 2025 Research Release / arXiv | Foundation Video WM + Action Conditioning | 當前偏 Manipulation | Foundation Latent WM | ⭐⭐⭐⭐ |
| One-Step WM | 2026 arXiv | One-Step Video WM + Planning | 直接相關 | Navigation WM 實時化 | ⭐⭐⭐⭐⭐ |
| AR Forcing | 2026 arXiv | Autoregressive Training for Navigation WM | 直接相關 | Long-horizon stability | ⭐⭐⭐⭐ |
| NavWAM | 2026 arXiv | World Model + Action Model | 直接相關 | WM → WAM 新路線 | ⭐⭐⭐⭐ |
| DreamerNav | Frontiers in Robotics and AI 2025 | DreamerV3 + MBRL | 直接相關 | 低成本系統參考,不作為高水平主 baseline 首選 | ⭐⭐⭐ |
| Atlas | World Labs 2026 Research Release | Generation + Reconstruction + Simulation | Spatial Intelligence | Foundation Spatial WM | ⭐⭐⭐⭐ |
5.論文方向佔比:每篇到底偏哪裏
下表是為了幫助快速建立直覺的研究重心估計,不是作者官方給出的數字。三個百分比相加為 100%。
| 論文 | CV / 3D Vision | World Modeling / Simulator | Planning / Policy / Robotics | 一句話定位 |
|---|---|---|---|---|
| Atlas | 55% | 40% | 5% | Foundation Spatial WM,核心是生成、重建、空間模擬 |
| MaskGWM | 65% | 30% | 5% | 最偏 CV,重點是駕駛視頻生成、長時預測和泛化 |
| GWM | 50% | 45% | 5% | 3D Gaussian 世界表示 + dynamics |
| DreMa | 35% | 45% | 20% | 3DGS Digital Twin + Physics + Robot Learning |
| Drive-WM | 45% | 35% | 20% | 多視角未來視頻 + image reward planning |
| NWM | 40% | 40% | 20% | Navigation Video WM + CEM planning |
| One-Step WM | 35% | 40% | 25% | 更快的 future imagination + optimization planning |
| AR Forcing | 30% | 55% | 15% | 解決 Navigation WM long-horizon rollout 漂移 |
| DINO-WM | 30% | 40% | 30% | 不生成 RGB,在 pretrained latent space 中預測並規劃 |
| V-JEPA 2-AC | 35% | 40% | 25% | Foundation representation + action-conditioned prediction |
| World4Drive | 25% | 35% | 40% | Latent WM 直接生成/評價候選軌跡 |
| WMNav | 20% | 30% | 50% | WM memory 服務 Object Goal Navigation |
| X-MOBILITY | 20% | 35% | 45% | Latent dynamics 服務 Learned Action Policy |
| DreamerNav | 10% | 35% | 55% | RSSM imagination + Actor-Critic + A* |
| MoSim | 5% | 55% | 40% | Future physical state + Model-Based RL |
| NavWAM | 20% | 30% | 50% | Future + Value + Action Chunk 聯合建模 |
6.Renderer / Simulator / Planner 功能強度
這一表與上面的“方向佔比”不同。一個系統可以同時是強 Renderer 和強 Simulator,因此三項不要求相加。
| 論文 | Renderer | Simulator | Planner / Policy | 主要輸出 |
|---|---|---|---|---|
| Atlas | ★★★★★ | ★★★★☆ | ★☆☆☆☆ | RGB / Depth / Point Cloud / 3DGS |
| MaskGWM | ★★★★★ | ★★★☆☆ | ★☆☆☆☆ | Future driving video |
| GWM | ★★★★☆ | ★★★★★ | ★★☆☆☆ | Future 3D Gaussian Scene |
| DreMa | ★★★★☆ | ★★★★★ | ★★☆☆☆ | Digital Twin imagined state / demonstrations |
| Drive-WM | ★★★★★ | ★★★★☆ | ★★★☆☆ | Future multiview video + trajectory choice |
| NWM | ★★★★★ | ★★★★☆ | ★★★★☆ | Future visual observation + CEM / ranking |
| One-Step WM | ★★★★☆ | ★★★★☆ | ★★★★☆ | One-step future prediction + optimized action |
| AR Forcing | ★★★★☆ | ★★★★★ | ★★★☆☆ | Stable long-horizon rollout |
| DINO-WM | ★☆☆☆☆ | ★★★★☆ | ★★★★☆ | Future pretrained visual latent |
| V-JEPA 2-AC | ★☆☆☆☆ | ★★★★☆ | ★★★★☆ | Future latent state |
| World4Drive | ★★☆☆☆ | ★★★★☆ | ★★★★★ | Future latent + selected trajectory |
| WMNav | ★☆☆☆☆ | ★★★☆☆ | ★★★★★ | WM memory / curiosity map → navigation action |
| X-MOBILITY | ★★☆☆☆ | ★★★★☆ | ★★★★★ | Probabilistic latent → velocity / path policy |
| DreamerNav | ★☆☆☆☆ | ★★★★☆ | ★★★★★ | RSSM imagination → navigation policy |
| MoSim | ☆☆☆☆☆ | ★★★★★ | ★★★★☆ | Future physical state → Model-Based RL |
| NavWAM | ★★★☆☆ | ★★★★☆ | ★★★★★ | Future observation + value + action chunk |
7.按 World Model 輸出形式分類
| 輸出 | 代表工作 | 特點 |
|---|---|---|
| Future RGB / Video | NWM、Drive-WM、MaskGWM、One-Step WM | 人可以直接觀察預測,但生成成本通常較高 |
| Future Latent State | DINO-WM、X-MOBILITY、V-JEPA 2-AC、World4Drive、DreamerNav | 不追求畫出未來,更適合 task-oriented prediction |
| Future 3D Gaussian Scene | GWM | 顯式三維幾何結構 |
| Digital Twin Future / Imagined Data | DreMa | 顯式場景 + Physics,主要服務數據生成和策略學習 |
| Future Physical State | MoSim | 直接預測機械人動力學狀態,幾乎不依賴 CV |
| Future + Action Chunk | NavWAM | 世界預測與動作生成開始融合 |
| Unified Spatial Output | Atlas | RGB、Depth、Point Cloud、3DGS 等統一空間內容 |
8.當前最值得關注的三條論文譜系
A. Visual Future → Planning
Drive-WM → NWM → One-Step WM / AR Forcing
B. Latent Future → Planning / Policy
DINO-WM → X-MOBILITY → World4Drive / NavWAM
C. Explicit / Physical World → Simulator → Decision
GWM / DreMa / MoSim
當前最值得重點追蹤的是 B:Latent Future → Planning / Policy。
原因:它正好位於 CV、World Model 與機械人規劃之間,不必把主要算力花在 photorealistic RGB generation(照片級 RGB 生成)上,同時能夠保留移動機械人 Navigation、Planning 和實機閉環優勢。
一個長期值得追的問題是:
對於移動機械人導航,什麼樣的 Future Representation(未來世界表示)最有利於 Planning?Future RGB、Future Latent、Future Occupancy、Future Geometry、Future Risk,誰真正能改善閉環導航?
9.論文鏈接總表
10.X-MOBILITY — ICRA 2025
10.1.一句話理解
RGB + Robot State
↓
World Model
↓
Probabilistic Latent State
↓
Action Policy
↓
Velocity + Optional Local Path
World Model 與 Action Policy 解耦。核心價值不是生成漂亮未來圖,而是讓 latent state 通過 world modeling 學到對導航有用的環境和動態信息。
原論文使用 Isaac Sim Nova Carter 數據,包含 Random Action Dataset 和 Nav2 Teacher Dataset,並提供 checkpoint、dataset、ONNX / TensorRT / ROS2 部署鏈。
當前定位:第一主 baseline 候選。
值得研究的問題:動態障礙、RGB-only geometry grounding、Navigation-oriented representation、輕量化、uncertainty-aware dynamics。
11.Navigation World Models(NWM)— CVPR 2025
11.1.一句話理解
Observation + Navigation Action
↓
Conditional Diffusion Transformer
↓
Future Visual Observation
↓
CEM / Trajectory Ranking
↓
Navigation
NWM 是 Visual World Model → Planning 的代表。它説明視頻生成型 WM 並不只是 CV:模型可以對不同候選 action 想像不同未來,再由 CEM 選擇更好的動作。
主要問題:推理成本、long-horizon drift、OOD mode collapse、pedestrian temporal dynamics。
當前定位:必須精讀,但大模型從零訓練成本過高,不優先作為第一篇完整重訓 baseline。
12.DINO-WM — ICML 2025
12.1.一句話理解
Image
↓
DINOv2 Features
↓
World Model
↓
Future Visual Latent
↓
CEM / Gradient Planning
核心問題:World Model 為什麼一定要重建未來 RGB?
DINO-WM 直接在 pretrained visual feature space(預訓練視覺特徵空間)預測未來,減少無關 pixel reconstruction,並讓 world representation 更 task-oriented。
當前定位:Latent World Model + Planning 必讀;非常適合啓發 Navigation-Oriented World Model。
13.World4Drive — ICCV 2025
13.1.一句話理解
Camera / Scene Features
↓
Vision Foundation Model
↓
Physical Latent World Model
↓
Generate Candidate Trajectories
↓
Predict Intention-aware Future Latents
↓
World Model Selector
↓
Best Trajectory
這篇最值得學習的是:World Model 直接參與 Trajectory Generation / Evaluation / Selection(軌跡生成 / 評價 / 選擇)。
雖然任務是自動駕駛,但方法思想非常容易遷移到移動機械人:
MPPI / Policy
产生候选轨迹
↓
World Model
预测每条轨迹对应的 Future State / Risk
↓
Trajectory Selection
當前定位:WM → Planning 橋樑重點參考。
14.WMNav — IROS 2025
14.1.一句話理解
WMNav 將 Vision-Language Model(視覺語言模型)、World Model Memory(世界模型記憶)與 Object Goal Navigation(目標物體導航)結合。
它維護 Curiosity Value Map(好奇價值地圖),並根據 world-model plan 與真實 observation 的反饋差異減輕 hallucination(幻覺)對導航決策的影響。
當前定位:明顯偏 Planner / Navigation。它説明 WM 不一定是視頻生成器,也可以是導航系統中的 memory / reasoning component。
15.Drive-WM — CVPR 2024
Current Multiview Images
+ Different Driving Maneuvers
↓
Drive-WM
↓
Multiple Future Videos
↓
Image-based Reward
↓
Choose Safer Trajectory
它是理解 Renderer / Simulator → Planner 鏈路的經典工作:先想像多個駕駛動作對應的未來視覺,再根據 reward 選擇更安全的軌跡。
當前定位:CV / Video WM 較重,但與 Planning 的連接非常明確。
16.GWM — ICCV 2025
GWM 把 World Representation 改成顯式 3D Gaussian(3D 高斯):
RGB Image(s)
↓
3D Gaussian Scene
↓
3D VAE
↓
Gaussian Latent + Action
↓
Latent Diffusion Transformer
↓
Future 3D Gaussian Scene
核心不是“換一種傳感器輸入”,而是讓世界表示具有顯式三維幾何結構。
當前定位:3D Vision + Simulator,適合學習 Geometry-aware World Model,不是移動導航主 baseline 首選。
17.DreMa — ICLR 2025
DreMa 使用 Gaussian Splatting + Physics Simulator 構建 Learnable Digital Twin(可學習數字孿生):
Few Real Demonstrations
↓
Learnable Digital Twin
↓
Imagination
↓
Synthetic Demonstrations
↓
Imitation Learning
↓
Robot Policy
當前定位:Simulator + Robot Learning。核心價值是理解 WM 不只用於在線規劃,也可以用於 imagination / data generation。
18.MoSim — CVPR 2025
MoSim 非常適合糾正“WM = 未來視頻生成”的誤解:
Current Physical State
+ Action
↓
MoSim
↓
Future Physical State
↓
Model-Based RL / Zero-shot RL
Physical State 包含關節位置、速度、空間位置與速度等。
當前定位:Simulator / Dynamics / Model-Based RL,幾乎不依賴 CV。
19.MaskGWM — CVPR 2025
MaskGWM 使用 Diffusion Transformer + Video Mask Reconstruction,重點解決 Driving World Model 的 Long-Horizon Prediction(長時預測)、Generalization(泛化)和 Multi-view Generation(多視角生成)。
當前定位:明顯偏 CV / Generative World Model。論文很強,但與當前移動機械人 Planning 主線距離比 World4Drive、DINO-WM、X-MOBILITY 更遠。
20.One-Step WM / AR Forcing / NavWAM:2026 新路線
20.1.One-Step World Model
重點解決多步 Diffusion + Autoregressive Generation 的高延遲問題,用 One-Step Generation 提升 Navigation WM 實時性,再接 optimization-based planning。
20.2.AR Forcing
重點解決訓練時使用 Ground Truth Context、推理時使用 Model-generated Context 導致的 Train-Test Distribution Shift 和 Long-Horizon Error Accumulation。
20.3.NavWAM
Observation + Goal
↓
World Action Model
↓
Future Observation
+ Goal Progress
+ Action Chunk
它把“預測未來”和“決定動作”聯合起來,體現 WM → WAM(World Action Model,世界動作模型) 的趨勢。
21.當前 Baseline 選擇判斷
21.1.第一梯隊:真正考慮拿來改
- X-MOBILITY — ICRA 2025:最貼移動機械人、ROS2、Sim2Real、Policy。
- DINO-WM — ICML 2025:最適合研究“Future Latent 是否比 Future RGB 更適合 Planning”。
- NWM — CVPR 2025:Navigation WM 標誌性工作,但完整訓練成本高。
- World4Drive — ICCV 2025:如果研究重點逐漸偏“WM 如何評價/選擇軌跡”,價值很高。
21.2.第二梯隊:強 Related Work / 方法參考
WMNav、Drive-WM、GWM、DreMa、MoSim、One-Step WM、AR Forcing、NavWAM。
21.3.不建議當前作為主 baseline
- MaskGWM:太容易滑向純視頻生成;
- Atlas:沒有公開完整訓練鏈,規模也不適合作為碩士第一篇 baseline;
- DreamerNav:工程和學習價值高,但如果目標明確要求高水平 venue,不再優先作為主 baseline。
22.一個值得長期追的 Research Question
Navigation World Model 真的有必要預測 photorealistic Future RGB 嗎?
可以逐漸形成這樣的實驗譜系:
Future RGB
vs
Future Latent
vs
Future Occupancy / BEV
vs
Future Geometry / 3DGS
vs
Future Risk
↓
同一个 Planning / Policy Framework
↓
比较:
Success Rate
Collision Rate
Generalization
Latency
VRAM
Real-Robot Performance
這比單純“給某個模型加一個模塊”更接近一個可以持續做下去的研究問題。
23.統一論文閲讀模板
以後每篇新論文優先回答以下問題。
23.1.基本信息
- Title:
- Venue / Year:
- Organization:
- Paper / Project / GitHub:
- 是否正式發表:
- 是否有 Code / Checkpoint / Dataset:
23.2.Observation / Input(觀測 / 輸入)
模型看什麼:RGB、Depth、LiDAR、Robot State、Camera Pose、Language Goal?
23.3.World Representation(世界表示)
模型腦子裏怎麼表示世界:Pixel / Video Latent、DINO Feature、RSSM Latent、BEV、Occupancy、3D Gaussian、Physical State、Digital Twin?
23.4.World Model Output(世界模型輸出)
Future RGB、Future Latent、Future Occupancy、Future 3DGS、Future Physical State、Reward / Risk / Value、Action Chunk?
23.5.Renderer / Simulator / Planner 定位
Renderer:
Simulator:
Planner / Policy:
23.6.CV ↔ Robotics 方向佔比
CV / 3D Vision:
World Modeling / Simulator:
Planning / Policy / Robotics:
23.7.Planner / Policy
CEM、MPC、MPPI、Gradient Planning、Learned Policy、Actor-Critic,還是 World Action Model?
23.8.Robot Output
移動機械人:vx / vy / wz、trajectory、local path?
機械臂:joint position、joint velocity、end-effector pose、action chunk?
23.9.Training / Compute
記錄 GPU、數量、顯存、訓練時間、Epoch / Steps、Precision、Multi-GPU Strategy。
23.10.Evaluation
記錄 Baselines、Main Results、Ablation、Efficiency、Generalization、Failure Cases、Real Robot、Sim2Real。
23.11.是否適合當我的 baseline
代码完整度:
Checkpoint:
Dataset:
算力可承受:
移动机器人相关性:
实机难度:
创新空间:
24.當前階段 Checklist
- 精讀 X-MOBILITY
- 精讀 NWM
- 精讀 DINO-WM
- 精讀 World4Drive
- 閲讀 WMNav
- 閲讀 Drive-WM
- 精讀 GWM
- 閲讀 DreMa
- 閲讀 MoSim
- 閲讀 MaskGWM
- 閲讀 One-Step WM
- 閲讀 AR Forcing
- 閲讀 NavWAM
- 理解 Renderer / Simulator / Planner
- 理解 RGB / Latent / Occupancy / 3DGS / Physical State 等不同 WM 輸出
- 跑通至少一個正式頂會 baseline 的 official checkpoint
- 建立 Failure Case Dataset
- 從穩定 failure 提出 Research Question
- 做最小修改驗證 hypothesis
- 再進入正式創新點設計