今日从 arXiv 订阅中筛选 10 篇论文。

⚡ GeoWAM: Visual Geometry World Action Models for Autonomous Driving

GeoWAM: Visual Geometry World Action Models for Autonomous Driving

⚡ ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding

ViSMoE: Visual-Aware Sparse Mixture-of-Experts for Embodied Referring Expression Grounding

⚡ Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

⚡ Correcting a learned physical invariant improves world-model rollouts

⚡ EchoWM: Open and Enterable Omnimodal World Models

EchoWM: Open and Enterable Omnimodal World Models

⚡ ReWorld: An Interactive World Model with Long-Horizon Memory

ReWorld: An Interactive World Model with Long-Horizon Memory

⚡ MOSH-WM: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models

MOSH-WM: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models

⚡ WAM-OPD: On-Policy Distillation for World Action Models

WAM-OPD: On-Policy Distillation for World Action Models

⚡ Perturb the Thought, Not the Pixels: Latent-Space Rollout Diversification for RLVR

Perturb the Thought, Not the Pixels: Latent-Space Rollout Diversification for RLVR

⚡ On the Capability Separation Between World-Model Policy Learning and Imitated World-Action Models


自动生成于 2026-08-25 · 基于 arXiv Daily Digest