今日从 arXiv 订阅中筛选 8 篇论文。
⚡ WorldDiT: A Unified Diffusion Architecture for World and Action Modeling

⚡ MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving

⚡ VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation

⚡ ObsDriveBench: Benchmarking Multimodal Understanding under Adverse Weather

⚡ InterOCF: Spatio-Temporal 2D-3D Interaction for Camera-Only 4D Occupancy Forecasting
⚡ Reasoning to Regulate: Chain-of-Thought for Traffic Rule Understanding

⚡ CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding

⚡ Mixture-of-Thought-Tokens: Unifying Perception and Reasoning for Free-form Multimodal Grounding

自动生成于 2026-07-28 · 基于 arXiv Daily Digest