今日从 arXiv 订阅中筛选 10 篇论文。

⚡ World in World: Explore the World with World Models

训练free的"控制证据接口":把多源视觉证据统一成冻结视频世界模型能直接读的干净状态,相机重渲染/重访/动作迁移共用同一骨干。

World in World: Explore the World with World Models

⚡ Thinking with Looped Flows

递归计算的新训练法:局部去噪目标 + 概率流积分推理;ARC-AGI-1 58.8 为 looped 家族新高。

Thinking with Looped Flows

⚡ MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in BEV

统一检测-预测架构的运动一致性补丁:train-only 损失、推理零开销、代码开源。

MC-DeTra: Motion-Consistent Joint Object Detection and Socially-Aware Trajectory Forecasting in BEV

⚡ Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs

代码世界补上"构造法":全局-局部-全局递归,单图→可执行 3D 世界,VLM coding agent 驱动细化。

Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs

⚡ Beyond Visual Quality: Evaluating Physical Consistency under Ego-Motion with EgoGenEval

ego-motion 物理一致性双轴尺:无敌两轴兼得;pairwise 监督被点名是瓶颈。

Beyond Visual Quality: Evaluating Physical Consistency under Ego-Motion with EgoGenEval

⚡ Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding

"视觉-文本二元性"落地成路由:语言背时序、像素只查属性;视觉访问=query 级成本。

Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding

⚡ From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models

视频生成"认知正确性"的分层诊断:渲染强、逻辑弱;prompt 外移认知负担有增益。

From Evaluation to Enhancement: Benchmarking and Improving Think-with-Video Reasoning for Video Generative Models

⚡ New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models

证据选择是独立盲区:95-100% 的配对重复同一动作,最优切换全对仅 5.9%。

New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models

⚡ TrajFusionNet+: Transformer-Based Prediction of Pedestrian Crossing Intention via Fusion of Trajectory Representations and Scene Graphs

行人过街意图三分支融合 + 跨数据集泛化新协议;v2 工程增量。

TrajFusionNet+: Transformer-Based Prediction of Pedestrian Crossing Intention via Fusion of Trajectory Representations and Scene Graphs

⚡ MindTopo: Can Foundation Models Reason in Topological Space?

拓扑空间推理量表:推理>规划 gap 全员存在;生成式观测不保拓扑。

MindTopo: Can Foundation Models Reason in Topological Space?

自动生成于 2026-09-12 · 基于 arXiv Daily Digest