今日从 arXiv 订阅中筛选 10 篇论文。

⚡ The Latent That Never Was: A Forensic Re-run of the CVAE Ablation in Action Chunking Transformer

把 ACT 最著名的消融(删 CVAE encoder:35%→2%)拿到原代码里重跑。

⚡ Right Direction, Wrong Step: Geometric Analysis of Finite-Step Failure in Looped Transformers

looped Transformer 加循环为什么有时变差:不是方向错了,是步长错了。

Right Direction, Wrong Step: Geometric Analysis of Finite-Step Failure in Looped Transformers

⚡ Modality-Autoregressive World-Action Models

WAM 应该预测哪些模态:点轨迹、DINO 特征、深度 > RGB。

⚡ CorrRisk-WM: Corridor-Conditioned Risk World Modeling for Safety-Critical Trajectory Planning

风险评估要条件化在候选轨迹上:同样他车运动,对每条自车候选轨迹风险不同。

CorrRisk-WM: Corridor-Conditioned Risk World Modeling for Safety-Critical Trajectory Planning

⚡ PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control

生成中途可交互的视频物理控制:控制信号用物理量而不是像素位置。

PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control

⚡ LoopSpec: Pipelined Self-Speculative Decoding for Looped Transformers

looped Transformer 的免训练自推测解码:中间循环状态天然是 draft。

⚡ Not Another Text Benchmark: Putting the “Visual” Back in Visual Question Answering for Large Video Models

换评测模态:当 query 是视觉的而不是文本选项时,视频模型会露馅。

Not Another Text Benchmark: Putting the

⚡ Dense to MoE Adaptation for Compact Vision Language Action Policies

稠密 FFN 转 MoE 来压缩 VLA 策略:保函数初始化 + 动态专家掩码。

Dense to MoE Adaptation for Compact Vision Language Action Policies

⚡ XPACE: Joint World and Action Modeling from Heterogeneous Experience

同一个视频骨干,既是 world-action model 也是 world simulator。

XPACE: Joint World and Action Modeling from Heterogeneous Experience

⚡ Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It

"循环模型不擅长召回"被拆成受控分解:是缺的卷积,不是循环本身。


自动生成于 2026-09-17 · 基于 arXiv Daily Digest