今日从 arXiv 订阅中筛选 10 篇论文。
⚡ Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation

⚡ Video-OPSD: Exploiting Privileged Visual Evidence for On-Policy Self-Distillation in Video LLMs

⚡ Reason in the Words You Speak: Idiolectal Paraphrasing Off-Policy Traces for Reasoning Distillation in VideoLLMs

⚡ SpatialCrafter: Single Image World Modeling with Generative 3D Proxies

⚡ RECAP-Forcing: Retaining Content Appearances for Long Video Generation

⚡ Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion

⚡ LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics

⚡ Thinking on Shots: Consistent Multi-Shot Video Editing with Agentic Reasoning

⚡ Retrieval Heads Meet Vision: Uncovering How VLMs Locate and Extract Visual Information

自动生成于 2026-08-29 · 基于 arXiv Daily Digest
