今日从 arXiv 订阅中筛选 10 篇论文。

⚡ Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

⚡ Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs

⚡ H3-World: Turning Language Understanding into World Control

H3-World: Turning Language Understanding into World Control

⚡ ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

⚡ IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training

IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training

⚡ CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction

⚡ What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models

What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models

⚡ TempCloze: Can Video-LLMs Identify the Missing Middle?

TempCloze: Can Video-LLMs Identify the Missing Middle?

⚡ Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures

Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures

⚡ Towards Generalizable Visually Grounded Exploration of Household Devices

Towards Generalizable Visually Grounded Exploration of Household Devices

自动生成于 2026-09-03 · 基于 arXiv Daily Digest