今日从 arXiv 订阅中筛选 10 篇论文。
⚡ Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

⚡ Latent Recurrent Thoughts: Recurrent Refinement of Proposed Latents for Reasoning with Frozen LLMs
⚡ ZimaBlue: Evolving Generalizable World Action Models through Scalable Video Pre-training

⚡ IMPACT: Attention Is the Interaction Map for Scalable Interaction-Aware World Model Training

⚡ CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction
⚡ What, Where, and How: Probing Spatiotemporal Representations in Video Foundation Models

⚡ Separating perception from reasoning in vision-language models: a model-free render ceiling for crystal structures

⚡ Towards Generalizable Visually Grounded Exploration of Household Devices

自动生成于 2026-09-03 · 基于 arXiv Daily Digest

