Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published 6 days ago • 59
Boogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation Paper • 2607.13125 • Published 8 days ago • 133
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 12 days ago • 81
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published 14 days ago • 64
ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Paper • 2606.19531 • Published Jun 17 • 24
DreamX-World 1.0: A General-Purpose Interactive World Model Paper • 2606.16993 • Published Jun 15 • 114
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper • 2606.14777 • Published Jun 10 • 214
InterleaveThinker: Reinforcing Agentic Interleaved Generation Paper • 2606.13679 • Published Jun 11 • 83
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models Paper • 2605.21573 • Published May 20 • 111
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining Paper • 2605.14747 • Published May 14 • 147