N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Paper • 2607.23782 • Published Jul 26 • 78
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published Jul 16 • 75
Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Paper • 2607.11643 • Published Jul 13 • 47
Humanoid-GPT: Scaling Data and Structure for Zero-Shot Motion Tracking Paper • 2606.03985 • Published Jun 2 • 43
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents Paper • 2605.12481 • Published May 12 • 28
Repurposing 3D Generative Model for Autoregressive Layout Generation Paper • 2604.16299 • Published Apr 17 • 12
Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development Paper • 2603.27460 • Published Mar 29 • 72
EgoActor: Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models Paper • 2602.04515 • Published Feb 4 • 39
Future Optical Flow Prediction Improves Robot Control & Video Generation Paper • 2601.10781 • Published Jan 15 • 19