PhiZero: A World Model Built Around Physical Language Paper • 2607.28624 • Published 4 days ago • 158
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 5 days ago • 127
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Paper • 2607.13429 • Published 19 days ago • 16
GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction Paper • 2605.23888 • Published May 22 • 17
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published 21 days ago • 84
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 24 days ago • 86
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space Paper • 2607.05373 • Published 28 days ago • 66
Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling Paper • 2607.01642 • Published Jul 2 • 39
Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views Paper • 2606.29513 • Published Jun 28 • 53
TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction Paper • 2605.26115 • Published May 25 • 52
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Paper • 2605.30280 • Published May 28 • 146
Geometry-Aware Representation Denoising for Robust Multi-view 3D Reconstruction Paper • 2605.26230 • Published May 25 • 41
SpatialBench: Is Your Spatial Foundation Model an All-Round Player? Paper • 2605.27367 • Published May 26 • 72
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Paper • 2605.27365 • Published May 26 • 146
CubePart: An Open-Vocabulary Part-Controllable 3D Generator Paper • 2605.28763 • Published May 27 • 14