ContextPilot: Teaching Agents for Proactive Context Management via Fine-grained RL Paper • 2608.28476 • Published 21 days ago • 28
What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents Paper • 2608.27260 • Published 22 days ago • 73
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published Aug 17 • 151
ConceptFormer: Learning Adaptive Latent Concepts for Query-Document Alignment in Visual Document Retrieval Paper • 2608.15698 • Published Aug 16 • 6
See2Think: Do Multimodal Models Really Use Intermediate Visual States? Paper • 2607.26769 • Published Jul 29 • 25
DataEvolver: Self-Evolving Multi-Agent Data Construction for Text-Rich Image Generation Paper • 2606.31537 • Published Jun 30 • 30
From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI Paper • 2606.14502 • Published Jun 12 • 118
Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding Paper • 2603.18472 • Published Mar 19 • 20
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention Paper • 2602.05847 • Published Feb 5 • 12
VidEgoThink: Assessing Egocentric Video Understanding Capabilities for Embodied AI Paper • 2410.11623 • Published Oct 15, 2024 • 49
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM Paper • 2501.00599 • Published Dec 31, 2024 • 45