-
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Paper • 2606.10917 • Published • 76 -
DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratch
Paper • 2606.10728 • Published • 36 -
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
Paper • 2606.32032 • Published • 29
Collections
Discover the best community collections!
Collections including paper arxiv:2606.10917
-
Multi-Agent Computer Use
Paper • 2606.01533 • Published • 7 -
OpenSkill: Open-World Self-Evolution for LLM Agents
Paper • 2606.06741 • Published • 29 -
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Paper • 2606.07412 • Published • 12 -
Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses
Paper • 2606.08348 • Published • 16
-
EvoMaster: A Foundational Agent Framework for Building Evolving Autonomous Scientific Agents at Scale
Paper • 2604.17406 • Published • 7 -
Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution
Paper • 2605.15301 • Published • 23 -
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
Paper • 2605.11739 • Published • 61 -
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
Paper • 2605.18401 • Published • 132
-
Large Language Models Orchestrating Structured Reasoning Achieve Kaggle Grandmaster Level
Paper • 2411.03562 • Published • 70 -
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
Paper • 2502.06060 • Published • 37 -
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Paper • 2502.14499 • Published • 196 -
SurveyX: Academic Survey Automation via Large Language Models
Paper • 2502.14776 • Published • 100
-
Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering
Paper • 2605.29648 • Published • 10 -
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
Paper • 2605.29548 • Published • 13 -
Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation
Paper • 2605.29861 • Published • 16 -
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
Paper • 2605.31264 • Published • 131
-
SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models
Paper • 2506.04180 • Published • 35 -
AniMaker: Automated Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
Paper • 2506.10540 • Published • 37 -
AutoMind: Adaptive Knowledgeable Agent for Automated Data Science
Paper • 2506.10974 • Published • 19 -
SPAR: Scholar Paper Retrieval with LLM-based Agents for Enhanced Academic Search
Paper • 2507.15245 • Published • 11
-
Role-Agent: Bootstrapping LLM Agents via Dual-Role Evolution
Paper • 2606.10917 • Published • 76 -
DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratch
Paper • 2606.10728 • Published • 36 -
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
Paper • 2606.32032 • Published • 29
-
Multi-Agent Computer Use
Paper • 2606.01533 • Published • 7 -
OpenSkill: Open-World Self-Evolution for LLM Agents
Paper • 2606.06741 • Published • 29 -
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Paper • 2606.07412 • Published • 12 -
Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses
Paper • 2606.08348 • Published • 16
-
Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering
Paper • 2605.29648 • Published • 10 -
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
Paper • 2605.29548 • Published • 13 -
Towards Verifiable Multimodal Deep Research: A Multi-Agent Harness for Interleaved Report Generation
Paper • 2605.29861 • Published • 16 -
COLLEAGUE.SKILL: Automated AI Skill Generation via Expert Knowledge Distillation
Paper • 2605.31264 • Published • 131
-
EvoMaster: A Foundational Agent Framework for Building Evolving Autonomous Scientific Agents at Scale
Paper • 2604.17406 • Published • 7 -
Solvita: Enhancing Large Language Models for Competitive Programming via Agentic Evolution
Paper • 2605.15301 • Published • 23 -
Learning to Foresee: Unveiling the Unlocking Efficiency of On-Policy Distillation
Paper • 2605.11739 • Published • 61 -
SkillsVote: Lifecycle Governance of Agent Skills from Collection, Recommendation to Evolution
Paper • 2605.18401 • Published • 132
-
SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models
Paper • 2506.04180 • Published • 35 -
AniMaker: Automated Multi-Agent Animated Storytelling with MCTS-Driven Clip Generation
Paper • 2506.10540 • Published • 37 -
AutoMind: Adaptive Knowledgeable Agent for Automated Data Science
Paper • 2506.10974 • Published • 19 -
SPAR: Scholar Paper Retrieval with LLM-based Agents for Enhanced Academic Search
Paper • 2507.15245 • Published • 11
-
Large Language Models Orchestrating Structured Reasoning Achieve Kaggle Grandmaster Level
Paper • 2411.03562 • Published • 70 -
Training Language Models for Social Deduction with Multi-Agent Reinforcement Learning
Paper • 2502.06060 • Published • 37 -
MLGym: A New Framework and Benchmark for Advancing AI Research Agents
Paper • 2502.14499 • Published • 196 -
SurveyX: Academic Survey Automation via Large Language Models
Paper • 2502.14776 • Published • 100