Collections
Discover the best community collections!
Collections including paper arxiv:2608.05987
-
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Paper • 2607.29209 • Published • 34 -
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper • 2608.05987 • Published • 94 -
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Paper • 2608.05131 • Published • 12 -
On-Policy Self-Distillation without Any Supervision
Paper • 2608.06296 • Published • 196
-
Multi-Agent Computer Use
Paper • 2606.01533 • Published • 7 -
OpenSkill: Open-World Self-Evolution for LLM Agents
Paper • 2606.06741 • Published • 29 -
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Paper • 2606.07412 • Published • 12 -
Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses
Paper • 2606.08348 • Published • 16
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 103 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75
-
Continuous Latent Diffusion Language Model
Paper • 2605.06548 • Published • 87 -
Scaling Latent Reasoning via Looped Language Models
Paper • 2510.25741 • Published • 234 -
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Paper • 2502.05171 • Published • 161 -
Pretraining Language Models to Ponder in Continuous Space
Paper • 2505.20674 • Published • 3
-
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 170 -
Mental World Modeling
Paper • 2607.27201 • Published • 104 -
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper • 2608.05987 • Published • 94 -
Progressive Agent Skill Generation via Reinforcement Learning
Paper • 2608.01678 • Published • 58
-
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
Paper • 2606.30406 • Published • 23 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 19 -
Trust Region Policy Distillation
Paper • 2607.04751 • Published • 35 -
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Paper • 2607.14777 • Published • 106
-
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Paper • 2504.06261 • Published • 110 -
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
Paper • 2510.11696 • Published • 183 -
AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference
Paper • 2502.04077 • Published • 2 -
An Embarrassingly Simple Approach for Wafer Feature Extraction and Defect Pattern Recognition
Paper • 2303.11632 • Published • 1
-
Continuous Latent Diffusion Language Model
Paper • 2605.06548 • Published • 87 -
Scaling Latent Reasoning via Looped Language Models
Paper • 2510.25741 • Published • 234 -
Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach
Paper • 2502.05171 • Published • 161 -
Pretraining Language Models to Ponder in Continuous Space
Paper • 2505.20674 • Published • 3
-
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks
Paper • 2608.01964 • Published • 170 -
Mental World Modeling
Paper • 2607.27201 • Published • 104 -
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper • 2608.05987 • Published • 94 -
Progressive Agent Skill Generation via Reinforcement Learning
Paper • 2608.01678 • Published • 58
-
SAF-OPD: Stable Advantage Fusion for On-Policy Distillation
Paper • 2607.29209 • Published • 34 -
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning
Paper • 2608.05987 • Published • 94 -
OPD-V: Visual On-Policy Self-Distillation with Modality Balance
Paper • 2608.05131 • Published • 12 -
On-Policy Self-Distillation without Any Supervision
Paper • 2608.06296 • Published • 196
-
MOPD: Multi-Teacher On-Policy Distillation for Capability Integration in LLM Post-Training
Paper • 2606.30406 • Published • 23 -
TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training
Paper • 2607.05804 • Published • 19 -
Trust Region Policy Distillation
Paper • 2607.04751 • Published • 35 -
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
Paper • 2607.14777 • Published • 106
-
Multi-Agent Computer Use
Paper • 2606.01533 • Published • 7 -
OpenSkill: Open-World Self-Evolution for LLM Agents
Paper • 2606.06741 • Published • 29 -
Socratic-SWE: Self-Evolving Coding Agents via Trace-Derived Agent Skills
Paper • 2606.07412 • Published • 12 -
Bayesian-Agent: Posterior-Guided Skill Evolution for LLM Agent Harnesses
Paper • 2606.08348 • Published • 16
-
Hogwild! Inference: Parallel LLM Generation via Concurrent Attention
Paper • 2504.06261 • Published • 110 -
QeRL: Beyond Efficiency -- Quantization-enhanced Reinforcement Learning for LLMs
Paper • 2510.11696 • Published • 183 -
AttentionPredictor: Temporal Pattern Matters for Efficient LLM Inference
Paper • 2502.04077 • Published • 2 -
An Embarrassingly Simple Approach for Wafer Feature Extraction and Defect Pattern Recognition
Paper • 2303.11632 • Published • 1
-
Diffusion Augmented Agents: A Framework for Efficient Exploration and Transfer Learning
Paper • 2407.20798 • Published • 24 -
Offline Reinforcement Learning for LLM Multi-Step Reasoning
Paper • 2412.16145 • Published • 38 -
REINFORCE++: A Simple and Efficient Approach for Aligning Large Language Models
Paper • 2501.03262 • Published • 103 -
SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Paper • 2502.18449 • Published • 75