Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents Paper • 2608.30322 • Published 4 days ago • 1 • 2
From Production Traffic to Post-Training: Building a Self-Hosted LLM That Covers the Corporate Request Mix Paper • 2609.01572 • Published 3 days ago • 33 • 3
Super Library Agent: Joint Generation and Maintenance of Multiple Applications Beyond the Single Codebase Paper • 2608.29310 • Published 6 days ago • 27 • 4
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Paper • 2608.28281 • Published 7 days ago • 100 • 5
What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals Paper • 2608.19269 • Published 10 days ago • 5 • 3
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published 9 days ago • 195 • 5
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents Paper • 2608.26530 • Published 8 days ago • 32 • 4
TorchMorph: CUDA-accelerated Morphological Transforms Paper • 2608.24738 • Published 10 days ago • 4 • 3
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs Paper • 2608.20953 • Published 14 days ago • 11 • 4
Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference Paper • 2608.20210 • Published 15 days ago • 8 • 4
QuoteBench: How Matched Scores Can Hide Command-Path Failures Paper • 2608.13547 • Published 22 days ago • 8 • 4
The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published 22 days ago • 15 • 5
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published 15 days ago • 65 • 4
SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents Paper • 2608.18852 • Published 16 days ago • 7 • 3
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks Paper • 2608.14905 • Published 21 days ago • 30 • 4
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published 21 days ago • 16 • 4
Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review Paper • 2608.12440 • Published 23 days ago • 10 • 5
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control Paper • 2608.12123 • Published 23 days ago • 2 • 3
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents Paper • 2608.08389 • Published 26 days ago • 12 • 3
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace Paper • 2608.08621 • Published 26 days ago • 22 • 5