LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering Paper • 2608.28281 • Published 4 days ago • 79 • 2
What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals Paper • 2608.19269 • Published 7 days ago • 3 • 3
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published 6 days ago • 185 • 4
PILOT in the Loop: Live Self-Improvement for Long-Horizon Agents Paper • 2608.26530 • Published 5 days ago • 30 • 3
TorchMorph: CUDA-accelerated Morphological Transforms Paper • 2608.24738 • Published 7 days ago • 4 • 3
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs Paper • 2608.20953 • Published 11 days ago • 11 • 4
Daedalus-150M: A Convolution-Attention Hybrid Designed for CPU Inference Paper • 2608.20210 • Published 12 days ago • 8 • 4
QuoteBench: How Matched Scores Can Hide Command-Path Failures Paper • 2608.13547 • Published 19 days ago • 8 • 4
The Embedder's Dilemma: LLMs Are Better, but at What Cost? Paper • 2608.12875 • Published 19 days ago • 15 • 5
SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science? Paper • 2608.19799 • Published 12 days ago • 65 • 4
SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents Paper • 2608.18852 • Published 13 days ago • 7 • 3
How Do Agents Fail on AutoResearch: End-to-End Diagnostic Evaluation on 100 Real-World Frontier Research Tasks Paper • 2608.14905 • Published 18 days ago • 30 • 4
Second Thought: Reasoning in Parallel as LLM Agents Act and Observe Paper • 2608.13667 • Published 18 days ago • 16 • 4
Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review Paper • 2608.12440 • Published 20 days ago • 10 • 5
Ready Cohorts: Bounding GPU Opportunity and Avoiding Host Round Trips in LLM-Agent Control Paper • 2608.12123 • Published 20 days ago • 2 • 3
Not Worth Another Token: Marginal Value Estimation for Efficient Deep Research Agents Paper • 2608.08389 • Published 23 days ago • 12 • 3
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace Paper • 2608.08621 • Published 23 days ago • 22 • 5
Omega-S: A Functional Resilience Index for LLM Fine-Tuning Paper • 2608.03887 • Published 28 days ago • 7 • 5
Stealing Reasoning Traces from Proprietary LLM APIs Paper • 2608.09867 • Published 22 days ago • 117 • 3
Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt Compression Paper • 2608.04569 • Published 27 days ago • 14 • 3