FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published 6 days ago • 140
Apodex 1.1: Scaling Agentic Intelligence for Complex Work Paper • 2608.23283 • Published 7 days ago • 201
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Paper • 2608.11341 • Published 20 days ago • 64
SFT Doesn't Always Hurt General Capabilities: Revisiting Domain-Specific Fine-Tuning in LLMs Paper • 2509.20758 • Published Sep 25, 2025 • 2
Good SFT Optimizes for SFT, Better SFT Prepares for Reinforcement Learning Paper • 2602.01058 • Published Feb 1 • 45