βοΈ The Open Quantum Challenge is live. Both seasons open today.
Quantum computers are expensive, queued, and reachable by only a few. So we removed the barrier. With a single GPU and a public harness, anyone can contribute to core quantum-computing problems with no quantum hardware at all. That is the point of this challenge: to move quantum research from a handful of labs to everyone.
β’ Season 1, Quantum simulation: reproduce a target quantum system on a GPU. β’ Season 2, QEC decoder: correct errors on coded quantum states.
Answers are held privately and every submission is auto-scored against a frozen ground truth, so the ranking is reproducible and hardware-independent.
π Prize: 2,000 USD total (1,000 USD per season, π₯600 / π₯300 / π₯100)
Getting started takes one step: copy the participation guide from the Space and paste it into Codex or Claude Code, then submit against the harness.
𧬠Darwin-180B-RSI β an AI that learns from itself and knows when it's right π FINAL-Bench/Darwin-180B-RSI
𧬠Darwin β crossbreed and evolve the parent Darwin diagnoses strong parent models like an MRI, inherits only their best parts, and evolves the weak spots β producing a child stronger than its parents. Father model: Qwen3.8-Flash-Next (180B MoE).
π RSI Γ ποΈ ZTC RSI (recursive self-improvement): solve β verify against real answers β learn only the correct reasoning β repeat. ZTC (Zero-Token Confidence): reads the model's internal state once, before answering, and returns the probability the answer is right β zero extra tokens. Returns answer + confidence as JSON. {"answer": "...", "confidence": 0.97, "truncated": false}
β¨ Synergy: ZTC finds where the model wavers β RSI learns exactly there β confidence gets sharper. Low confidence = stop, so agents don't act on wrong answers. β‘ Same accuracy, 11% shorter reasoning β faster and cheaper.
π The result β #1 on five Hugging Face official leaderboards π₯ AIME 2026 100% (first perfect score on the board) π₯ HMMT Feb 2026 100% (first perfect score on the board) π₯ GPQA Diamond 94.44% π₯ MMLU-Pro 88.12% π₯ MMMU-Pro 79.48%
π 131K-token thinking budget Β· bf16 Β· samples per benchmark listed on the model card. π
Run it on defaults and it takes 244 s. Switch to 3 steps and it's 48.6 s. Add VAE tiling and it's 46.4 s.
The biggest culprit was the default. Z-Image Turbo is distilled to paint in few strokes, but the tool's default is 20. We were throwing away 5Γ for no reason. So were we, at first.
3 is the floor. Put 4 and 3 side by side and you cannot tell them apart. At 2 it collapses β water droplets and wood grain vanish, and the surface turns cloth-like.
The cost of a judging gate is usually quoted as a number. This puts it on a Tetris board.
Three boards get the same piece order, and on every move the same proposal and the same noise β a paired comparison. The gate decides one thing: keep this move, or draw again. Each board gets the same 60 seconds of gate time.
The text-writing gates get through 15β22 moves. The generation-free gate gets through 40β50. The boards that stop simply run out of clock.
It does not win on accuracy: on the same 2,018-question LODO set, JEV scores AUC 0.7350 against ZTC-Judge-27B's 0.7289. The separation is elsewhere. Clock β 2.1 s vs 0.0615 s per call, and on a 200-candidate agent screen one judging call measured 3.206 s generative vs 0.033 s readout, same server. Calibration β a gate is a threshold, and at ECE 0.4985 (vs ZTC 0.0245) a threshold stops carrying information. Mechanism β a text judge can name option 42 when there is no option 42; a scoring readout cannot. Not a lower error rate. No path.
The curve in the ZTC panel is real online fitting, scored prequentially β predict first, learn after β with base weights untouched. Not recursive self-improvement.
Limits, also stated on the page: Laya's AUC and latency are not our measurements and are set equal to JEV's, so calibration is the only measured axis it differs on. The page is a simulation driven by measured constants.