Join the conversation

Join the community of Machine Learners and AI enthusiasts.

Sign Up
SeaWolf-AIΒ 
posted an update 2 days ago
Post
5904
🧬 Darwin-180B-RSI β€” an AI that learns from itself and knows when it's right
πŸ‘‰ FINAL-Bench/Darwin-180B-RSI

🧬 Darwin β€” crossbreed and evolve the parent
Darwin diagnoses strong parent models like an MRI, inherits only their best parts, and evolves the weak spots β€” producing a child stronger than its parents.
Father model: Qwen3.8-Flash-Next (180B MoE).

πŸ”§ Rewired paths
πŸ”Ή 12 full-attention layers Β· πŸ”Ή 36 linear-attention layers Β· πŸ”Ή 48 shared-expert layers β€” precision-strengthened
πŸ”’ 512 routed experts Β· router Β· vision encoder β€” untouched
β†’ Only 0.02% of the weights changed.

πŸ” RSI Γ— πŸ›οΈ ZTC
RSI (recursive self-improvement): solve β†’ verify against real answers β†’ learn only the correct reasoning β†’ repeat.
ZTC (Zero-Token Confidence): reads the model's internal state once, before answering, and returns the probability the answer is right β€” zero extra tokens. Returns answer + confidence as JSON.
{"answer": "...", "confidence": 0.97, "truncated": false}

✨ Synergy: ZTC finds where the model wavers β†’ RSI learns exactly there β†’ confidence gets sharper. Low confidence = stop, so agents don't act on wrong answers.
⚑ Same accuracy, 11% shorter reasoning β€” faster and cheaper.

πŸ“„ https://arxiv.org/abs/2605.14386
πŸ€— FINAL-Bench/Darwin-180B-RSI
πŸ›οΈ https://huggingface.co/collections/FINAL-Bench/ztc-models-jev-ecosystems

πŸ† The result β€” #1 on five Hugging Face official leaderboards
πŸ₯‡ AIME 2026 100% (first perfect score on the board)
πŸ₯‡ HMMT Feb 2026 100% (first perfect score on the board)
πŸ₯‡ GPQA Diamond 94.44%
πŸ₯‡ MMLU-Pro 88.12%
πŸ₯‡ MMMU-Pro 79.48%

πŸ“ 131K-token thinking budget Β· bf16 Β· samples per benchmark listed on the model card. πŸš€

#Darwin #RSI #ZTC #AIME #HMMT #GPQA #MMLUPro #MMMUPro #OpenSource
In this post