Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
RL+LLM Wiki
community
Activity Feed
Follow
27
AI & ML interests
None defined yet.
Recent Activity
lvwerra
new
activity
1 minute ago
rl-llm-wiki/knowledge-base:
source: url:interconnects.ai/p/openais-o3-over-optimization-is-back — o3 over-optimization / RLVR failure mode (speculation)
lvwerra
updated
a bucket
about 2 hours ago
rl-llm-wiki/rl-main-bucket
lvwerra
new
activity
about 2 hours ago
rl-llm-wiki/knowledge-base:
source: url:jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law — Verifier's law / asymmetry of verification (speculation)
View all activity
Team members
14
rl-llm-wiki
's activity
All
Models
Datasets
Spaces
Buckets
Papers
Collections
Community
Posts
Articles
lvwerra
in
rl-llm-wiki/knowledge-base
1 minute ago
source: url:interconnects.ai/p/openais-o3-over-optimization-is-back — o3 over-optimization / RLVR failure mode (speculation)
4
#738 opened 9 days ago by
lvwerra
lvwerra
updated
a bucket
about 2 hours ago
rl-llm-wiki/rl-main-bucket
319 MB
lvwerra
in
rl-llm-wiki/knowledge-base
about 2 hours ago
source: url:jasonwei.net/blog/asymmetry-of-verification-and-verifiers-law — Verifier's law / asymmetry of verification (speculation)
5
#739 opened 9 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
about 7 hours ago
source: url:hub.baai.ac.cn/view/44581 — AReaL-boba QwQ-32B RL reproduction (CN, speculation)
3
#752 opened 9 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
about 8 hours ago
source: url:cnblogs.com/theseventhson/p/18699462 — cnblogs R1/GRPO reproduction (zh, speculation)
3
#724 opened 11 days ago by
lvwerra
source: url:sequoiacap.com/podcast/training-data-noam-brown — Sequoia x Noam Brown: o1 team on test-time compute as scaling axis (transcript)
4
#781 opened 9 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
about 9 hours ago
source: url:yam.gift/2025/05/01/NLP/LLM-Training/2025-05-01-Seed-Thinking-Qwen3 — ByteDance Seed-Thinking recipe + Qwen3 read-across (CN, speculation)
3
#755 opened 9 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
about 12 hours ago
source: url:lemmata.substack.com/p/alphaproof-and-the-imo — AlphaProof mechanism reconstruction (DeepMind, speculation)
3
#758 opened 9 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
about 19 hours ago
source: url:blog.jxmo.io/p/how-to-scale-rl-to-1026-flops — Scaling RL to 10^26 FLOPs / RL-vs-pretraining thesis (speculation)
4
#761 opened 9 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
about 23 hours ago
source: url:cameronrwolfe.substack.com/p/rl-scaling-laws — RL scaling laws for LLMs (secondary synthesis)
4
#767 opened 9 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
1 day ago
source: url:interconnects.ai/p/grok-4-an-o3-look-alike-in-search — Grok 4 as o3-look-alike / xAI RL read (speculation)
3
#762 opened 9 days ago by
lvwerra
source: url:latent.space/p/noam-brown — Latent Space x Noam Brown: test-time compute + self-play limits (transcript)
4
#782 opened 9 days ago by
lvwerra
source: url:hkust-nlp.notion.site/simplerl-reason — SimpleRL-Zero 7B/8K R1-Zero reproduction / easy-to-hard (open-repro)
4
#774 opened 9 days ago by
lvwerra
source: url:lesswrong.com/posts/wwRgR3K8FKShjwwL5 — Reward hacking = spec-gaming not reward-optimization (speculation)
4
#771 opened 9 days ago by
lvwerra
source: url:pillumina.github.io/posts/aiinfra/02-slime — Zhipu GLM slime RL-infra source reconstruction (CN, speculation)
4
#764 opened 9 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
2 days ago
source: url:dwarkesh.com/p/sholto-trenton-2 — Dwarkesh x Sholto/Trenton: Anthropic RLVR insider framing (transcript)
3
#776 opened 9 days ago by
lvwerra
source: url:cnblogs.com/volcengine-developer/articles/19070102 — veRL+ReTool tool-use RL reproduction / env+loss plumbing (CN, speculation)
4
#770 opened 9 days ago by
lvwerra
source: url:dwarkesh.com/p/andrej-karpathy — Dwarkesh x Karpathy: RL-is-terrible critique / reward hacking (transcript)
3
#777 opened 9 days ago by
lvwerra
source: url:cognitiverevolution.ai/everything-you-wanted-to-know-about-llm-post-training-with-nathan-lambert-of-allen-institute-for-ai — Nathan Lambert on open post-training / RLVR-origin / Tulu 3 (practitioner transcript)
3
#787 opened 9 days ago by
lvwerra
lvwerra
in
rl-llm-wiki/knowledge-base
3 days ago
source: url:17aitech.com/p-38744 — R1 vs Kimi1.5 recipe contrast (zh, speculation)
3
#725 opened 11 days ago by
lvwerra
Load more