Learning with Verifiable Rewards
An empirical study of GRPO, controlled ablations, uncertainty, and failure modes in reinforcement learning with verifiable rewards.
Personal project
🔨 In Development — model foundations complete
An empirical study of GRPO, controlled ablations, uncertainty, and failure modes in reinforcement learning with verifiable rewards.