math500

Here is 1 public repository matching this topic...

shaheennabi / rlvr_grpo-experiment-with-math500

A small experiment repository comparing a base reasoning model against RLVR-GRPO checkpoints on the Math500 dataset. It includes evaluation results, short-form observations, and a local temp_clone of the full open-posttraining-system codebase for reference.

reinforcement-learning post-training evaluating-models policy-optimization sparse-rewards reasoning-models rlvr-grpo math500 grpo-checkpoint open-posttraining-system

Updated Jun 17, 2026
Jupyter Notebook

Improve this page

Add a description, image, and links to the math500 topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the math500 topic, visit your repo's landing page and select "manage topics."

Learn more

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

math500

Here is 1 public repository matching this topic...

shaheennabi / rlvr_grpo-experiment-with-math500

Improve this page

Add this topic to your repo