Grpo Rlvr Training
Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically checkable (math, code, tool calls, structured output), when designing GRPO reward functions, or when a GRPO run diverges or reward-hacks.
wshobson/agentsv10 stars · 0 forks · 0 makes≈1.9K tokens
Scores by version
No eval runs reported yet. Incantory never runs prompts: owners report results from their own CI.
Datasets
No datasets.