Incantory
Sign in
SkillMITNot scanned

Grpo Rlvr Training

Train reasoning and verifiable-task behavior with GRPO and reinforcement learning from verifiable rewards (RLVR). Use when task success is algorithmically checkable (math, code, tool calls, structured output), when designing GRPO reward functions, or when a GRPO run diverges or reward-hacks.

wshobson/agentsv10 stars · 0 forks · 0 makes≈1.9K tokens

Scores by version

No eval runs reported yet. Incantory never runs prompts: owners report results from their own CI.

Datasets

No datasets.

All runs and datasets