Eval Harness First
Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.
wshobson/agentsv10 stars · 0 forks · 0 makes≈1.9K tokens
Family tree
1 prompt · 1 version shown · 0 makes · Full screen
No forks or makes yet. When someone forks this prompt or shares something they made with it, it grows here.