Eval Driven Dev
Improve AI application with evaluation-driven development. Define eval criteria, instrument the application, build golden datasets, observe and evaluate application runs, analyze results, and produce a concrete action plan for improvements. ALWAYS USE THIS SKILL when the user asks to set up QA, add tests, add evals, evaluate, benchmark, fix wrong behaviors, improve quality, or do quality assurance for any Python project that calls an LLM model.
Awesome GitHub Copilotv10 stars · 0 forks · 0 makes≈4.3K tokens
Family tree
1 prompt · 1 version shown · 0 makes · Full screen
No forks or makes yet. When someone forks this prompt or shares something they made with it, it grows here.