Arize Evaluator
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and continuous monitoring. Use when the user mentions create evaluator, LLM judge, hallucination, faithfulness, correctness, relevance, run eval, score spans, score experiment, trigger-run, column mapping, continuous monitoring, or improve evaluator prompt.
Awesome GitHub Copilotv10 stars · 0 forks · 0 makes≈7.9K tokens
Scores by version
No eval runs reported yet. Incantory never runs prompts: owners report results from their own CI.
Datasets
No datasets.