Evals and the CI gate
Report eval results from CI with the CLI or SDKs, and block label moves until a version has a passing eval run.
An eval run records how one version of a prompt performed on one model: per-case outputs, scores and pass/fail, plus a summary (pass rate, cost, latency). Runs show up on the version's page, drive the "verified on" badges, and can gate releases.
The model
- Datasets and cases: an eval dataset belongs to a prompt; each case holds input variables and, optionally, an expected output and tags. Datasets can be imported from CSV or JSONL.
- Runs: a run is for one prompt version and one model, optionally on a dataset. It is reported by CI, an SDK or by hand, and ends as passed or failed.
- Results: one per case: the output, scores, passed, latency and cost.
You run the model yourself (in CI, in a notebook, anywhere); Incantory stores and compares the results. Nothing is sent to a model by Incantory.
Reporting results
Write one JSON object per line:
{"caseId": "c1", "input": {"topic": "autumn"}, "output": "Crisp leaves drifting", "passed": true, "score": 0.92, "latencyMs": 812}
{"caseId": "c2", "input": {"topic": "rain"}, "output": "...", "passed": false, "score": 0.41}Fields: caseId, input, output, expected, score, passed, metrics, model, latencyMs (all optional).
With the CLI
export INCANTORY_TOKEN=inc_... # a token with the evals:write scope
npx incantory eval report results.jsonl --ref alice/haiku@v3 --model claude-opus-5-5
npx incantory eval report more.jsonl --run <run-id>--ref creates a new run for that version; --run appends to an existing one. Results are uploaded in batches.
With the TypeScript SDK
import { Incantory, parseJsonl } from '@incantory/sdk';
import { readFile } from 'node:fs/promises';
const inc = new Incantory();
const results = parseJsonl(await readFile('results.jsonl', 'utf8'));
const { runId } = await inc.evals.report(results, { create: { ref: 'alice/haiku@v3', model: 'claude-opus-5-5' } });With the Python SDK
from incantory import Incantory
with Incantory() as client:
results = client.evals.read_jsonl("results.jsonl")
client.evals.report(results, create={"ref": "alice/haiku@v3", "model": "claude-opus-5-5"})Under the hood these call POST /api/v1/evals/runs and POST /api/v1/evals/runs/{id}/results; see the REST API reference.
Use a CI token
Create a token with only the evals:write scope for CI. It can report eval runs and read, but it cannot commit versions or move labels, so a leaked CI secret cannot publish anything.
The label gate
Set requirePassingEval on a label and it can only move to a version that has a passing eval run:
npx incantory label set alice/haiku production v3 --require-passing-evalFrom then on, moving production to a version without a passing run fails with 403 eval_required, in the UI, the API, the SDKs and the CLI alike. Combine it with --protected so only you (or an admin) can move the label at all.
A typical pipeline:
- Push a new version (
incantory push -m "..."), which becomesv4. - Run your evals against
alice/haiku@v4and report them withincantory eval report. - Move the label:
incantory label set alice/haiku production v4. It succeeds only if the run passed.
The repository contains a complete GitHub Actions example of this pipeline in docs/examples/github-action-eval-gate.yml.
Getting notified
Subscribe a webhook to eval.run.completed and eval.run.failed to hear about finished runs, for example to move a label automatically or post to chat.