Incantory
Sign in

Evals and the CI gate

Report eval results from CI with the CLI or SDKs, and block label moves until a version has a passing eval run.

An eval run records how one version of a prompt performed on one model: per-case outputs, scores and pass/fail, plus a summary (pass rate, cost, latency). Runs show up on the version's page, drive the "verified on" badges, and can gate releases.

The model

  • Datasets and cases: an eval dataset belongs to a prompt; each case holds input variables and, optionally, an expected output and tags. Datasets can be imported from CSV or JSONL.
  • Runs: a run is for one prompt version and one model, optionally on a dataset. It is reported by CI, an SDK or by hand, and ends as passed or failed.
  • Results: one per case: the output, scores, passed, latency and cost.

You run the model yourself (in CI, in a notebook, anywhere); Incantory stores and compares the results. Nothing is sent to a model by Incantory.

Reporting results

Write one JSON object per line:

json
{"caseId": "c1", "input": {"topic": "autumn"}, "output": "Crisp leaves drifting", "passed": true, "score": 0.92, "latencyMs": 812}
{"caseId": "c2", "input": {"topic": "rain"}, "output": "...", "passed": false, "score": 0.41}

Fields: caseId, input, output, expected, score, passed, metrics, model, latencyMs (all optional).

With the CLI

bash
export INCANTORY_TOKEN=inc_...   # a token with the evals:write scope
npx incantory eval report results.jsonl --ref alice/haiku@v3 --model claude-opus-5-5
npx incantory eval report more.jsonl --run <run-id>

--ref creates a new run for that version; --run appends to an existing one. Results are uploaded in batches.

With the TypeScript SDK

ts
import { Incantory, parseJsonl } from '@incantory/sdk';
import { readFile } from 'node:fs/promises';

const inc = new Incantory();
const results = parseJsonl(await readFile('results.jsonl', 'utf8'));
const { runId } = await inc.evals.report(results, { create: { ref: 'alice/haiku@v3', model: 'claude-opus-5-5' } });

With the Python SDK

python
from incantory import Incantory

with Incantory() as client:
    results = client.evals.read_jsonl("results.jsonl")
    client.evals.report(results, create={"ref": "alice/haiku@v3", "model": "claude-opus-5-5"})

Under the hood these call POST /api/v1/evals/runs and POST /api/v1/evals/runs/{id}/results; see the REST API reference.

Use a CI token

Create a token with only the evals:write scope for CI. It can report eval runs and read, but it cannot commit versions or move labels, so a leaked CI secret cannot publish anything.

The label gate

Set requirePassingEval on a label and it can only move to a version that has a passing eval run:

bash
npx incantory label set alice/haiku production v3 --require-passing-eval

From then on, moving production to a version without a passing run fails with 403 eval_required, in the UI, the API, the SDKs and the CLI alike. Combine it with --protected so only you (or an admin) can move the label at all.

A typical pipeline:

  1. Push a new version (incantory push -m "..."), which becomes v4.
  2. Run your evals against alice/haiku@v4 and report them with incantory eval report.
  3. Move the label: incantory label set alice/haiku production v4. It succeeds only if the run passed.

The repository contains a complete GitHub Actions example of this pipeline in docs/examples/github-action-eval-gate.yml.

Getting notified

Subscribe a webhook to eval.run.completed and eval.run.failed to hear about finished runs, for example to move a label automatically or post to chat.