Full directory
BENCHMARKS

Score your model
on your task.

Most eval tools answer “how does this model do on MMLU”, which tells you nothing about your work. evalkit is a small, dependency-free harness: point it at any OpenAI-compatible endpoint, name a pack, and get a report of what passed, what failed and why. It grew out of this catalog’s stance, and it lives in its own repo.

GET IT

One command, no install.

Node 20+. Plain fetch, no framework, no Python. Nothing you run is uploaded; results stay on your machine.

npx evalkit list
npx evalkit run packs/receptionist-v1 \
  --model qwen3 \
  --base-url http://127.0.0.1:11434/v1
PACKS

The packs that exist today.

Each pack is a fixed dataset and its scoring rules, versioned. Results are only comparable within one pack. Open one to read its dataset and checks.

AI receptionist: does it follow the rules?

receptionist v1 · 8 items · 6 kinds of failure
Open in the repo

What it’s for. For anyone running a voice or chat agent on real calls: a garage, a clinic, a booking line. Run it after every change to the prompt, the tools or the model, and see whether the agent still keeps the rules it must not break.

What it measures
Invented pricesUnconfirmed slotsThird-party confirmationEscalationStructured extractionOff-topic drift
evalkit run packs/receptionist-v1 --model <name>

Coding from a spec: does the model implement what was asked?

dev-spec v1 · 6 items · 6 kinds of failure
Open in the repo

What it’s for. For triaging a coding model before you wire it into an agent, and for catching regressions when you change a prompt or swap a model. It reads the code the model returns; it does not run it.

What it measures
Wrong signatureIgnored edge casesForbidden shortcutsUnsafe SQLSwallowed errorsWrong return shape
evalkit run packs/dev-spec-v1 --model <name> --max-tokens 1500
THE RULES

How it stays honest.

Four rules, borrowed from people who benchmark on real hardware and say so.

A fixed pack, or it is not a comparison.

Results are comparable only within one pack version. A pack is a dataset plus its scoring rules, versioned.

The failures are the result.

Every report lists what failed and why, not just a score. "17/20" tells you little; "missed the three items about dates" tells you where to look.

One variable at a time.

A pack measures a setup: model, quantization, runtime and settings together. Change one thing, run again, compare.

A number without its limits is a claim.

Every report carries its caveats: one run per item, temperature 0, setups not models.

WRITE YOUR OWN

A pack is a file, not a submission.

Your real task is not in either pack above, and that is the point. A pack is a folder you keep: a manifest.json and a dataset.jsonl. Write one from the work you actually do, run it locally, and it is yours. Nothing is uploaded, and no account is involved.

packs/my-task/
  manifest.json   # name, version, dataset file, scoring
  dataset.jsonl   # one item per line: input + expect[] checks
What this is not.

A pack scores the shape of an answer, not execution (for code) or a live phone call (for an agent). It is a fast screen to catch regressions and triage, not a certificate. A pack must be publishable: real customer data, private transcripts or personal information never belong in one. Run those locally and keep the results.

Learn the method