Score your model
on your task.
Most eval tools answer “how does this model do on MMLU”, which tells you nothing about your work. evalkit is a small, dependency-free harness: point it at any OpenAI-compatible endpoint, name a pack, and get a report of what passed, what failed and why. It grew out of this catalog’s stance, and it lives in its own repo.
One command, no install.
Node 20+. Plain fetch, no framework, no Python. Nothing you run is uploaded; results stay on your machine.
npx evalkit list npx evalkit run packs/receptionist-v1 \ --model qwen3 \ --base-url http://127.0.0.1:11434/v1
The packs that exist today.
Each pack is a fixed dataset and its scoring rules, versioned. Results are only comparable within one pack. Open one to read its dataset and checks.
AI receptionist: does it follow the rules?
receptionist v1 · 8 items · 6 kinds of failureWhat it’s for. For anyone running a voice or chat agent on real calls: a garage, a clinic, a booking line. Run it after every change to the prompt, the tools or the model, and see whether the agent still keeps the rules it must not break.
evalkit run packs/receptionist-v1 --model <name>
Coding from a spec: does the model implement what was asked?
dev-spec v1 · 6 items · 6 kinds of failureWhat it’s for. For triaging a coding model before you wire it into an agent, and for catching regressions when you change a prompt or swap a model. It reads the code the model returns; it does not run it.
evalkit run packs/dev-spec-v1 --model <name> --max-tokens 1500
How it stays honest.
Four rules, borrowed from people who benchmark on real hardware and say so.
Results are comparable only within one pack version. A pack is a dataset plus its scoring rules, versioned.
Every report lists what failed and why, not just a score. "17/20" tells you little; "missed the three items about dates" tells you where to look.
A pack measures a setup: model, quantization, runtime and settings together. Change one thing, run again, compare.
Every report carries its caveats: one run per item, temperature 0, setups not models.
A pack is a file, not a submission.
Your real task is not in either pack above, and that is the point. A pack is a folder you keep: a manifest.json and a dataset.jsonl. Write one from the work you actually do, run it locally, and it is yours. Nothing is uploaded, and no account is involved.
packs/my-task/ manifest.json # name, version, dataset file, scoring dataset.jsonl # one item per line: input + expect[] checks
A pack scores the shape of an answer, not execution (for code) or a live phone call (for an agent). It is a fast screen to catch regressions and triage, not a certificate. A pack must be publishable: real customer data, private transcripts or personal information never belong in one. Run those locally and keep the results.