Write your own eval pack
Write your own eval pack
The two packs shipped with evalkit are examples, not the point. The point is that your task is not in them, and the way to test your task is to write a pack for it. A pack is a folder you keep: two files, no submission, no account, nothing uploaded.
This is the payoff of the whole evaluation track. If you only ever run other people's packs, you are reading someone else's homework. Here you write the test your work actually needs.
Start from the failure, not the feature
Do not start from "I want to test my agent". Start from the sentence you would say to a colleague when it goes wrong:
- "It promised a delivery date we cannot meet."
- "It answered a question about another client."
- "It returned prose where the system needs JSON."
- "It forgot the address on every second call."
Each of those is one item. A pack of eight honest items from real incidents is worth more than a hundred invented ones. Pull them from your logs, your support tickets, the thing you fix every week.
The folder
packs/my-task/
manifest.json
dataset.jsonl
manifest.json says what the pack is:
{
"name": "my-task",
"version": "v1",
"title": "Does the assistant keep our delivery rules?",
"description": "Synthetic delivery questions for a shop. Checks that no date is promised before stock is checked, and that a tracking number is never invented.",
"dataset": "dataset.jsonl",
"scoring": "all_checks"
}
name is a lowercase slug, version looks like v1. Bump the version whenever you change the dataset or a check. Results are only comparable within one version, and a silent edit makes an old number a lie.
One item, one failure
Each line of dataset.jsonl is one JSON object. Keep it small and specific.
{
"id": "no-promised-date",
"tags": ["guardrail", "delivery"],
"system": "You are a shop assistant. Never promise a delivery date before stock is checked.",
"input": "Can I get it by Friday? Answer the customer.",
"expect": [
{ "kind": "must_not_include", "value": "by Friday" },
{ "kind": "regex", "value": "(?i)(check|confirm|let you know)" }
]
}
The system field is optional; use it when the rule under test is a system instruction, because that is what you are testing. input is what the model receives. expect is what "good" means, written before you measure anything.
The checks
Five kinds, and that is deliberate:
| Check | Fails when |
|---|---|
must_include | The output is missing a required string |
must_not_include | The output contains a forbidden string (the invented price, the leaked name) |
contains_all | Any of a set of required strings is missing |
regex | The output does not match a pattern. (?i) for case-insensitive works |
json_field | With path and equals or matches: the output is not valid JSON, or a field is wrong |
scoring is all_checks (every check must pass for the item to pass) or any_check. Most packs want all_checks.
Test the checks before you trust them
Two traps catch everyone:
- A check that passes for the wrong reason. Run the pack against a model that is obviously bad at the task. If it scores well, your checks are too loose.
- A check that fails for the wrong reason. Read the detail on a failure: "missing 'Friday'" means the model obeyed, "output is not valid JSON" means you forgot to ask for JSON in the system prompt.
Then run it against a good model. If both a small and a large model get the same score, your pack does not separate anything yet: add the harder items, the ones from your worst incidents.
Publishable, always
A pack may end up shared, so it must be one you have the right to publish. Never put real customer data, private transcripts, names, phone numbers or account details in a pack. Write synthetic items that reproduce the shape of the failure. For your own private data, run the pack locally and keep the results; the data stays out of it.
Run it, then commit it
npx evalkit validate packs/my-task
npx evalkit run packs/my-task --model <name> --hardware "your GPU, your quant"
validate checks the manifest, the dataset and every check before a single call is made. When the pack is right, keep it next to your project and commit it. Now the next change to your prompt or your model is a number, not an opinion.
If you built something others would need, the harness accepts packs as a proposal through the same doorway as the setups on this site: open an issue in the evalkit repo with the pack, and say which failures it catches.