Evaluation for AI agents
Your agent sounds right.
Only your expert knows if it is.
Cotune lets your domain expert review agent outputs in minutes, and shows you what each prompt change fixed or broke.
No product to log into yet. Leave your email and we set up the first run with you.
Try it — click a verdict or press F / D.
Task
A basement flooded after a pipe burst overnight. Is the water damage covered under a standard homeowner policy?
Agent output
Yes. Sudden discharge from a plumbing system is covered. Gradual seepage that happens over weeks is not, and the adjuster should confirm the pipe failed suddenly.
The loop today
Expert feedback runs through a spreadsheet, and it comes apart.
The person who can judge the output isn't the person who built it. So the answer travels back and forth by file and email — until the thread is too tangled to follow.
- 1
You write a prompt and dump the outputs into a CSV.
- 2
You email the file to your expert.
- 3
They leave notes in one column and send it back.
- 4
You change the prompt and re-run everything.
- 5
Nobody can tell which cases got better and which got worse.
- 6
After a few rounds, everyone quietly stops looking.
Three moves, and the loop closes.
Upload outputs
Bring a batch as CSV, JSONL, or straight from the API.
Send your expert a link
No login, no setup, no project to pick — just the review screen.
Change the prompt, see the difference
Which cases got fixed, which broke, side by side with the last run.
The expert's side
Built for the person who actually knows the answer.
Your expert opens a link and lands on one screen: the output, and two verdicts. They read it, say correct or needs fix, and add a one-line reason. They move through the queue on the keyboard. That is the entire job.
You already did it — the screen at the top of this page is the whole thing an expert sees.
What they never touch
- No account
- No setup
- No project to pick
- No dashboard
After a change
See exactly what your last prompt moved.
Every change is measured against the run before it. The wins are easy to celebrate. The point is that the regressions can't hide.
+12
cases fixed
−3
cases broke
Broke — worth a look
- CLM-0917correct→flagged
Total-loss threshold on a 12-year-old vehicle
- CLM-1101correct→flagged
Flood exclusion vs. burst-pipe coverage wording
- CLM-1188correct→flagged
Grace period after a lapsed premium
Request early access
We're working with a small number of teams right now, and we set up the first run together. Tell us where you'd point it and we'll take it from there.