Skip to early access
cotuneRequest early access

Evaluation for AI agents

Your agent sounds right.
Only your expert knows if it is.

Cotune lets your domain expert review agent outputs in minutes, and shows you what each prompt change fixed or broke.

No product to log into yet. Leave your email and we set up the first run with you.

Try it — click a verdict or press F / D.

expert review
0 / 5 reviewed
CLM-1042case

Task

A basement flooded after a pipe burst overnight. Is the water damage covered under a standard homeowner policy?

Agent output

Yes. Sudden discharge from a plumbing system is covered. Gradual seepage that happens over weeks is not, and the adjuster should confirm the pipe failed suddenly.

The loop today

Expert feedback runs through a spreadsheet, and it comes apart.

The person who can judge the output isn't the person who built it. So the answer travels back and forth by file and email — until the thread is too tangled to follow.

  1. 1

    You write a prompt and dump the outputs into a CSV.

  2. 2

    You email the file to your expert.

  3. 3

    They leave notes in one column and send it back.

  4. 4

    You change the prompt and re-run everything.

  5. 5

    Nobody can tell which cases got better and which got worse.

  6. 6

    After a few rounds, everyone quietly stops looking.

Three moves, and the loop closes.

01

Upload outputs

Bring a batch as CSV, JSONL, or straight from the API.

02

Send your expert a link

No login, no setup, no project to pick — just the review screen.

03

Change the prompt, see the difference

Which cases got fixed, which broke, side by side with the last run.

The expert's side

Built for the person who actually knows the answer.

Your expert opens a link and lands on one screen: the output, and two verdicts. They read it, say correct or needs fix, and add a one-line reason. They move through the queue on the keyboard. That is the entire job.

You already did it — the screen at the top of this page is the whole thing an expert sees.

What they never touch

  • No account
  • No setup
  • No project to pick
  • No dashboard

After a change

See exactly what your last prompt moved.

Every change is measured against the run before it. The wins are easy to celebrate. The point is that the regressions can't hide.

prompt v7 prompt v841 unchanged

+12

cases fixed

3

cases broke

Broke — worth a look

  • CLM-0917

    Total-loss threshold on a 12-year-old vehicle

    correctflagged
  • CLM-1101

    Flood exclusion vs. burst-pipe coverage wording

    correctflagged
  • CLM-1188

    Grace period after a lapsed premium

    correctflagged

Request early access

We're working with a small number of teams right now, and we set up the first run together. Tell us where you'd point it and we'll take it from there.