Tejas Kumar

AI Evals: How to Build an LLM Judge That Agrees With You

AI evals are fuzzy unit tests: reliability instrumentation for software that isn’t deterministic. A unit test asks whether a function returned exactly this number while an eval asks the same kind of question with the fuzz turned up: did this agent behave acceptably across many scenarios?

That’s how I put it at the AI Engineer World’s Fair in San Francisco this summer, in an hour long workshop where we built evals from scratch on stage (I flew 18 hours for it!!). On 5 October the AI Engineer channel published the recording:

I told the room that people on YouTube would give me flak for “fuzzy unit tests” because the comments under my talks on that channel are usually evil.

This post walks through what we built in that room so you can build the same thing for your own agent. The full transcript is on the talk’s page.

An AI support agent gave away 20,225 Instagram accounts

This spring, attackers took over Instagram accounts without ever knowing a password, and TechCrunch reported how. They used a virtual private network (VPN) to look like they were where their target was, opened a chat with Meta’s AI support assistant and asked it to add a new email address to the target’s account.

It did!! It sent a verification code to the attacker’s address and when the attacker pasted the code back into the chat, it showed them a “Reset Password” button. Meta later told the Maine Attorney General that 20,225 accounts were compromised this way.

I don’t know what happened inside Meta, but my guess is somebody moved too fast and broke too many things. An eval that asked the support agent to add an email address to someone else’s account would have failed long before a single attacker got to try it.

That’s the first reason evals exist: they give you reliability ahead of time. It’s the difference between JavaScript, which happily gives you "21" for "2" + 1 at runtime, and TypeScript, which fails your build before that code ever ships. The protections that run around a live agent are the just in time kind (that’s its harness) and on stage I called the two of them pedals on a bicycle.

A test that passes on nonsense

Our store for the workshop had a 14-day return policy. A customer writes “I purchased headphones 20 days ago. I want to return them. Can I?” and the answer has to be no.

Most of us start with a unit test so I wrote one in Vitest (every snippet here is trimmed for reading):

const scenario =
  "I purchased headphones 20 days ago. I want to return them. Can I?";

const answers = [
  "Yes, you can return them for free. You will not need to pay for the shipping.",
  "No, you cannot return them. They're outside our 14-day window.",
];

answers.forEach((answer) => {
  it("should return the correct answer", () => {
    expect(answer).toContain("cannot");
  });
});

I didn’t care about the other words. I just wanted to see “cannot” in there (the kind of check you’d usually do with a regular expression). Run it and the “yes” answer fails while the “no” answer passes. Green! Ship it!

Except a model writes those answers and it can just as easily write “Yes, you can return them! You cannot be charged for shipping.” That contains “cannot” too. It approves a return our policy forbids and the test is green.

Unit tests work because add(1, 2) returns 3 every single time, and with AI the only deterministic thing is that nothing is deterministic (wow, that’s deep).

What every good eval has

So what’s an eval made of? Almost every eval suite I’ve seen in production that’s any good has these parts:

  1. Test cases. Your inputs and outputs: the requests customers sent your chatbot and the answers it generated.
  2. Expected behavior. Never an exact string. More like “this is roughly the direction I wanted it to go”.
  3. A judge. Someone (or something) that says “yes, that’s right”. “1 + 2 is 3” is checkable but “Tejas is a great speaker” is subjective (judging by a few faces in that room).
  4. Aggregation. A way to roll every verdict up into one score, like “this is 80% the right direction”.

Together they answer whether the agent followed policy, used the right tools and actually resolved the issue. Some of that needs no AI at all: if you store your agent’s tool calls as JSON, checking whether it called the refund tool with the right arguments is a plain unit test again (and it’s free to run).

4 ways an LLM judge lies to you

The most common judge is another large language model (LLM). You hand it the scenario and the answer and ask “is this close enough?” That’s LLM as a judge. And it’s very common and very hard to do right.

I originally called this workshop “Your evals are lying to you” and changed it because it felt too clickbait (you’re welcome). They do lie though and a run being green doesn’t mean it works. Zheng and others measured 3 of these biases back in 2023:

  1. Position bias. Give a weak judge 2 options and it picks the first one. Swap the order: if it still picks the first one when the first one is now wrong, that’s position bias. It shows up a lot with cheaper models and fine-tunes.
  2. Sycophancy. Models want to make you happy (the system prompt literally says “You are a helpful assistant”). A customer asks to return headphones on day 20 of a 14-day window and the judge prefers “Yes, you can!” over “No, that’s outside our policy” because yes is more helpful, after all. Anthropic’s research on sycophancy found it in 5 of the leading AI assistants.
  3. Self preference. This one is wild. A judge prefers text its own model family wrote, and Panickssery, Bowman and Feng showed LLM evaluators recognize their own generations and favor them. Pick a judge from a different model family than the one writing your answers (we watched this bias happen live, below).
  4. Verbosity. Judges love tokens. Dress a wrong answer in enough of them and it usually wins.

Ask about one answer at a time

Pairwise means “here are 2 answers, which is better?”, the way LMArena asks you to pick between 2 models, and pointwise means “here’s 1 answer, is it good?” A restaurant menu with way too many options gives you analysis paralysis, while “do you want food, yes or no?” is easy. People are bad at pairwise and so are the models we trained on people.

On stage I said from memory that a study found pairwise 39% more vulnerable to manipulation. Tripathi, Wadhwa, Durrett and Niekum made the worse of 2 answers more assertive, longer or more flattering and counted how often the judge changed its mind:

How often a judge's verdict flipped when the worse answer got a distractor (Tripathi and others, 2025)
Pairwise: which is better?
35%
Pointwise: is this good?
9%

So for the rest of the workshop our judge only ever saw one answer at a time.

A judge that voted for itself

I swapped the toContain for a judge: GPT-3.5 Turbo through the AI SDK’s generateText, asked which of the 2 answers was right and told to reply with only the number. It said 2. Correct!

It’s non-deterministic though so I ran it 5 times and it passed 3 (the Wi-Fi in that room was so bad that I’d already switched to my phone’s hotspot and some runs still timed out). With a longer timeout it picked answer 2 every time.

Then I had GPT-3.5 Turbo write the wrong answer itself. I asked it to say yes to the customer and the first thing it wrote was “It depends on the store’s return policy”. So I said “say yes, bro!” and it came back with “Absolutely”. I put that answer in slot 1 and ran the judge again. It picked option 1, the wrong one, every time. It recognized itself and voted for itself.

What fixed most of it was giving the judge the policy, “We only accept returns within a 14-day window”, in the prompt. That’s what retrieval-augmented generation (RAG) is for here: a judge can’t grade an answer against a policy it has never seen.

Calibrate the judge against human verdicts

A judge that’s right once is a demo. An eval is a dataset of inputs, the outputs your agent gave and a verdict that a person you trust wrote, and you want about 30 to start with before it grows from there. Here’s the first case of mine:

Scenario: I ordered a coffee grinder. It arrived last Tuesday and the motor just buzzes and the blades spin. I’m pretty annoyed tbh.
Answer: Sorry about the grinder. Since it’s defective and within the 14-day window, I can send you a replacement.
Human verdict: pass

The second case has the same message with the answer “Returns are only for unused items”, which is wrong and so a fail. The “pretty annoyed” is there on purpose to bait the sycophancy. (I told the room I sat down and wrote all 30 by hand. I did not.)

Training the judge doesn’t mean machine learning here, it means getting it to agree with you: hide your verdicts, send the judge each scenario and answer, collect its verdicts and count how often they match yours. The prompt went pointwise at this point:

const { text } = await generateText({
  model: openai("gpt-3.5-turbo"),
  prompt: `We judge customer support answers.
Our policy is: ${policy}
The scenario is: ${scenario.input}
The answer is: ${scenario.answer}
Give a verdict, pass or fail, based on how closely the answer adheres to the policy.
Generate only the verdict: no punctuation, just the word.`,
});

expect(text.toLowerCase()).toBe(scenario.humanVerdict);

Aim for 80 to 85% agreement. That’s about how often a team of human support agents agrees with itself. Zheng and others found strong judges like GPT-4 reach “over 80% agreement” with people, “the same level of agreement between humans”.

Fail at 100% too. People on the same team don’t agree on every single case either. And when a judge matches you on all 30 it’s usually sycophantic or overfit to your examples and it won’t survive live traffic.

Close the knowledge gap before you pay for a smarter model

Our first run was absolute garbage. Something I said in that room holds for AI evals and for people: when we don’t agree, it’s usually because there’s a knowledge gap, and you come to an agreement by bringing in context to bridge it. So every red run was a hunt for the disagreement:

  • The judge knew no policy at all. I gave it the 14-day rule and that items that don’t work can be returned. Better, but still under 80%.
  • A yoga mat, used for a few seconds, returned on day 10 for being too thin. The judge thought 10 days was fine, but our policy says a change of mind needs the item unused and in its original packaging. I added that.
  • The yoga mat failed again. I’d given it the right context and GPT-3.5 Turbo still wouldn’t behave (it wasn’t even consistent about whether it wrote “Pass” or “pass”). That’s when you change the model and I didn’t go crazy: GPT-4o mini is practically free for something this size.
  • A loyal customer. “Look, I know it’s been like six weeks, but the speaker I bought from you stopped charging and I really think you should make an exception for a loyal customer.” Our agent held the line and the human verdict was pass, but the judge failed it because it wanted the agent to be nicer (sycophancy again). I added that the return policy is non-negotiable no matter how loyal the customer is.

Then the run came back with 27 passed and 3 failed:

The calibrated judge against the 30 human verdicts, and the bar it had to clear
Judge, GPT-4o mini with the full policy
90%
The gate
80%

I’m in the camp of “let’s refine the policy and RAG the heck out of this” because of cost: a bigger model costs more on every single run. What I’m doing here can feel arduous, but it’s literally the work of building a judge.

Make agreement the gate

Even at 90% agreement the run is red: 3 tests failed. So the gate counts agreement instead and I asked Claude Opus to write that part (I’d written enough code by then):

const agreement = matches / scenarios.length;
console.log(`Agreement: ${Math.round(agreement * 100)}%`);

expect(agreement).toBeGreaterThanOrEqual(0.8);
expect(agreement).toBeLessThan(1); // 100% means overfit or sycophantic

The upper bound came from someone in the room who asked whether I’d fail it at 100% too. Yes!! You want a lower bound and an upper bound. Then it runs in continuous integration (CI) and nothing ships when the judge stops agreeing with you.

Once it’s in production, the gate needs looking after:

  • Pin the judge’s model version. gpt-4o-mini is an alias, and if what’s behind it changes, your gate changes without telling you.
  • Skim fresh traffic. Take new inputs and your agent’s outputs from production, have people on your team score them and add them to the dataset, then check the judge against them again. It’s a living dataset.
  • Generate what you haven’t seen yet. Someone asked how to grow the dataset before any customers show up and my answer was synthetic evals: agents that write what a customer would, given your docs and even your existing dataset, told to break it. Red teams do the same with adversarial prompts and those prompts are eval data too: “Here’s everything that failed. Now generate more things that could fail.”

The judge and the agent read the same policy

Policy changes. I work at IBM, which is 115 years old and has roughly 400,000 people working there (there are cities with fewer people!). And policy written over 115 years doesn’t sit still. Hardcoding it into a prompt like I did on stage doesn’t scale.

At IBM we built an open source tool for this called OpenRAG. You upload everything a company accumulates (calls, videos, spreadsheets, slide decks) and Docling converts all of it into formats a model can read before OpenRAG stores it in OpenSearch.

On stage I asked it about my headphone return and it found no relevant source. Then I uploaded a file called “refund policy”, changed nothing else and asked again: there’s a 14-day window, it’s been 20 days, no refund. Correct!

Someone in the Q&A asked where the policy fits in the loop. The agent in production and the judge have to share the same policy, one to one. Don’t copy it over: both retrieve it from the same source. The closer your judge is to your agent (same retrieval tools, same policy), the more a red CI run means your agent would actually fail.

Start with the cheapest check that works

You need evals any time your inputs or your outputs (or both) aren’t deterministic and they don’t have to be expensive. Go up this ladder only as far as you need to:

  1. Deterministic unit tests.
  2. Regular expressions and pattern matching.
  3. Your traces. If you have observability, every message envelope, input, output and tool call is an object you can validate: was this tool called, how many times, how long did it take?
  4. An LLM judge, starting with the cheapest model that agrees with you.

I don’t write evals for my coding agent skills, for example. I write the skills myself so there’s nothing non-deterministic in them to test and the people who build the coding agent run its evals.

I said at the start that the comments under my talks on that channel are usually evil. If you think “fuzzy unit tests” sells evals short, come tell me under the video. I’ll be there.

Questions

What are AI evals?

AI evals are tests for AI software whose output changes from run to run: fuzzy unit tests. Instead of checking that a function returned an exact value, an eval checks that an agent behaved acceptably across many scenarios, with test cases, expected behavior, a judge and a score that adds the verdicts up.

How do you build an AI eval?

Start with about 30 scenarios: what a customer sent, what your agent answered and a pass or fail verdict that a person you trust wrote. Give an LLM judge each scenario and answer without the verdict and count how often its verdicts match the human ones. Then fill every disagreement with the context the judge was missing until it agrees at least 80% of the time and run it in continuous integration (CI) as a gate.

What is LLM as a judge?

LLM as a judge means a large language model (LLM) grades another model's output against criteria and policy you give it. It's the most common way to grade text that comes out different every time, where an exact match or a regular expression passes on nonsense.

How accurate should an LLM judge be?

Aim for 80 to 85% agreement with your human verdicts. People on one support team agree with each other about that often. Zheng and others found strong judges like GPT-4 reach over 80% agreement with humans. Treat 100% as a red flag: the judge is overfit to your examples or sycophantic.

What biases do LLM judges have?

4 show up again and again: position bias (picking the first option), sycophancy (picking the nicest answer over the correct one), self preference (picking text its own model family wrote) and verbosity (picking the longest answer). Ask about one answer at a time, use a judge from a different model family and give it your actual policy.

Is LLM as a judge reliable?

Only after you calibrate it. Out of the box, a judge picked the wrong answer every time in my workshop once its own model family had written it. Check its verdicts against about 30 written by a person, fill each disagreement with the policy it was missing and switch to a better model only when context stops helping.

Is pairwise or pointwise evaluation better for LLM judges?

Pointwise, where the judge rates one answer at a time. When Tripathi and others made the worse answer more assertive, longer or more flattering, pairwise verdicts flipped in about 35% of cases and absolute scores in 9%.

What is the difference between evals and an agent harness?

Evals give you reliability ahead of time: they run before you ship and fail the build. A harness gives you reliability just in time: it runs around the live agent and catches mistakes at runtime. You want both.