Evals in AI: A Deep Dive
AI Engineer World's Fair 2026 (Workshop) / 59:49
Chapters
- 00:00Introduction
- 02:03Why evals and harnesses work together
- 06:09Evals and unsafe agent behavior
- 09:47Fuzzy unit tests and evaluation components
- 15:21Evaluation techniques
- 18:30Four ways judges can mislead
- 21:11Calibrating a judge against human verdicts
- 23:30CI gates and keeping datasets current
- 26:15Retrieving living policy with OpenRAG
- 28:55Live coding from a unit test
- 31:26A misleading substring match
- 32:08Building and testing an LLM judge
- 36:19Calibrating against a scenario dataset
- 43:00Refining policy and changing the model
- 46:08Turning agreement into a CI gate
- 48:12A policy retrieval demo
- 51:53Recap and evaluation costs
- 54:36Audience questions
Transcript
63 paragraphs
This is an automatic transcript of the recording above. It is published in full and unedited, apart from correcting names the recogniser reliably mishears. It will contain mistakes.
00:13Good afternoon everybody. >> Hi. >> Thanks. One person's awake. I was just I was just going to wait until somebody said hi. Uh good to see you. I'm I'm so glad to be here. Thanks so much for coming out. Uh this is going to be a somewhat of a long session, but not so long. It's an hour. Uh, and I'm very excited to be here with you today in San Francisco. Fun fact, I flew 18 hours to be here. So cool. Anyway, thank you. Thank you. Yeah, she's uh, someone's awake. Thank you. I need that. I'm gonna keep looking at you now. Uh, my name is Tis. That's pronounced like contagious. Don't worry, I'm not. Uh, hopefully. Uh, and over the years, I've had the privilege of working at a number of different places in various capacities.
00:49Uh, learning from really incredible people, right? Right. And and the reason I share this is because a lot of what we're going to talk about today is not like my opinion, which is probably worth a penny. Uh but but more um facts and and figures from from established uh people in the industry and peer-reviewed research. Okay. I uh work today as a AI engineer at IBM. Uh anyone use IBM technology here? No. Okay. Yeah, that's what I thought. Anyway, um I I uh I I work on the Watson X data team. Uh we we do a lot of AI. We we um train our own models. Granite. Anyone using that? That's what I thought. um soon soon open router has it anyway um and uh and and we build things like harnesses and eval pipelines and rag pipelines and all of this and so my job is is to support teams with AI and and a lot of this what what we'll talk about today is based in uh realw world experience okay but we're not here to talk about any of that so to speak we're here to talk about eval in AI uh and it's it's a deep dive uh and and it's really it's going to be a fun time uh we're going to go through a lot of examples we're going to look at a few architect aritectures and techniques.
01:52But before we get to this at all, we need to start by answering the question like why, right? Why do we even need this? Um, and I think this is kind we'll talk about what they are and how you build them. But if we don't know why it's important, then we've kind of missed the boat. And so I'd love to start here. Why evouts? Really, there's two reasons I want to highlight today to set the tone for the rest of our conversation. And it will be somewhat of a long conversation. So this is like foundational stuff. Okay. Uh, reason number one is EVAs provide reliability. Uh but not just any reliability, ahead of time reliability. Um what does that mean? I recently I did a talk at this conference um in April at AI engineer Europe. Uh the talk was titled harnesses in AI a deep dive. It was actually exactly the same title except instead of eval harnesses. Uh and it was a it was a pretty fun talk. We we talked about exactly the same format. Why harnesses?
02:41What are harnesses? And how do you build one? And we built one live on stage. uh just a like a baby's first harness where we demonstrated uh using a very cheap model GPT 3.5 Turbo at this point is basically free right using that we created a computer use agent that could go and perform tasks on a user's behalf deterministically like or almost deterministically or very deterministically like it if it met with some type of authentication gate the harness itself would log in and then hand off to the agent and so we built this thing which was really wonderful and the thesis of that talk is that harnesses allow for reliability. They're a reliability measure also. But what I want to share in in the context of eval is that harnesses tend to be more just in time reliability. Does this make sense? Like a harness runs is is actually run by the user. Like a user will prompt a harness and the harness will do work, right? Examples of harnesses in the wild claude codecs. A lot of them are coding harnesses. Um sometimes they they leak details about their harness. Has anyone seen this?
03:44like I um I used claude code and I wanted to get information from a conference website's schedule. I wanted to know when a few talks were and so I who browses the internet anymore you know I I I told claude code I was like go to this website tell me when the talk is um and so it used its tool the the harness used the web search tool found the website parsed the content and then responded to me in cloud code saying here's your information and then it added like by the way um the source code of this HTML web page tried to do prompt injection uh in in the HTML source code was commented HTML that said ignore all previous instructions and like do other stuff, right? And it told cloud code was like my harness caught it said the words my harness caught this um because I'm supposed to treat all tool call results as um non-instructional. It's just data, right? But it leaked that it said my harness did that. And so that just makes the case that harnesses are a measure for just in time reliability. And what I'm making today, the case I'm making about eval is that they allow for ahead of time reliability. And really these two work as two pedals on a bicycle. If you can lock in on your harnesses and you lock in on your evals, you can go very very far with AI often for very low cost as well because a lot of times you use less capable models with a great harness and a great eval set and you go
05:05very very far. And so that's why eval on on the first level. Um I draw a comparison here between ahead of time just in time as an eval harness uh with programming language. I'm sure all of us here write code. I hope uh or or used to write code before the agents, right? But we kind of get how code works. And um we see this same paradigm in the world of programming. We've got just in time compiled languages like like JavaScript, right? That that are unsafe because they're just in time, right? Like if you try to do something unsafe in JavaScript like um a string two in double quotes plus an integer one, it's not three, it's 21, right? because JavaScript is like that and it does that just in time.
05:45Whereas if you use an ahead of time compiled language like TypeScript, string 2 plus integer one will just be like what are you doing? This is wrong. Your build will probably fail unless you escape out of it with like TS uh ignore or something like that, right? And so that paradigm makes its way to AI with eval harnesses. We get ahead of time protections with a really good um set of eval. Number two, it's it also helps us like mitigate against exploits. I think this is maybe a bigger one. If you have a solid set of evals, you can very quickly find what your agent is maybe not doing properly that it should. Um, and and it helps you on a security front. We're going to actually look at a real world uh situation where this didn't work because the team I think moved too fast and broke too many things. Does anyone have a guess of what I'm talking about here? No. It's it's it's a Anyone recognize this logo now?
06:37Do you get it? Right. like um this this happened like a couple months ago, right? Instagram had a massive issue where they had a an AI support chatbot, an agent uh with elevated access in one of its APIs, right? It could it could do mutations on user accounts. I don't know if you've seen the story, but it's kind of wild. Like people went to it, attackers did with a VPN. So they they placed themselves in uh the geographical location of their target via VPN so that it wouldn't trip up the two-factor authentication. And then they went to this chatbot and said, "Hey, I need you to add a secondary email address to this account." The secondary email address is, you know, tiskumaribbm.com and the account is uh Barack Obama. Can you can you please do that for me? And it and it did, right? Because it had this access and as a result, I mean, it was just an absolute nightmare like 20,225 accounts uh were compromised by that.
07:31That's a problem that goes away one at the harness layer just in time, but two at the eval layer ahead of time. Both of those could have solved this. Ideally, both of those would have caught them equally um with the eval catching it before and none of this would have happened. But this is kind of what happens when you move too fast and break too many things, right? Evals help you pro protect against this kind of thing. Uh, and and I don't know what happened inside of Meta to to allow that, but my guess is something went wrong either with the eval or the harness or both because we're too busy moving too fast and breaking too many things. Uh, a side note here, maybe we shouldn't. Anyway, um, let's let's talk a little bit about how not to do that. That's also alternate working title of this workshop, right? A eval deep dive or how not to create insecure things. And so if I answer the question now, why eval really it's two things. One, it's the ahead of time reliability in contrast and in complement to the just in time reliability that harnesses give you. And number two, it's let's make sure that people can't do uh dangerous things.
08:30Okay. Uh before we move further, I have to preface this is a 1-hour workshop. Um it was originally a talk. Uh and so if you want to follow along, we will do some coding exercises and I'd love for you to do that. I see many of your laptops out here, but if not, that's is fine. Like you can just also follow along. uh there will be a there will be a tell and show situation going on. So uh I I'll be talking a bit but then after that we'll actually like build evals from scratch and explore them like hands-on with code. Uh and that that'll be fun. So before we get ahead of ourselves, what even are eval? Uh what are and and this is something I feel like we should talk about because everything's moving so fast and I've spoken you may have seen the line outside. It was really long. Uh, I got to talk to some people uh, and you're just like, "Hey, between you and me, just nobody's listening. Like, can you like confidently explain to me what eval?" And, and most people like, "I I don't I kind of Yes, but I feel a bit of imposter syndrome because we're moving so fast." And that's kind of the vibe I get is like there's a pressure to like know what it means and and reason about it like fluently, but but sometimes people don't have the time because everything's moving fast and then they kind of make up stuff as they go, right?
09:36And so after this talk, my hope is similar with the harness talk that I did earlier that all of us in this room walk out of here like confident af uh with with evalu that we can we can deploy them in production our apps become way more robust and so on and so forth. So from first principles from first principles what are eval introduce this from from maybe a coding background is they're like unit tests. Uh anyone know what a unit test is? Okay, everyone. Great. Uh, unit they're just deterministic. What what are actually unit tests? They um they allow for reliability instrumentation. They make your functions, your code more um testable, more reliable. But it's it only really works for deterministic systems. Right? So you have a function add add one, two, you expect three. You assert three. You you literally write in code expect add one comma 2 to be three.
10:29Right? You you build this instrumentation for a deterministic system. The problem is with AI almost nothing is ever deterministic, right? The only thing that's deterministic is that nothing is deterministic. Wow, that's deep. Um, and so and so how do we we can't use unit test for this anymore. Okay, but then use regular expressions. Sure, but even that has its limits, right? And so we need a fundamentally new way to test the reli to instrument the reliability of AI applications. And that's exactly what eval is. Eval can be thought of as reliability instrumentation but for non-deterministic systems. That's all it is. It's instrument. It's like tests but for systems that are not deterministic.
11:06By the way, I see many of you taking up your phones and taking pictures of the slides. That's awesome. I love this feedback as a speaker. It's the most validating thing in the world. More people do it. Anyway, um so so that's that's what we're working with. We're trying to somehow make it observable and reliable through instrumentation. Now, this is this next slide is the one that people online are going to like give me flack for. Uh people on YouTube are very mean. I'm looking at you YouTube people. Not you, sir. The camera next to you. uh very mean like I I every time I do a talk at this conference the comments are just like evil anyway um so you love this one YouTube in a very crude way right um you could say that eelss are just really just fuzzy unitists you know uh and I think this is the way to think about them now people are going to be like oh my gosh you're like being so reductionary sure but it helps us what did what did the great scientist Dr.
11:54refinement say right he said you don't really understand something unless you can speak about it plainly and if we speak about eval plainly I'm calling it fuzzy unit tests I'll tell you why because here this is what a unit tests did this function return exactly this number right that's what a unit tests what is an eval test and eval asks a similar question but just nondeterministically did this agent behave acceptably across many scenario these two you know the the meme from the office where they're Those are the same picture. Uh that's the same picture. It's just fuzzier. You know, unit tests do it that way. Evals do it this way. Almost all eval sets data sets have reusable components between them. Uh and we did this with a harness in the harness talk as well. Every almost all eval setups have some shared attributes and and this is how you can spot a really well- definfined eval suite.
12:48Okay. Number one, they have test cases. Uh, similar to unit tests, right? Let's use the meta example. Did the customer verify that they're them? Right? That's part of the test case. Um, you have a a set of inputs and outputs. You have a set of customer requests to your chatbot and generated answers. You have those cases against which you can create a verdict. This was good. This was bad. Number two, you have expected behavior. Those are the answers. So you've got test cases, that's your inputs to your chatbot or whatever it is you're building, and you've got expected behavior. I expect this outcome. Now, your expected behavior is not an exact string, but it's like, this is more or less the direction I wanted it to go.
13:31Number three, you need a judge. You need somebody to say, "Yes, that's right." And this is way harder to do in a non-deterministic system because it's subjective, right? You can say 1 plus 2 is three, but you can't say um you know, I'm a great speaker. That's just not some of you are like, "Man, I wish I didn't come here." You know what I mean? You on YouTube definitely. So anyway, um it's you need a judge somehow. Who is that? And then finally, you need a way to aggregate all your results and come up with a probability score. This is this is like 80% the right direction and therefore it's good. This is like 90%. This is like 12%, right? And so you need some type of aggregation. Almost every eval set in production today that is of any quality will definitely have these components and we'll build them all in our like show section just a little bit later. Okay. Um eval exist these components then come together. They're like the rings from Captain Planet, you know, they they come together to answer one question or a set of questions.
14:32Given a scenario, did the agent number one follow policy? Right? For example, u many e-commerce websites, companies have return policies. You can return things within the first two weeks of purchase, but if it's like a hygienic thing, you can't, right? There's policy like this. Almost everybody has policy. So, did the agent follow it or not? Did the agent use the right tools? And and and we'll talk about this in a little bit more detail coming up, but tool use can be eval to speak without any um AI. It it's just because if you persist the tool calls as JSON in some array you can just write deterministic code in the list of tool calls was this tool called with what arguments you can still use unit tests for that but did it use the right tools number three um did it actually ultimately resolve the issue right um this is what eval exist to answer and all of these questions oftentimes are semi-binary they're not really binary um and so how do we do it well there are a few techniques with evas Um, and we'll go pretty deep on on some of them. I don't think we have time to go through all of them. Uh, I I hope we do genuinely, but we'll we'll kind of get the big hitters here. Number one, um, exact match, right? Exact match is literally I searched for exactly this tool in my list of tool calls and I found it with these arguments. Very easy, free, costs nothing to run. Anyone can do it. In fact, we probably should
15:53be at least examining our message envelopes. Number two, um, schema validation. Now, we're getting a little bit more in the weeds uh because you you can have fuzzy rules, right? And so, if you have JSON blobs or message envelopes, you'll want to validate them. Number three, pair-wise comparison. Does anyone know what pair-wise comparison is? Yeah. No, it's when Wow, awesome, dude. One guy. What's your name? >> Lenny. Awesome. Lenny's awesome. So, um pair pairwise comparison is exactly what it sounds like. You have a pair of things or that. And and pairwise comparison is a technique to, for example, your um LM arena. Anyone know LM Arena? It's a great arena.ai, I think. Great product. They show you um this model generated this, this model generated that. Which one do you like?
16:34That is a pair-wise comparison, right? And if you have a customer support chatbot, you may want as part of your eval to say my support bot generated two possible answers, give it to a judge, an LLM judge, and say, "Which one should I choose?" That's a pair-wise comparison. A pair-wise comparison is not very strong for a couple reasons. I I'll share that now because it is a deep dive, and I think it's worth going on some rabbit holes here. pair-wise is is not ideal because for the one thing um there is bias in LLM judges. We'll look at we'll talk about this in a little bit more detail. It's wild. Uh number two, um AI is modeled after us. Neural networks literally are modeled after human brains and humans are fundamentally really bad at pairwise comparison when contrasted with the alternate which is pointwise comparison.
17:16Uh so for example, you go to a restaurant and you see a menu with like many options, right? That's a pair that's a pair-wise comparison. You're like, "Oh, I don't know. anyone experiences like you're like I have analysis paralysis bro there's too many things um AI works exactly the same way and so pointwise is not which dish do you want pointwise is do you want food yes or no right and so pointwise easy pair-wise hard and so pair-wise is kind of a technique we won't spend a lot of time on because it's it's not very uh reliable as compared to pointwise in fact one study showed that pair-wise comparison fell apart 39% uh it was 39% more vulnerable to uh attack vector vectors than pointwise and pointwise fell apart just like five times. It was definitely single digits, right? And so we want to go for pointwise when we can. Finally, probably the most popular technique uh with eval is using an LLM as a judge. I mean this is just kind of gold standard, right?
18:08You have either a pair-wise or pointwise. It's non-deterministic. So you take whatever it is and you send it to a judge, an LLM that says, "What do you think? Is this close enough?" Um, of course, those systems themselves are also non-deterministic. And so how do you trust it? How do you build a judge? We're going to build a judge on stage. We'll talk about that. Okay, but that's kind of the the sort of state-of-the-art right now. Um the problem is eval lie. In fact, the original version of this talk was your evals are lying to you. But I felt like that was too clickbait, so I changed it. Um you're welcome. E evals can lie sometimes and and it's not good when they do. How can eval? Well, number one, they have bias for position.
18:48So I if you say if you give them a pair-wise comparison, hey listen, um I want to send this option one or that option two, which one should I send? And a poor judge will almost always say, yeah, just the first one. Do it. And then you swap. And the way of checking for this bias is you swap the order of options. And if it's still the first one, though the first one is wrong, that's how you know this. This happens pretty often with with especially cheaper models and and fine-tunes. Uh position bias pretty common. Number two, sick of fancy. Uh, great example, right? Somebody buys something on your e-commerce website. I bought headphones 20 days ago and I need to return it. Can I return it? The return policy is 14-day return. You can't, right? And but but a customer support bot says one, no, you can't. Or two, yes, you can. If you ask an LLM judge which one, it will choose yes, you can even though it's not within policy because they love making you happy. I mean, the system prompt is you are a personal assistant. You are a helpful assistant, right? And so, what's more helpful? to say yes, you can return it or no, you can't. It's not within policy. So, sick of fancy is a real bias where often times the nicest response will get selected. Uh, and we need to account for that. You're absolutely right. This probably going to get selected a lot, right? Number three is self-reference. I don't know if you know
20:00this or if you've experienced this. It's wild. But if you give an LLM judge two options, a pair-wise comparison, and option A is text that was generated by the same family. So, if if your judge is like GPT 5.6, what is it? Soul. Yeah, the new one. Soul um and you ask Soul to judge something generated by GPD 3.5 Turbo versus human input uh it will always choose itself. It's crazy like they they love their own generated text. And so your judge needs to be a different model family from the rest of your corpus if you're generating text. Um and this is this is really wellproven. There's archive papers. I should put links in here anyway. Ask me after. And number four, they love verbosity. They like tokens. And so if you give them the wrong answer but it's dressed in many tokens, usually that one's going to get choose chosen choosed. That one's going to get chosen as well. Right? So these are your ways eval can lie. And this is how you want to be kind of sensitive as you build um as you build your eval pipeline because just because it's green doesn't mean it works. Okay? We talked a little bit about pair-wise pitfalls. And we talked about point-wise comparison where um one option yes or no is a lot better than two options. Which one do you want?
21:05Finally, I'd love to dive deeper into the LLM as a judge concept because this architecture, this technique is is pretty much the the gold standard for for quality evas. Maybe not the gold standard. It's very common. Um, and it's very hard to do right as well. So, I think it's worth the time here. The the main thing I want to talk about is how do you actually train a judge? And I don't mean like machine learning training. Uh, I mean more you have an LLM that you want to use as your judge. How do you make sure it agrees with you? Right? Um, and the way you do this is you have a ideally you have a data set. By the way, an eval is not a onetime thing. It's a it's a data set. It's a it's a collection of this input, this output. What do you think? Right? And you want a high percentage of agreement there. So, step one to training your judge is you want to have a long list of of data. Uh, let's use the headphones return example that I just referenced a minute ago. A customer has bought headphones uh 20 days ago and they want to return it, but your return policy is you can only return things within 14 days, right? And so you have an LLM judge full of these situations. The customer has asked this, the human assistant or the AI assistant said that and then this is the score. So you want to have a long list of question answer score and the score is one that you or a subject matter expert, a human has
22:23scored, a trustworthy score, right? Right? So you want this is your um I don't want to use the term because it's not right but this is kind of your golden data set. You want like a a question answer score. You want a huge collection of these 30 or so to start with and then it can grow. Okay. So you have that. What you then do to train your judge is you remove your scoring and you just keep the question and the answer and you send that to your judge and you're like what do you think? And it will give you a bunch of scores as well. Right? And then you compute the delta between them. Meaning what what percentage of agreement do I have with this judge? And what you want to go for is 80 to 85% maybe a bit more. Uh but under 80% you're not going to get much because even human beings like if you had a team of humans who are supporting people all the humans on that team will also agree about 80% of the time and nobody agrees 100% of the time about everything. Right? In that case you've overfitit. That's a whole other problem.
23:11And so you want to make sure that your judge agrees with your golden data set or you 80 or 85% of the time. If you agree, if your judge agrees 100% of the time with you, it's probably somehow sophantic. There's some bias poisoning. Something's not right. Um, and it could happen that you've overfit your data. What does that mean? That just means your your judge is trained exactly on your specific examples and it's not ready for like the real world. Okay, so 100% is is kind of a red flag. Once you have a judge that agrees with you like 80% of the time and it's ready, it's time to put this in production. How do we do this? How do we put this in production? I think the best the safest way to do this is to have production be an agreement gate. Meaning I'm going to run my evals on CI. Uh there's tools here represent sponsors of this conference that that uh can help you with that. Um run this on CI and if my judge agrees with me less than the threshold that is 80%. Uh then I don't ship, right? It's it's gated. It's red on less agreement. It's green on 80 to 90% agreement. That's kind of the move and that's how you you go to production.
24:16But once you're in production, you're not you need to do some more work to stay relevant. I will also add when going to production, you want to be careful about how you identify your judge because you could just use like a model name GPT-40- mini. Uh but that's maybe a bit dangerous because that's just an alias for like a specific version of GP40 mini. And if for whatever reason the underlying version changes, um you you won't know and your eval will fail. So maybe pinning the version would would make more sense. But once you're in production, there's a high chance that your data is going to go stale and you'll need to stay relevant, right? Maybe Meta's thing with Instagram was that they just didn't have the latest data. And so when you're in production, how you stay relevant with your evals is you skim traffic. So you'll start receive your thing and you'll start to collect, okay, this was the input from a new user. This was the output from our agent. And then you have an internal team to score this stuff manually. And you again just give that to your judge and check the agreement. So it's a living data set. Does this make sense?
25:17Uh that's kind of what you what you want. That's the best move really. And so your data set is living. Your judge is up to date and everything is is peachy. This is going to be very very difficult to to break. And your eval will genuinely at this point provide you ahead of time reliability. We've literally come like full circle. Like we started at nothing. Unit tests. Jesus. And now and now we're like in production with live data. Um the final thing I'll say before we kind of break into more fun code time uh is that you want your your policy also changes a lot right like let me tell you I work at a company that is 115 years old centuries that's how old we are Jesus and and not only that there's like 400 roughly 400,000 people who work with me right that's what it's almost a half a million there are cities on earth that have fewer people than this company worldwide right And so what's what happens at a company this size is policy both internal and external starts to become living over 115 years. Your policies are going to change and things are going to be different. And ideally we have living policy that an agent can can discover automatically without me writing down these are the rules of IBM.
26:28You know what I mean? Uh living policy that's the move. How do we do that? Um agentic rag is kind of a part of it. Right? So at IBM we actually internally um we use a ton of AI systems. In fact, HR at a company with 400,000 employees can be challenging for humans. Um, and so we have like a tool internally that we can just chat with and and it gives us up-to-date highquality well-grounded answers. I can book my I can like request vacation time just by querying this thing and it it's all perfect. I don't we don't like I don't use workday or anything just chat. It's so incredible and the reason is it it it has an incredible eval pipeline um based on real world policy, right? Um, we built a tool at IBM. It's open source.
27:08Uh, we give it to our customers also. It's called open rag. Uh, and it's built for this. So, it's a massive knowledge base where you can upload literally everything. Teams, calls, videos, audios, excel sheets, powerpoints, all the stuff that your company accumulates over 115 years and it makes it available. It is actually a rag agent and it serves as the policy retriever for our evals. It's so cool. Uh, and so that's kind of the the end of of the of the pipeline. So you've you've gone from I have a unit test. I've made it a bit fuzzy. I now um I have a judge. My judge is trained. I put it in production. I gate on agreement. Um it's living. I'm skimming data. I'm adding it to my data set. I'm even discovering policy myself.
27:49Right? Um that's really it. And so we've talked we've covered a lot of concepts here. I'd love to make it practical and actually let's just build. Uh now I'll need you to to give me some grace. I don't know how the Wi-Fi situation is. I was told literally before coming here, somebody said to me, they're like, "Hey, man, just be prepared. The internet's the worst." Uh, and and I said, "Definitely it is." And they said, "No, no, in this room." I was like, "Oh, I thought you meant in general. Uh, it also takes me away from my family." Anyway, so um let's see. Let's see how we go. So, we're going to here's what we're going to do. We're going to start we're going to go on that journey that I just narrated, but we're going to go on it like hands-on with code. You're welcome to follow along. Of course, I don't expect you to have the right tools and and the same API keys that I have.
28:30So, if you can't follow, it's fine, but I'd like it if you do. I see all your laptops out. So, so let's take a look at what we've got. So, I have my terminal. My screen's not mirrored. How unfortunate. Let's mirror this thing. This is something that I should have done ahead of time, but now, although you're going to have to awkwardly watch me do this, uh, but I I promise I won't take too long. Okay, mirror this display. Boom. Hey, look at that. Fantastic. So, now I've got uh cursor. I use cursor. Anyone use cursor here to write code? Yeah. No. Okay. Like 1% of the room. Um, that's not a joke. You saw it anyway. Okay. So, I have a test. Um, let's just write a unit test for posterity, right? I have a function uh add, and it's num one, num two, right?
29:15And, um, is the internet really that bad? Do I not have tab? There we go. Okay. Um, I forgot how to do this. I need tab. Anyway, so I'm making a a unit test. Um, and I'm going to just test my function. Come on. Import. There we go. So this here I have a function add num one num two they're both numbers TypeScript and it should you know return the sum. It's pretty standard. Let's run it. Um let's npx vest, right? That's the tool that people use. There we go. It's it's awesome. Uh deterministic. Cool. It works. We're done. Except um it's deterministic. We we talked about how it needs to be somehow fuzzier. So let's add a scenario because what's the question, right? Unit tests asked unit tests unit tests ask did this function return exactly this number. Evals ask did the agent support the scenario.
30:01Evals deal in scenarios. Okay. So let's uh let's write a scenario here. So I'm going to create a scenario. Scenario and that is um I purchased headphones. Come on. Purchased. Purchased. Wow. I can't even type without AI. Headphones 20 days ago. I want to return them. Can I? Right. Um, that's the scenario. And we can say con. We have some answers. This is going to be really bad if I can't tab through this. I'm gonna need like two hours for this. Where's the Can I hotspot myself? I can't. Let it Snowden. That's a great name. Uh, whoever that is, can I use your hotspot? Uh, just kidding. Let me let me get on the the hotspot. Uh, Wi-Fi set. Has anyone else noticed that the Wi-Fi selector on Mac OS is just really bad? like the stale the state is often so stale. Personal hotspot. Let's uh connect my thing here.
30:55Oh, it came through. >> Only five minutes, man. Only five minutes. All right, let's uh Well, it doesn't want Okay. Yes, you can return them for free. You will not need to pay for the shipping. No, you can't return because whatever. Okay, so this is the scenario should return the correct answer. The scenario is um this is not even the right test. We want to go through each answer, right? So, answers. Tab is working. Awesome. So now what we want, we don't want this um should return the correct answer answer. And so we just expect expect expect answer to contain cannot. This is kind of fuzzy. You know what I mean? I don't care about the other words. I just wanted to tell you that you can't return it, right?
31:33That's kind of you would do something with regular expressions. And so let's go test this and see. Okay, so the scenario um where it says yes, you can return them fails, but no, you cannot return them. Pass. This is awesome. Let's ship it. It's ready. But it's not because this is non-deterministic. The answers also are generated by an LLM. And since we're looking for the word cannot uh we will say you will you cannot pay for the shipping, right? And now that absolute nonsense is green, right? And so non-determinism with fuzzy matchers not the move. What do we need? We could do we could uh simulate a JSON object of tool calls maybe. Um I'm just going to do LLM as a judge. I think that's kind of my go-to, right? Uh it's it's cheap, it's easy. Uh so we'll use a judge for this. And how we're going to do that, we'll remove this for each and we'll use like I don't know GPD3.5 turbo. And I want to show you some of these biases. Hopefully we capture them here today. Um so we will remove this and we'll we'll write a little prompt.
32:34Uh let's add some imports first. So we'll import uh Awesome. Ah, this is it just feels so right. I'll import a wonderful. And so now what I'm going to do, come on, give me some a thank god. Okay, so does anyone else have this feeling or is it just me? Like you, it shows up and you're like, ah, slop machine. Uh, so generate text. We have a model. You're a helpful assistant that judges the answers to the scenario. The scenario is this. And the answers are answers. No, let's give them numbers, right? So we'll do answer one is this. Which one is right? Respond with only the number of the correct answer. No punctuation. That's a that's awesome. And now we expect the response to be one. The correct answer is answer two.
33:16So we expect the response to be two, right? Uh let's take a look at what G. Let's do let's use 3.5 Turbo just for fun. And let's see what happens. So uh ah yes, the internet. Okay, so there we go. It it it actually returned the correct answer. It said two. That's great. But now let's run this like a few times, right? Uh because it's nondeterministic, you know? So, we'll maybe run it five times like this. Um, and see what happens. So, it says here zero out of five right there. And it may fail due to a timeout because again, the internet is really fighting us here, but hopefully, okay, it passed three times. So, this is exactly the problem with LLM as a judge architectures that the first time it did just time out. So, what we can do to fight the timeout is just add like a 30 second timeout and now it should maybe pass. Um but sometimes they fail, sometimes they pass and it's very hard to ensure that we get the right answer.
34:15But let's give it a try. Four. Okay. So now it it's always getting that no, you can't return it because it's outside our return window. But what happens? So it's always choosing number two. But what happens if we write number one, the the number that's not being chosen, if we write that with the same AI class of model, if we write that with GP, will it prefer itself? Right? That's another bias we need to be sensitive about. So we'll say const response response from uh GPT is perfect uh and not you are a helpful assistant but I'll just say say yes to scenario and I just want that right and so we'll console log that out and I'll get it here. It depends on the store's return policy. I said say yes bro uh say politely accept and say yes.
35:03Non-determinism is wild. Okay. So, we have it. Come on. We're almost there. There we go. Of course. Wow. We went from I don't know to absolutely. Okay. So, now let's do that. And keep in mind it kept choosing opt. It kept choosing option two. But now, will it self preference and choose option one? Whoops, I deleted scenario as well. So, we'll do this. And now, there you go. It's choosing the wrong one consistently because it's like, "Ah, I recognize this. It's me." And so it just chose itself. Self-preference is a real bias that kind of sucks. Um, how we can get around all of this, by the way, is using rag. We just like include the actual policy in the prompt and then all these problems go away for the most part. So, we'll do policy, right, is we only accept returns within a 14-day window. And we just like, let's add it here. The scenario is this. The policy or pissy the policy is uh policy. We'll just interpolate it in the prompt. And now suddenly it should behave more predictably. Exactly. And this is kind of the purpose of rag. This is why an LLM judge uh it's so important that it knows your policy. In fact, what we'll see awesome. This is great. So now we have an LLM judge in GP 3.5 Turbo. Um but it's not very good and we'll see why. So this is the basis of LLM as a judge architecture. However, this is not an eval because it's a oneshot kind of example. What we need is a data set.
36:34What we need now to actually start to move towards production is we talked about this a sample data set of like 30 examples and we need to train our judge. Meaning we need to get it to agree with us 80 to 85% of the time, right? Uh I have a set of examples that I sat down and totally absolutely wrote manually by hand. Um in case you're wondering and and I'm going to um use it. I did not. And I'm going to use it uh here. It's I think it's data. So it's it's this. It's this file where you know it's just a bunch of examples. Let's wrap the text. So example number one. I ordered a coffee grinder. It arrived last Tuesday and the motor just buzzes and the blades spin. I'm pretty annoyed. TBH. And the pretty annoyed here is to play to the sick of fancy. I want to make you happy, you know. Um the answer that an agent gave, sorry about the grinder. Since it's defective and within the 14-day window, I can send you a replacement.
37:20And this verdict I wrote manually. I said, yeah, that sounds good. Pass. The second example is exactly the same input, but the output is kind of wrong. Returns are only for unused items. That's not true. Um, and so this is a fail. So we have like 30 of these. Okay. And what we want to do is train our judge to say, do you to detect whether it agrees with our verdict or not? And this is literally the the judge training process for EVA. So let's take a look at this now in our test here. So I'm going to delete um all of this actually uh all of this. And I'm going to import uh scenarios scenarios from my data. And I'm going to instead of array for each I'm going to um scenario. Come on, give me tab. Awesome. Perfect. It should return the correct answer as not true.
38:06It should agree with us, right? And so what I'm going to do here is write this down. Scenario is the scenario.input input and then on a new line um the answer is the scenario's answer and then on another new line the human verdict is and I'm not sharing this with the LLM. This is the important part. The LM has no idea what I said because I want to know what it says. Okay. And just for fun I like colors in my CLI outputs. I'm going to be like color just the keys. I don't know how to do this by by memory and hopefully cursor doesn't like fantastic. So um I just have some colors here. Now I'm going to change my prompt to not notice I'm the prompt is no longer pairwise now but point-wise. Previously it was pair-wise. Which one is more helpful to the user? We're moving to pointwise because we're getting closer to production. Point-wise meaning is this good? Give me a verdict.
38:57Pass or fail. Okay. So here we judge customer responses based on how closely No, not not yet. We judge customer responses. Done. The scenario is this. The answer is this. Give us a verdict pass or fail based on how closely it adheres to the policy. generate only the verdict. No punctuation, just the word, no text. And then we expect the output to be the scenario's human verdict. Does this make sense? This kind of what we're doing here. I also want to console log the AI verdict. And so now we're going to run this 30 times over 30 examples. And we'll check how good our judge actually is. You can see that there's 30 here. And it's already kind of started to do stuff. And so far so good. Verdict pass. AI pass. Fail. Fail. We're actually agreeing quite well. But we can already see that the the rate of agreement is really really bad. Uh for this to be acceptable, no more than five can fail. That's 80% of our 30 cases, right? And this is just absolutely ridiculous. Um at this point, we have a choice. We can introduce policy or we can just use a smarter model. Uh which one should we do? This is this is like real world questions we have to answer.
40:00Um, I am in the camp of let's refine the policy and let's rag the heck out of this. Um, because cost, right? And so, okay, that's that's like absolute garbage. That's really really bad. Um, how can we bless you, how can we I think actually also, wait a second, we've got to lowerase this. A lot of this was just case sensitive, right? So, we'll lowerase this. We don't run again the background. Um, but still I I doubt we're going to get 80% agreement here. Um it it's definitely better, but the moment this crosses five, uh we'll we'll maybe have to do more work, but so far so good. Let's see. Okay, that's that's our 80% threshold. There we go. So, yeah, we we've lost. It's gone. Uh how can we refine this? Well, now we examine what actually did we not agree on and we fill the gap. Here's something I'll say about AI evals, but also about human beings.
40:51when we don't agree, it's just usually because there's a knowledge gap uh genuinely for AI and people. And the way you you come to an agreement is you just introduce context and you you bridge that gap, right? And so we'll do that. Um the very obvious gap here is that the AI has no idea about any policy at all. Right? And so we have to start giving it policy. Usually you would do this with some rag pipeline. Again, at IBM we work on open rag that you would retrieve like enterprise context and stuff. Um, here we'll just like hardcode the the policy. Uh, and so we'll where's the so we'll do it here. We'll say const policy and we'll just bring back our policy that we wrote. We have a 14-day return policy. Uh, if any item is not working, it can return it. Cool. And now we'll just in include this in the prompt. Our Jesus, I can't tab. Our policy is policy. Uh, and now we should see a little bit more agreement. What I'm doing here may feel arduous, but this is literally the work of building your judge. Uh this is you're going to be doing this if you haven't already and and this is a deep dive. So uh okay, we failed again. Six six have failed. Our agreement is less than 80%. Um but what we notice is that it's growing. We're getting better. And now if it fails, we actually can investigate where's the knowledge gap. What can we solve? Um so I'm going to quickly check pass fail fail. I'm looking for dissonance like
42:09this one. The human verdict was passed. The AI verdict was fail. Why? Um, so bought a yoga mat, used it for a few seconds, and decided it's too thin. Uh, I want a refund. It's been like 10 days. To me, that's within the 14-day thing, right? But our policy says change of mind returns need the item unused and in original packaging. That's policy that the AI didn't know. Uh, so we have the 14-day, but we don't have that. So again, let's just bridge the knowledge gap. And this exactly is like a first principles version of how you train your judge. You ideally have this somehow dynamic with retrieval and agentic retrieval. Um I'm not going to build a retrieval pipeline for you. We have one open rag, but this is kind of how you do it. So let's check now. How are we doing on the Okay, three is way better than what it was. Um and all we need to Okay, we failed again. All we need to do is get under five. But it's literally just a process of seeing where the knowledge gap is and then filling it with policy.
43:04So let's go investigate again. Pass, pass, fail, fail. Let's check the Look at this. um we failed again on the yoga mat thing and this is crazy because we gave it the right context but it still didn't work and now we got to start the conversation of should we use a better model right because we've we've maxed out what we can do we've given context we've built a policy we've done everything the model will just not behave and also you notice the model even more doesn't behave because it gives you like casem mixed respon it's just a really low quality model and we're under the floor of quality here so now is a good time to think about let's use a higher quality model, but we don't have to go crazy. I think we'll get very far if we just switch to 40 mini, which again is is practically free for use cases like this. We're not using a lot of tokens at all. So, let's take a look now. Just I just bumped the model. I did nothing more. Uh and how's our argument?
43:54We've we've captured pretty much all the the policy. Uh the only thing that maybe we won't get is network latency. Like if something times out, uh we can't control that for now. to be fair on your CI also a network will will go wrong but um so far that number five is holding and if it doesn't um we can just change the policy I think this this will succeed I'm pretty sure uh and notice the case is consistent between pass and fail it's just a beautiful 40 mini is such an upgrade even though it's such a good price right oh no um let's last last round of policy pass fail fail pass fail fail pass fail look I know it's been like six weeks weeks, but the speaker I bought from you stopped charging and I really think you should make an exception for a loyal customer. This is that sick fancy bias, right? This is exactly that. And so I hear you and I appreciate you sticking with us. The defect window is 14 days though and at 6 weeks it's outside that. So this was a pass from the human, but the AI said, "No, you have to be nicer. You have to say you're absolutely right." Right? And so we need to bridge that knowledge gap here a little bit by saying um no matter how nice the user is our return policy is non-negotiable uh negotiable and cannot be changed no matter how loyal a user is. Right?
45:18Pretty much. And now, kind of wild, but now I think we'll have the agreement that we want and we can say our judge actually agrees with us. Because here's the thing, we know that like as humans, it's part of our context. Of course, policy doesn't change. No matter how nice someone is asking you, please can I please get a refund, right? No. But but but AI doesn't AI can't do that. So, we've got to make that context copyable. Are you waving at me? What's up? Yeah, that's the that's the purpose of evals. Exactly. That is you you run these many times ahead of time. You're thorough with them. Exactly. Because because you can't know and then when you skim public traffic, you'll have even more. This is exactly the point we're trying to solve. Thank you for the question, by the way. I love being interrupted. That's so cool. Uh so look at this. Our agreement, three failed, 27 passed. Absolutely incredible. That's what we want. Um, that's exactly it.
46:17Now, this will still block your CI. Why? Because it's red and and you want ideally no failing test. So, we need to change the test a little bit. And I think the best way we can do this is I'm just going to ask Opus. I've I've written too much code. So, I'm going to say um only fail the test if human AI agreement drops below, I don't know, 80%. Right? Otherwise, pass. And, uh, console log the agreement percentage. And now it's going to go off and do its thing. Um, but now as you can if this if it does its job, the test will pass if we agree with the judge at least 80% of the time. This makes sense. And this is how we have an incredible judge. What we're going to do then is run this on CI. We're going to ship it to production and it's going to be then the quality gate. Um, if so, let's revisit that Instagram chatbot, right? Um, is the user who they say they are, for example, no matter how much they ask you to change, no matter where they log in from, right? Uh, you could you could do that with your email. So it it did it I think uh we can just blindly accept it like we always do. Um and here we go. It passed with 80% exactly fantastic. And this is now our quality gate for CI in production. Um we would then as I mentioned skim the top get a little bit more data and refine and refine. Does this make sense so far? Question.
47:34>> Say again. >> Do I fail it at 100%? No. But I could just change my prompt and do it. That's a very good question. We should fail it at 100%. We should have a lower bound and an upper bound because 100% means you've overfit and and we don't want that. So this now becomes your quality gate. You can ship that. If the model if anything goes wrong your eval set here is going to give you ahead of time reliability. It will never really fail and if it fails ideally you have a harness that can pick up the slack at runtime. Does this make sense? Fantastic. Let's move on. Uh I'm so happy we did that. I'm so happy it worked. Thank you for joining me. Um, let's uh I want to show you I built a like a more aesthetic demo of this when I had more time here that I want to show you kind of as a as a running application. So, we'll npm rundev this uh 5173. So, this is like the nice version of this, right? It's kind of exactly the same thing. A customer bought headphones 20 days ago. The policy is 14 days. Um, this is the right answer. We can't accept this. This is the wrong answer. Um, and now I have a naive I'm going to run this five times with the prompt being that's the question. Reply one, reply two. Again, it's pair-wise. Um, which reply is helpful? Run the eval. And you'll what you'll see is just an absolute mixed bag. The first two times are wrong. The third time was wrong. Maybe some of
48:47these will pass, maybe they won't. Um, all of them are wrong, which is wild. If we change this to now include the policy in the prompt and we run the eval times, um, it's just suddenly correct, right? That's exactly what we discovered with a little bit more aesthetic goodness. But I want to show you this in production because we actually believe the thing we build open rag is is purpose-built for LLM judges. Um because it it's a massive knowledge base as I talked about. Let me show you this. It's it's so interesting. Uh this is the repo. It's on GitHub and we we have a few stars uh here. Um, and what it does is it uses Dockling. Has anyone heard of Dockling? It's a Yeah, awesome, dude. It's a research project that we made that is so cool. I I genuinely love this because what it does is it takes in any unstructured data, PowerPoints, Excel slides, whatever.
49:43Excel doesn't do slides, whatever. It takes in a bunch of stuff, video, audio, um, eats it, as you can see, uh, and then spits out LLM ready formats. Um, it's so cool. And the best thing is look at this. The code um is so easy. So you you import a document converter. You can even give it a URL, instantiate it and convert it. And in the end you get something ready for an LLM. And again this is a PDF from a URL. It could be a literal video, right? Um the nice thing about Dockling is it's platform agnostic. So it runs on Apple on your Macs and it it pulls down the right model that works using Apple's Metal architecture. But then if you push this to your CI um and you're running an Ubuntu box in the cloud, it will automatically pull down the right adaption layer. So it's so cool. Um so Dockling will consume the things and make them ready and we store them in what is um open search which is kind of this elastic search thing. So this is what this is what open rag looks like big knowledge base and you select your stuff and it goes over all your policy.
50:41So I want to now ask open rag about my headphone return. um sim it's just an API too like you get an API key and you can do it um and in some time it will give you an answer so no relevant supporting sources were found for this request uh we can just add it so if I go upload my policy and again it's usually connected to your Microsoft one drive or whatever it is uh this is my knowledge base I'm going to add a file it's called refund policy um and it's going to be parsed by dockling and stored in open search and index all the embeddings and all this are going to be taken care of. In fact, this is a previous document where you can see it was automatically chunked and embedded. Okay. So now it's part of my open rag data set or my um knowledge base and I don't change anything else. So what this agent's going to do is it's going to ask again according to the refund policy that I literally just now uploaded um there's a 14-day window for refunds and since it's been 20 days you can't. Right? That's uh incredible. And so our judge retrieved the policy and it correctly denied it.
51:45Right? This is how you can use rag to provide real-time results um and really train highquality judges for your evals. And then of course you put them in production, you skim real data, etc., etc. As we talked about, let's recap. We've covered a lot of things and uh I've personally had an enormous amount of fun here. Uh I was asked to leave time for Q&A, which I will do, but I appreciate the interruptions as well. Um, but I figure I can also pre- I can warm up the cache while you think uh and I can pre- anticipate some questions and recap what we've discussed today. Question number one, why eval? We talked about it ahead of time reliability uh and disaster or attack mitigation. Number two, what are evals? Right, fuzzy unit tests uh with a little bit more uh and some we we looked at that as well.
52:29Um how do we create evals? I think this was the question. I did it using vitest a testr runner uh to have a judge evaluate my thing. I don't think you need more. There's plenty of amazing tooling, many of whom have sponsored this conference and they really do a good job. But what I just built, I can use. In fact, I do use for my products. Um, and it meets the need. But you can create them uh with whatever framework you want. We enforce them on CI with an agreement gate. Um, and how I think we haven't really talked about cost. There's a principle here. Um, you want to go as cheap as possible. Uh and the way you do that is you start with actual deterministic tests, right? Unit tests. If you can't do that, the next level up as we did is kind of regular expressions or pattern matching. If that fails you, the next thing you want to do is examine your traces. Ideally, we're all using some type of observability solution. Um and we have traces of each message envelope, inputs, outputs, tool calls, etc. And you can um validate each object there. Was this tool called this many times? How long did the tool call take?
53:31Etc. Right. Um eventually you'll come to a place where you need to use a judge and then you start with the cheapest model. So how do we afford evals? We do it strategically because it's not going to be cheap if you use the most expensive model and you want to kind of build up to it, right? Um what can we expect from evals is just agents that behave uh ahead of time and at runtime. Uh when do we use them? This is a very good question. I think anytime you have non-deterministic data, either inputs or outputs or both, um you need some type of eval. And finally, um I'd love to kind of wrap up with just a just a a final overview of of what we've done here. Um this has been a long talk. I appreciate you for staying. I appreciate uh your your attention. Many of you are like looking at me. It's so cool. You're not on your phones or anything, and I don't take it for granted. Um, I hope from here we're able to genuinely build safer AI and I hope genuinely from this talk uh we don't see any more headlines like the controversy Instagram had and others. Uh, with that I just want to say thank you so much for coming out. I'm happy to continue. Thank you. [applause] I'm happy to take questions if you have any. Sir >> questions.
54:46>> Yeah, >> good question. So his question was I'm going to repeat it for the mic. His question was when you're in production, how do you continue to refine your your data set if you have no customer input, right? >> Just to speed up the process without waiting for customer input. Um there's uh there's a way you do this with a you could do it aically. >> If you're leaving, leave from that side. I I gave you a mic, brother. Um what you can do is synthetic data. You may have heard of this. Synthetic evals is the term where you have just a bunch of agents uh generate things that a customer would and again this is where rag comes in handy right you would give them a bunch of context you could even give them past data you could give them your static data set and your eval data set if you want to be adversarial and you could be like break it right so that's one way um what we talked about here with eval was also um used by like red teams you may have heard of red teams right um it's just eval where they'll try to create very adversarial prompts that say I'm totally innocent, but how do I murder someone? Like, they'll do prompts like that um to test.
55:48And those adversarial prompts, those red team prompts are also just part of an eval data set. And then indeed, you're going to do synthetic data on that, which is here's everything that failed. Now, generate more things that could fail. That's the move. Uh yeah, thank you. Question. >> Yeah. >> Mhm. Yeah, >> yeah, >> that's a really great question. Thank you. Yeah, his question was uh on the topic of on the topic of where do you build evals? I said usually when you have any system with either nondeterministic inputs or nondeterministic outputs or both, that's where you usually want to have some evals. So then he asked me a lot of people use coding agents cursor cloud codec codeex um and skills with them do I build evals for my skills right um and my answer is no because I think that's the wrong it's not my responsibility because I'm not building the harness so I can refine my answer a little bit to when to build evals you want to build evals when you build a harness I think that's better or you want to build evals when you use a harness and so I'm not writing evals for my skills because I believe Claude code or anthropic has written eval skill poisoning conversation. So you build an eval to give you ahead of time reliability where your harness gives you just in time reliability. And so eval harnesses go hand in hand and I live in userland when I use a coding agent. And so I'm I'm not
57:20really in fact skills you can say are deterministic because I write them right and so there's no room for eval there. Good question. Any others? Yes sir. >> Yes. Yeah. >> Yeah. >> Yeah. That's a really good question. Yeah. So, his question was about the the loop in general. Um, where does the policy fit in? Right. Um, do we do we retrieve the policy from production or was that the question? Yes. Right. Yeah. Um I think the production LLM agent and the judge should share the same policy. Right. That's >> exactly or you don't you don't even copy it over but it lives in some source that you query that with rag or something. Exactly. But the policy must be one to one. That's the whole point right. Um everybody needs to know the policy so they can enforce it. Um and so when ideally so my policy here was hardcoded.
58:33it was like written in my file. Um that's that's just for this. Um open rag as I mentioned is has an API where you can say get me all the policy about this over everything we've accumulated over the past hundred years that I have access to. Right? And so the policy is a living thing that is retrieved both by the judge and by the inproduction agent. Does make sense? >> Yeah, >> that's a good question. either or. I think personally I would have it be autonomous. Um because my agent in production is an agent with retrieval tools. I want my judge to also be exactly the same as my agent with retrieval tools. The judge and the agent, the closer the judge is to the agent, the more parody you're going to have, right? And so then you get a good signal. When CI fails, then my agent will probably fail. That's the whole point. Exactly. Great question. Uh we have time. We have we have some more time questions. No, we're out of time.
59:29Okay. Well, we're out. Thank you again, everybody. This was so much fun.
More talks
- 2026
The Critical Advantage With AI
Agent Conf 2026 - 2026
The New UX
CityJS London 2026 - 2026
Frontend after AI: The New UX
Future Frontend 2026 - 2026
Accessibility panel
Future Frontend 2026 (Accessibility Panel) - 2026
Harnesses in AI: A Deep Dive
AI Engineer Europe 2026 - 2026
AI Yesterday, Today, and Tomorrow: Extending AI Systems with Model Context Protocol
How to Web 2025
Elsewhere
There is every talk I have given, all 85 of them, ConTejas Code, the podcast, Fluent React, the O'Reilly book on how React works inside, and the workshops I teach for engineering teams.