← Back
AI Engineer October 6, 2026 17m

Why Building an Eval Platform Is Harder Than It Looks — Braintrust

Read full transcript 15 segments
  1. Thank you for joining Thank you for joining this session on this session on this session on why it's difficult why it's difficult why it's difficult to build to build to build quality agent platforms. My quality agent platforms. My quality agent platforms. My name is Hussein. I name is Hussein. I name is Hussein. I lead the lead the lead the solution development organization at the solution development organization at the solution development organization at the BrainTrust event. Hmm, I BrainTrust event. Hmm, I BrainTrust event. Hmm, I spent 15 years spent 15 years spent 15 years developing solutions between developing solutions between developing solutions between Salesforce and Databricks. I've Salesforce and Databricks. I've Salesforce and Databricks. I've been a technologist my whole life. been a technologist my whole life. been a technologist my whole life. I've been building computers since I was I've been building computers since I was I've been building computers since I was 5 years old. So I have an 5 years old. So I have an 5 years old. So I have an insatiable appetite for insatiable appetite for insatiable appetite for technology, which technology, which technology, which puts me at the puts me at the puts me at the forefront of forefront of forefront of artificial intelligence artificial intelligence artificial intelligence here at BrainTrust. For those here at BrainTrust. For those here at BrainTrust. For those of you, I know of you, I know of you, I know some of you were in some of you were in some of you were in the room earlier when the room earlier when the room earlier when Jess was doing her Jess was doing her Jess was doing her session. I saw several session. I saw several session. I saw several hands dedicated to hands dedicated to hands dedicated to evaluations. How many of evaluations. How many of evaluations. How many of you know what you know what you know what BrainTrust is and what we BrainTrust is and what we BrainTrust is and what we do as a do as a do as a company? Just company? Just company? Just raise your hands. raise your hands. raise your hands. Perfectly. I see one Perfectly. I see one Perfectly. I see one or two. So, at or two. So, at or two. So, at first glance, BrainTrust first glance, BrainTrust first glance, BrainTrust as a platform is as a platform is focused on the focused on the quality of agents. And the quality of agents. And the quality of agents. And the main idea is to main idea is to main idea is to help help help teams build teams build teams build and maintain and maintain and maintain confidence in the confidence in the confidence in the AI ​​functions or AI ​​functions or AI ​​functions or AI agents they AI agents they AI agents they deliver. And there are two deliver. And there are two deliver. And there are two key pillars of key pillars of key pillars of agent quality.

  2. agent quality. The first is the assessments The first is the assessments we talked about we talked about we talked about earlier. Evaluation ( earlier. Evaluation ( Evals) is what you Evals) is what you Evals) is what you do before your do before your do before your agent goes into agent goes into agent goes into production. This is production. This is production. This is when teams when teams when teams experiment, experiment, experiment, test behaviors, test behaviors, test behaviors, and build and build and build the confidence they the confidence they the confidence they want to have before want to have before want to have before launching the agent into launching the agent into launching the agent into production. The second production. The second production. The second pillar is pillar is pillar is observability. observability. observability. Observability Observability Observability occurs when your occurs when your occurs when your agent is in agent is in agent is in production and production and production and interacting with interacting with interacting with real real real users. This users. This users. This creates interactions. creates interactions. The goal is to The goal is to take the hypothesis take the hypothesis take the hypothesis you built in the you built in the you built in the offline offline development process and then development process and then development process and then reinforce it in reinforce it in reinforce it in your production your production your production scenario scenario scenario through continuous through continuous through continuous monitoring. monitoring. Evaluation and Evaluation and observability are observability are observability are closely related to the same closely related to the same problem. One problem. One happens in happens in happens in production, the other in production, the other in production, the other in development, and both development, and both are about understanding how to are about understanding how to are about understanding how to improve the quality of the improve the quality of the improve the quality of the agent. Now I won't agent. Now I won't agent. Now I won't waste any waste any waste any more time here.

  3. more time here. more time here. Let me Let me Let me finish this. Let's finish this. Let's finish this. Let's talk about talk about talk about why we're here, why are why we're here, why are evaluations (evals) evaluations (evals) important? You know, important? You know, important? You know, grading grading grading matters because LLMs matters because LLMs matters because LLMs are inherently are inherently are inherently non-deterministic. non-deterministic. non-deterministic. They are very changeable. They are very changeable. It is this variability that It is this variability that makes them so makes them so makes them so powerful. They powerful. They powerful. They can reason in can reason in can reason in many different many different many different areas. They can areas. They can areas. They can solve different solve different solve different types of problems and can types of problems and can types of problems and can handle a wide handle a wide handle a wide range of range of range of user needs. But that same user needs. But that same user needs. But that same flexibility also flexibility also flexibility also creates risk. Agents creates risk. Agents creates risk. Agents use LLM as use LLM as use LLM as their brain. And the their brain. And the their brain. And the agent experience agent experience agent experience is starting to become the is starting to become the is starting to become the primary way primary way users interact with users interact with companies. Therefore, companies. Therefore, companies. Therefore, teams need teams need teams need confidence in their confidence in their confidence in their agents that they agents that they agents that they will behave will behave will behave reliably and achieve the reliably and achieve the reliably and achieve the expected expected expected results. And without results. And without results. And without evaluations, companies evaluations, companies evaluations, companies face face face real risks.

  4. real risks. Brands become Brands become a risk if there is a risk if there is a risk if there is inconsistent inconsistent inconsistent behavior. behavior. Compliance if the Compliance if the agent says or does agent says or does agent says or does something wrong. And something wrong. And something wrong. And there is a cost and there is a cost and there is a cost and maintenance risk if maintenance risk if maintenance risk if systems are difficult or systems are difficult or systems are difficult or unreliable unreliable unreliable to set up or to set up or to set up or monitor. The goal monitor. The goal monitor. The goal of assessments is to reduce this of assessments is to reduce this of assessments is to reduce this uncertainty before uncertainty before uncertainty before launch so that customers launch so that customers launch so that customers have a good experience have a good experience have a good experience and agents behave as and agents behave as and agents behave as you expect. you expect. So you might be So you might be thinking, "Okay, Evals thinking, "Okay, Evals is just a is just a is just a user interface in a user interface in a user interface in a spreadsheet, spreadsheet, spreadsheet, right?" And I guess right?" And I guess right?" And I guess I should ask I should ask I should ask the audience. Just the audience. Just the audience. Just raise your hands, raise your hands, raise your hands, how many of you have how many of you have how many of you have run assessments and run assessments and run assessments and recorded recorded recorded the results in a the results in a the results in a spreadsheet? Has spreadsheet? Has spreadsheet? Has this ever happened? this ever happened? this ever happened? Okay, I see one or Okay, I see one or Okay, I see one or two cases. So, by two cases. So, by two cases. So, by the way, there is nothing shameful about this the way, there is nothing shameful about this the way, there is nothing shameful about this . This is . This is . This is actually a good actually a good actually a good first step. It is important for first step. It is important for first step. It is important for teams to recognize teams to recognize teams to recognize that this is a real that this is a real that this is a real problem and they problem and they problem and they need a way need a way need a way to understand how agents to understand how agents to understand how agents behave based on behave based on behave based on different inputs.

  5. different inputs. At the simplest At the simplest level, an assessment system level, an assessment system level, an assessment system tests or has three tests or has three tests or has three different types of criteria. different types of criteria. It has a way It has a way to launch agents to launch agents to launch agents based on test based on test based on test input data. input data. Second, a way Second, a way Second, a way to view to view to view results or results or results or grades, even if grades, even if grades, even if they are recorded in a they are recorded in a they are recorded in a spreadsheet. And spreadsheet. And spreadsheet. And third, a set of third, a set of third, a set of test test test inputs or examples inputs or examples inputs or examples that can trigger the that can trigger the that can trigger the agent. An example of agent. An example of agent. An example of input data is input data is input data is any information any information any information needed to needed to needed to run an agent. This run an agent. This run an agent. This could be a prompt, a could be a prompt, a could be a prompt, a request, or a context request, or a context request, or a context that forces the agent to that forces the agent to that forces the agent to act. act. act. Spreadsheet-based estimates Spreadsheet-based estimates Spreadsheet-based estimates are not are not are not wrong. However, wrong. However, wrong. However, they are often the first they are often the first they are often the first step to a useful step to a useful step to a useful version of the version of the version of the assessment workflow. But of assessment workflow. But of assessment workflow. But of course, as you course, as you course, as you can imagine, the can imagine, the can imagine, the iceberg is more iceberg is more iceberg is more than just what than just what than just what we just talked about. And we just talked about. And we just talked about. And if evaluations were just if evaluations were just , you know, running an , you know, running an , you know, running an agent, viewing agent, viewing agent, viewing the result, and then the result, and then the result, and then scoring it in a scoring it in a scoring it in a spreadsheet, spreadsheet, spreadsheet, this conversation would be this conversation would be this conversation would be over now. But there is over now. But there is much more going on behind the scenes or under the iceberg.

  6. There are many There are many support teams that support teams that support teams that ultimately need to ultimately need to ultimately need to build better build better build better datasets, datasets, datasets, scoring systems, scoring systems, scoring systems, verification process, verification process, debugging tools, and a debugging tools, and a way to connect way to connect way to connect pre-production pre-production pre-production testing with testing with production behavior. We'll production behavior. We'll touch on some of touch on some of touch on some of them today, and if I them today, and if I them today, and if I don't cover everything, you don't cover everything, you don't cover everything, you can come back can come back can come back later and we can later and we can later and we can discuss it. And this is where things discuss it. And this is where things discuss it. And this is where things start to get start to get start to get complicated. The complicated. The complicated. The underlying technology is underlying technology is underlying technology is complex. complex. complex. Master of Master of Master of Laws (LLM) programs are not Laws (LLM) programs are not Laws (LLM) programs are not simply simply simply deterministic. deterministic. deterministic. The quality of an agent is a The quality of an agent is a The quality of an agent is a problem for many problem for many problem for many people. It's not just people. It's not just people. It's not just software engineers software engineers software engineers and and and artificial artificial artificial intelligence engineers. There are intelligence engineers. There are intelligence engineers. There are project managers project managers project managers who used to who used to who used to create PRDs and now create PRDs and now create PRDs and now conduct evaluations (evals). conduct evaluations (evals). conduct evaluations (evals). There are small and medium-sized There are small and medium-sized There are small and medium-sized businesses that businesses that businesses that have expertise in the have expertise in the have expertise in the subject area of ​​the subject area of ​​the subject area of ​​the product you are product you are product you are creating. And creating. And creating. And the assessments themselves are just one the assessments themselves are just one the assessments themselves are just one part of the part of the part of the development and development and development and operation workflow. And operation workflow. And operation workflow. And gone are the days when gone are the days when gone are the days when you built once you built once you built once with with with unit unit unit testing and testing and testing and regression regression regression testing, testing, testing, shipped to shipped to shipped to production, and didn't production, and didn't production, and didn't think about it. Ratings think about it. Ratings are a way are a way are a way to continue moving to continue moving to continue moving upwards towards one goal, upwards towards one goal, upwards towards one goal, which is agent quality.

  7. which is agent quality. So, let's So, let's talk about the different talk about the different talk about the different stages of stages of stages of assessment platforms that we see. assessment platforms that we see. And before we And before we do that, it's the North do that, it's the North do that, it's the North Star for a lot of Star for a lot of Star for a lot of teams. They want teams. They want teams. They want to build an to build an to build an improvement cycle, meaning improvement cycle, meaning improvement cycle, meaning I have a feature or I have a feature or I have a feature or program that program that program that is in is in is in production. I can production. I can production. I can observe the observe the observe the failure modes that failure modes that failure modes that occur based on the occur based on the occur based on the measurements that I measurements that I measurements that I define. I want to be define. I want to be define. I want to be able to detect able to detect able to detect these failure modes and these failure modes and these failure modes and create test create test create test cases that I can cases that I can cases that I can iterate offline so that iterate offline so that iterate offline so that I can improve the I can improve the I can improve the quality of my agent quality of my agent quality of my agent without introducing new without introducing new without introducing new regressions. And then you regressions. And then you regressions. And then you keep keep keep iterating over time, iterating over time, iterating over time, because, like in because, like in because, like in classical machine classical machine classical machine learning, drift becomes a learning, drift becomes a learning, drift becomes a real thing. real thing. real thing. The way The way The way users interact with users interact with users interact with your agents your agents will change over time. The will change over time. The way you way you way you create your create your create your agents agents agents will change over will change over will change over time. So what does the time. So what does the time. So what does the first phase look like? first phase look like? first phase look like? There were some hands that There were some hands that There were some hands that went up around went up around went up around spreadsheets, spreadsheets, spreadsheets, and that's not the case. There is and that's not the case. There is and that's not the case. There is nothing to be ashamed of about this.

  8. nothing to be ashamed of about this. nothing to be ashamed of about this. This is where This is where This is where many many many people start. And since the people start. And since the people start. And since the basic configuration is basic configuration is basic configuration is quite simple. This is a quite simple. This is a quite simple. This is a four-cycle loop four-cycle loop four-cycle loop with a set of input with a set of input with a set of input examples and an examples and an agent execution method. And with this, you agent execution method. And with this, you can run the can run the can run the same examples that you same examples that you same examples that you configure and configure and configure and see how the agent see how the agent see how the agent responds with outputs over responds with outputs over responds with outputs over time. The biggest time. The biggest time. The biggest advantage is advantage is advantage is accessibility and time. You accessibility and time. You accessibility and time. You can start this can start this can start this with virtually zero with virtually zero with virtually zero barrier to entry. But barrier to entry. But barrier to entry. But the returns the returns the returns diminish quite quickly. diminish quite quickly. diminish quite quickly. At this stage, you are At this stage, you are At this stage, you are actually just actually just actually just documenting. You can't documenting. You can't documenting. You can't do do do real real real experiments. And even more so experiments. And even more so experiments. And even more so to conduct to conduct to conduct any analysis over any analysis over any analysis over time. Common time. Common time. Common limitations are that limitations are that limitations are that analytics are analytics are analytics are complex, it is manual, complex, it is manual, complex, it is manual, human evaluation is human evaluation is human evaluation is valuable but difficult, valuable but difficult, valuable but difficult, collaboration is very collaboration is very collaboration is very weak, and domain experts are weak, and domain experts are usually usually excluded from excluded from excluded from obtaining such obtaining such obtaining such results. So, results. So, results. So, again, again, again, spreadsheets spreadsheets are a great starting are a great starting are a great starting point, but they're hard point, but they're hard point, but they're hard to build on. Then to build on. Then to build on. Then someone, usually a someone, usually a someone, usually a product engineer, product engineer, product engineer, sees this problem and sees this problem and sees this problem and thinks, "Okay, I can thinks, "Okay, I can thinks, "Okay, I can create a nice create a nice create a nice UI UI UI , a , a , a custom custom custom UI UI UI , that , that , that will help solve this will help solve this will help solve this problem." And while this problem." And while this problem." And while this may be nice at may be nice at may be nice at first glance, it first glance, it first glance, it helps you helps you helps you engage with different engage with different engage with different characters that characters that characters that you wouldn't otherwise have you wouldn't otherwise have you wouldn't otherwise have access to. This allows

  9. access to. This allows access to. This allows project managers to project managers to project managers to now now now immerse themselves in the immerse themselves in the immerse themselves in the evaluation process. But evaluation process. But evaluation process. But the limitation you the limitation you the limitation you may have thought of may have thought of may have thought of remains. The fact that remains. The fact that remains. The fact that the documentation, the the documentation, the the documentation, the review process, although it review process, although it review process, although it may may may look better visually, look better visually, look better visually, still makes still makes still makes collaboration difficult. And while collaboration difficult. And while collaboration difficult. And while iteration cycles iteration cycles iteration cycles are getting faster, it's are getting faster, it's are getting faster, it's still still still difficult for users to store difficult for users to store difficult for users to store this data and conduct this data and conduct this data and conduct long-term long-term long-term analytics. It's still analytics. It's still analytics. It's still just a just a just a reporting tool. Then the reporting tool. Then the reporting tool. Then the next evolution next evolution next evolution becomes: I want becomes: I want becomes: I want my non-technical my non-technical my non-technical users to have an users to have an users to have an interface where they don't interface where they don't interface where they don't just log just log just log results, but also results, but also results, but also have the ability to have the ability to have the ability to experiment. I experiment. I experiment. I want to be able to want to be able to want to be able to customize customize customize system queries. I system queries. I system queries. I want to change the base want to change the base want to change the base model. I want to model. I want to model. I want to iterate over some iterate over some parameter values. So, parameter values. So, being able to do being able to do being able to do this kind of side-by- this kind of side-by- side comparison becomes the next side comparison becomes the next side comparison becomes the next evolution of the evolution of the evolution of the evaluation cycle that we see.

  10. evaluation cycle that we see. But there is still one But there is still one major drawback. major drawback. major drawback. Namely, the test Namely, the test Namely, the test cases you cases you cases you use for use for use for such comparisons such comparisons such comparisons are still a curated are still a curated are still a curated set that emerges set that emerges set that emerges during your offline during your offline evaluations. But what evaluations. But what evaluations. But what happens in happens in happens in production? I've done production? I've done production? I've done my iterations here. I my iterations here. I my iterations here. I have a good idea have a good idea have a good idea of ​​what I can of ​​what I can of ​​what I can expect in expect in expect in production, but once production, but once production, but once I send I send I send it to production, I'm it to production, I'm it to production, I'm working blind. So I working blind. So I working blind. So I have no idea what have no idea what have no idea what steps my steps my steps my agent is taking and when these failures agent is taking and when these failures agent is taking and when these failures actually actually actually happen in my happen in my happen in my real environment real environment . And when we think . And when we think . And when we think about about about going back to the going back to the going back to the flywheel, when we flywheel, when we flywheel, when we think about what the think about what the think about what the best systems are to best systems are to best systems are to allow for this, well, you allow for this, well, you allow for this, well, you have to be have to be have to be able to start at able to start at able to start at the top, observe the top, observe the top, observe failure modes, failure modes, failure modes, log every log every log every input, output, input, output, input, output, every trace execution step every trace execution step every trace execution step that that that your agent does, your agent does, your agent does, analyze, analyze, analyze, understand what went understand what went understand what went wrong. You create wrong. You create wrong. You create dimensions of failure and dimensions of failure and dimensions of failure and success because success because success because you know the results. you know the results. you know the results. Then you catch these Then you catch these Then you catch these failure modes, these failure modes, these failure modes, these different scenarios where different scenarios where different scenarios where your users your users your users had a bad had a bad had a bad interaction, and you interaction, and you interaction, and you build scores on them build scores on them build scores on them in your in your in your offline process, offline process, offline process, iterating on them so that you iterating on them so that you iterating on them so that you can build can build can build improvements without improvements without improvements without developing new developing new developing new regressions. And then regressions. And then regressions. And then you deliver it to you deliver it to you deliver it to production and production and production and climb the mountain climb the mountain climb the mountain against it. And when you against it. And when you against it. And when you think about the end think about the end think about the end result of that, well, result of that, well, result of that, well, you get teams you get teams you get teams that build a flywheel.

  11. that build a flywheel. You have this whole You have this whole development cycle development cycle development cycle inside your inside your inside your platform, but then the platform, but then the platform, but then the problem is that problem is that problem is that you have to you have to you have to have, you have to have, you have to have, you have to maintain it. You maintain it. You maintain it. You own this own this own this product that you product that you product that you created. And this is created. And this is created. And this is especially easy especially easy especially easy to do when it comes to do when it comes to do when it comes to smaller scale to smaller scale to smaller scale or POC and or POC and or POC and demonstration environments. But demonstration environments. But demonstration environments. But agent traces are agent traces are agent traces are unpleasant. They are unpleasant. They are unpleasant. They are semi-structured semi-structured semi-structured JSON. In some cases, JSON. In some cases, JSON. In some cases, we see teams we see teams we see teams logging logging hundreds of megabytes of movie data during each interaction. Therefore, hundreds of megabytes of movie data during each interaction. Therefore, being able to being able to being able to query this type of query this type of query this type of data can be data can be data can be challenging. And your challenging. And your challenging. And your typical cloud typical cloud typical cloud data storage seems to data storage seems to data storage seems to collapse under this collapse under this collapse under this volume and volume and volume and load. And this is the same load. And this is the same load. And this is the same problem we problem we problem we faced. I wo faced. I wo faced. I wo n't spend n't spend n't spend a lot of time a lot of time a lot of time talking about Brain Trust, talking about Brain Trust, talking about Brain Trust, but when you have but when you have but when you have production production production runs coming in, you want to be runs coming in, you want to be runs coming in, you want to be able to able to able to query that data in query that data in query that data in real real real time as soon as it time as soon as it time as soon as it gets there.

  12. gets there. gets there. Because if your Because if your Because if your users are having a users are having a users are having a bad experience, you bad experience, you bad experience, you want to be want to be want to be able to know in able to know in able to know in real real real time what's going on time what's going on , why they're having this , why they're having this , why they're having this bad experience, and how bad experience, and how bad experience, and how can I fix it can I fix it can I fix it as quickly as possible? You as quickly as possible? You as quickly as possible? You also want to be also want to be also want to be able to able to able to run long-running run long-running run long-running queries. If you're queries. If you're queries. If you're thinking about thinking about thinking about being able to fine being able to fine being able to fine -tune -tune -tune your agent experience or your agent experience or your agent experience or even having even having even having people do people do the matching by annotating the matching by annotating your your your judges' results. Two different judges' results. Two different judges' results. Two different workloads workloads . So we created an . So we created an . So we created an abstraction layer abstraction layer abstraction layer called BTQL, which is called BTQL, which is called BTQL, which is another form of another form of another form of complexity because complexity because complexity because you need an you need an you need an interface for people to interface for people to query this data. So, it's not So, it's not necessarily necessarily necessarily just a UI just a UI just a UI or UX issue, but or UX issue, but or UX issue, but creating an creating an creating an assessment platform really becomes a assessment platform really becomes a assessment platform really becomes a systemic issue. systemic issue. And a new And a new set of problems emerged when the set of problems emerged when the GPT boom in chat happened a few years ago GPT boom in chat happened a few years ago when I needed to when I needed to when I needed to perform perform real- real- time data loading. I have time data loading. I have time data loading. I have huge huge huge payloads, payloads, payloads, sometimes sometimes sometimes tens of megabytes and tens of megabytes and tens of megabytes and sometimes hundreds of sometimes hundreds of sometimes hundreds of megabytes, compared to megabytes, compared to megabytes, compared to traditional traditional heartbeat monitoring, which is heartbeat monitoring, which is only only only kilobytes in size. The structure kilobytes in size. The structure kilobytes in size. The structure and form of the data and form of the data and form of the data are different.

  13. are different. Deeply nested Deeply nested semi-structured semi-structured semi-structured full text is difficult full text is difficult full text is difficult to query. And to query. And to query. And reading patterns are also reading patterns are also reading patterns are also different. You different. You different. You want to be want to be want to be able to able to able to aggregate large aggregate large aggregate large volumes, as well as be volumes, as well as be volumes, as well as be able to take able to take able to take snapshots of the data in snapshots of the data in snapshots of the data in real time. And real time. And real time. And building the right building the right building the right system should system should system should allow you to not allow you to not allow you to not only empower the only empower the AI ​​engineer, but also AI ​​engineer, but also project managers, as well as project managers, as well as project managers, as well as small and medium-sized small and medium-sized small and medium-sized businesses. And businesses. And businesses. And recently we've recently we've recently we've seen agents seen agents seen agents become first-class become first-class become first-class citizens of these citizens of these citizens of these evaluation platforms, evaluation platforms, evaluation platforms, where you want to be where you want to be where you want to be able to able to able to use use use natural language in a natural language in a natural language in a headless experience, headless experience, headless experience, where you can say to where you can say to where you can say to your chosen your chosen your chosen coding agent, " coding agent, " Find me all the traces Find me all the traces Find me all the traces in the last 24 hours in the last 24 hours in the last 24 hours where a user had a where a user had a where a user had a bad experience." " bad experience." " Do the assessment Do the assessment for me." And because for me." And because for me." And because these coding agents these coding agents these coding agents have access to all of have access to all of have access to all of your underlying your underlying your underlying infrastructure and infrastructure and infrastructure and your codebase, your codebase, your codebase, they become the they become the they become the mechanism for mechanism for mechanism for running assessments running assessments running assessments and logging those and logging those and logging those results. And then the results. And then the results. And then the question arises: question arises: question arises: so what? Why is this so what? Why is this so what? Why is this important? Why is this important? Why is this important? Why is this important? Well, the point is that important? Well, the point is that everything I've shown you everything I've shown you so far requires so far requires so far requires people, the engineer, the people, the engineer, the people, the engineer, the project manager, project manager, project manager, to be deeply to be deeply to be deeply involved in the involved in the involved in the process, and at times that process, and at times that process, and at times that can be can be can be time-consuming. Here's how I

  14. time-consuming. Here's how I time-consuming. Here's how I think about it at BrainTrust: think about it at BrainTrust: think about it at BrainTrust: we want to we want to we want to help you help you help you work at scale. work at scale. Instead of you Instead of you having to think having to think having to think about these measurements and the about these measurements and the about these measurements and the success and failure of success and failure of success and failure of your AI agent, your AI agent, your AI agent, BrainTrust can BrainTrust can BrainTrust can automatically draw automatically draw automatically draw these conclusions. Because these conclusions. Because these conclusions. Because we log your we log your we log your tracking data on tracking data on tracking data on our platform, we our platform, we our platform, we can draw can draw can draw conclusions based on it and conclusions based on it and conclusions based on it and inform you about inform you about inform you about unknown unknown unknown unknowns. What are the unknowns. What are the unknowns. What are the scenarios where my scenarios where my scenarios where my agent silently fails agent silently fails , or what scenarios , or what scenarios , or what scenarios cause cause cause frustration for my frustration for my frustration for my users who users who ask my ask my agent the agent the agent the same query over and over again? And these are a same query over and over again? And these are a same query over and over again? And these are a lot of things lot of things lot of things we haven't talked about. we haven't talked about. we haven't talked about. Things like the underlying Things like the underlying Things like the underlying database that powers database that powers database that powers it, or it, or it, or the need to the need to the need to embed our embed our embed our internal interface internal interface internal interface into the system to into the system to into the system to manage manage manage controls and controls and controls and permissions, or permissions, or permissions, or data masking. data masking. data masking. You know, each of these You know, each of these You know, each of these features proves that features proves that features proves that it's not just a it's not just a user interface and a user interface and a spreadsheet, but spreadsheet, but spreadsheet, but rather a systemic rather a systemic rather a systemic problem that problem that problem that makes it work.

  15. makes it work. And when we And when we think about the think about the think about the evolution of the evolution of the evolution of the improvement cycle, I've improvement cycle, I've improvement cycle, I've talked about this before, talked about this before, talked about this before, where people used to do the where people used to do the where people used to do the improvement cycle, and improvement cycle, and improvement cycle, and now we see now we see now we see coding agents coding agents coding agents that can that can that can iteratively make iteratively make iteratively make changes and suggest changes and suggest changes and suggest what improvements you what improvements you what improvements you should make to your should make to your should make to your program. And ultimately, it is the program. And ultimately, it is the person's responsibility to person's responsibility to review the outcome review the outcome . If I have different . If I have different . If I have different iterations of these e-vals that iterations of these e-vals that iterations of these e-vals that my encoding agent runs my encoding agent runs my encoding agent runs , I can , I can , I can review the review the review the output and output and output and decide which one is decide which one is decide which one is best best best for what I want to for what I want to for what I want to release to production. release to production. So, this So, this concludes my report. I concludes my report. I concludes my report. I appreciate all of you appreciate all of you appreciate all of you coming and coming and coming and learning why it's hard learning why it's hard learning why it's hard to create to create to create assessment systems. Thank you.

Summary

The session focused on the challenges and best practices for building quality AI agent platforms, highlighting the critical pillars of evaluation (pre-production testing) and observability (in-production monitoring). The core takeaway is that while LLMs offer powerful flexibility, their non-deterministic nature necessitates rigorous assessment and continuous monitoring to ensure reliable and confident agent performance.

View original episode ↗