Full Transcript

·YouTLDR

The Production AI Playbook: Deploying Agents at Enterprise Scale — Sandipan Bhaumik, Databricks

36:49EnglishTranscribed Jul 23, 2026
0:07

[music]

0:15

>> All right. Um, thank you for joining my

0:17

session.

0:17

>> [applause]

0:20

>> Thank you, man.

0:21

Uh, I'm Sandy. Uh, I'm a technical lead

0:24

uh, for data and AI at Databricks. Um,

0:27

prior to working in Databricks, I worked

0:30

in Amazon Web Services uh, for 5 years

0:32

as a principal architect for data and

0:33

AI.

0:34

Uh,

0:35

in the past few years, I worked

0:36

extensively uh, building and scaling

0:39

data and AI platforms using distributed

0:41

systems and technology.

0:43

And in the past couple of years,

0:46

specifically, I've been working with

0:48

customers trying to figure out what we

0:49

do with this new AI technology.

0:51

Uh, when I say new AI, AI has been here

0:54

for a long time, but we all started

0:57

experimenting quite exponentially uh, in

1:00

in the past couple of years, right? And

1:03

I have learned a great deal of lessons

1:05

on from building demos and how to take

1:08

those demos to production working with

1:10

different customers in

1:12

uh, B2B software industries and then uh,

1:14

regulating industries like financial uh,

1:16

services.

1:18

So, in this session,

1:20

I want to share

1:22

a playbook, a framework that I put

1:24

together from lessons that I have

1:27

learned working in the trenches

1:29

uh, that you can take and apply on when

1:32

you think about how to put your AI

1:34

systems into production. And I think

1:36

this session is nicely placed in the

1:38

afternoon because what you can do now is

1:40

in this framework, you can fit the

1:42

different

1:43

um, knowledge, the knowledge that you've

1:45

gathered attending these different

1:46

sessions throughout the day and see

1:48

where they fit in each of these

1:50

you know, elements in the framework.

1:52

So, when I started 2 years ago, this is

1:54

the pattern I noticed in every customer

1:56

conversation, right? So, everyone wanted

1:59

to do something with AI. Uh, there was

2:01

immense pressure from the top to do

2:04

something, to build a demo.

2:06

And every conversation started with

2:07

let's choose the model, right? And it

2:09

was nobody's fault because the market

2:11

was like that. We were talking about

2:12

models, the models were new technology

2:14

for us, right? And every conversation

2:16

started, shall we use GPT? Shall we use

2:19

Claude?

2:20

You know, there was huge debate with

2:22

with within organizations. Then, you

2:24

would choose a model, you'll build some

2:26

features, offsets of over what features

2:28

to build for that application.

2:30

Uh, you would build that in a controlled

2:32

environment, so predictable data sets,

2:35

you know, um, limited scenarios, and

2:39

then it looked great as a demo, and then

2:41

leadership would get happy, they would

2:42

sign it off, and they'll put it into an

2:45

environment in a production environment.

2:48

Then,

2:49

after a few weeks, people would start

2:51

asking questions that what the hell is

2:54

AI doing?

2:55

Right? Why is it not answering the

2:57

questions the way we expected it to

2:58

answer when we were doing the demos?

3:01

Uh, it would result in not only

3:04

less you know, no realization in return

3:06

on investment, but also loss of money

3:08

and effort in building these demos that

3:10

can never scale to production.

3:15

Throughout these uh, meetings, I

3:17

gathered three insights that connect to

3:20

everything that you we are talking about

3:22

when thinking about taking uh, AI to

3:25

production. The first one is the

3:26

observability gap, right? When we use AI

3:30

and put it into production, if we can't

3:32

see what it is actually doing, if we

3:34

can't trace every decision that it's

3:35

making, it's no use in production.

3:38

Second is the evaluation gap gap. A lot

3:41

of these conversations that we were

3:42

doing, we were not actually thinking

3:44

about what is what is that one thing

3:47

that we are measuring. Yes, we talk

3:49

about accuracy, we talk about latency,

3:52

we talk about groundedness, but we were

3:55

not defining what is that exact thing

3:57

like that matters to the business, and

4:00

how can we build a system that can

4:01

continuously measure that, whether it's

4:04

improving, whether it's not improving,

4:05

like what what is that system that we

4:07

need to build. And that was that

4:08

evaluation gap that I noticed. And the

4:10

third is the governance gap. Like we

4:13

were not actually thinking what happens

4:15

when AI fails in production. Who's

4:17

accountable? Who do I go to when

4:19

something happens at 3:00 a.m. in the

4:21

morning, right? Who needs to own the

4:24

data assets that feed some AI responses?

4:27

What happens if AI

4:30

you know, uh, talks um, uh, um,

4:34

nonsense to a customer, right? What what

4:35

happens, right? So, there is no

4:37

accountability, no governance around it.

4:39

And these three insights led me to

4:42

build a framework

4:44

on how

4:46

I think AI should be taken to

4:48

production, and this has been

4:49

implemented across multiple customer

4:51

organizations, and I think this is

4:53

something that you can pick up from

4:54

here.

4:55

These are the five pillars, and these

4:57

are absolutely what you need to think

4:59

about even before starting a project,

5:01

right? Then you start build them

5:03

gradually,

5:04

preferably in sequence, but in real

5:06

life, I know that this sequence don't

5:08

work, but these are the pillars that you

5:10

have to know about and you have to think

5:11

about when start building. First one is

5:13

evaluation.

5:15

Before touching any code, before

5:16

discussing about any models, any

5:18

features, you have to think about when

5:20

we build this system, how do we measure?

5:23

What does success look like, and what is

5:25

that system that will help us

5:26

continuously measure what success looks

5:29

like for us?

5:31

Second is how do we trace each and every

5:34

decision that AI makes. It's not only

5:36

important for the performance of the AI

5:39

system, it is also important for the

5:40

regulators. In Europe or in a lot of

5:44

companies, especially in regulated

5:45

industry,

5:46

you cannot even onboard AI into

5:49

production without having tracing and

5:51

observability in place.

5:52

So, this is a must-have.

5:54

The third is the data foundation,

5:56

right? Uh, I I I I think of data

5:58

foundation in in two ways. One is the

6:01

question data, so that is basically

6:04

the data needed for the AI to answer

6:06

questions that users ask to it. So, it

6:09

could be your pre-training data, post

6:11

post-training data, data that you use

6:13

APIs to hook onto and and get to the

6:16

answer that the user needs. The other

6:18

one is the tracking data, related to the

6:19

tracing data in observability, but when

6:21

you think from the data foundation and

6:23

and uh, data strategy perspective, this

6:26

needs to be handled in this pillar

6:28

because you need a whole data strategy

6:30

now with tracing data, especially when

6:32

you run hundreds of agents in your

6:34

organization.

6:36

Fourth is orchestration. One agent would

6:39

work pretty well. You don't need to

6:40

think about orchestration. But when you

6:42

onboard five agents, the the complexity

6:45

increases exponentially, right? You will

6:48

have multiple coordination patterns

6:49

between these agents, they will need to

6:52

talk to each other in multiple different

6:54

ways, they will need to each wait for

6:55

each other's responses, there's a lot of

6:57

complexity that comes in. And that's

6:59

where orchestration patterns and

7:01

thinking about how you will orchestrate

7:02

your agents in a particular system

7:05

becomes really important.

7:06

Fifth is governance. This is where

7:08

you think about what happens when

7:10

something fails. Who's accountable? How

7:12

do we govern data? How do we secure it?

7:14

How do we secure our systems? How what

7:16

do we make sure how do we make sure that

7:17

no one injects into our agent and leads

7:21

uh, to you know, misbehavior, right? Or

7:23

loss of reputation.

7:27

So, in the rest of the session, I will

7:28

dive a bit deeper into each of these

7:30

pillars and tell you how how you can

7:31

think about when you start working with

7:32

them, right? The first one is

7:34

evaluation.

7:35

Evaluation is basically specification

7:37

for your AI system.

7:39

You define success. As I mentioned, it's

7:41

not like

7:42

talking about accuracy. You have to

7:43

define it with numbers, like what

7:45

accuracy is is is good for your business

7:49

use case,

7:50

right?

7:51

Define it in numbers.

7:52

Uh, define what kind of you know, false

7:54

positives you can handle. What should be

7:57

the deflection? So, this is this is an

7:58

example from a

7:59

a retail chatbot, right? A banking

8:01

chatbot where

8:02

when you implement a chatbot with an AI

8:04

agent, one of the main goals is to

8:07

deflect simple queries

8:09

um, to the agent so that a human agent

8:12

don't need to uh, deal with them, right?

8:15

And so, you need to uh,

8:16

you need to track those queries and

8:18

track those numbers and put that system

8:20

in place. Second is building those test

8:23

test cases, like the evaluation data

8:25

set. You've heard about golden data sets

8:26

in evaluation. Talk with the domain

8:28

experts and find what is actually

8:30

happening in real life on the ground.

8:32

Like what answer would a support human

8:34

support agent

8:36

um, give to a customer on a particular

8:38

question. Collect those information.

8:40

What happens in gray areas, in edge

8:42

cases, like what happens when a human

8:45

sees a customer asking a confusing

8:46

question, right? Collect those into a

8:48

data set, and then automate your AI

8:52

testing, right? So, you put a question

8:54

to AI, it answers, take that answer,

8:57

compare against the test set, and

8:59

automate this whole pipeline so that

9:01

when you put AI in production, that

9:03

pipeline can actually take live

9:04

responses and evaluate against the test

9:07

data set that you're building, and then

9:09

give you the result in terms of how AI

9:12

is performing against those numbers and

9:14

the goals that you've defined.

9:17

When we talk about evaluation,

9:19

there are three main layers that I see

9:22

appear across organization, and this is

9:24

an architectural decision that you need

9:26

to make when you build these evaluation

9:27

systems, right? The first layer is

9:29

deterministic. These are the easy stuff,

9:31

like you know, checking formats, you

9:33

know, checking email formats, phone

9:35

formats, the regular expression things

9:37

that we have already been doing with our

9:39

coding systems, right? The uh, the the

9:41

other other is like you know, you could

9:43

use a classic ML models for name entity

9:46

recognition to for intent

9:47

classification,

9:49

for understanding what is first name,

9:52

last name, PII detection, etc.

9:54

So, the these these are easy stuff,

9:56

cheap stuff, you should get them out of

9:57

the way. We have already been doing this

9:59

for years. The second layer is the

10:01

non-deterministic semantic stuff, all

10:03

right? This is where groundedness comes

10:04

in. This is where we implement

10:06

technologies like LLMs judges. We all

10:08

know what LLMs judges are, right?

10:11

Everyone? Okay, I see a lot of nods. So,

10:14

um

10:15

Again, this This is a pretty simple

10:18

version of how a uh how a prompt would

10:21

look for an LLM as a judge. Um

10:24

So, with LLM as a judge, you you you use

10:27

a separate LLM from the primary LLM to

10:29

judge the response of the primary

10:32

model.

10:33

And when you do that, you tell the

10:35

secondary, the judge model,

10:37

on how it should uh judge the primary

10:40

model's output. So, it could be around

10:41

safety, groundedness, you know,

10:44

relevance to the answer, etc. etc.,

10:46

right? Again, that can feed from a lot

10:49

of uh the evaluation data set that you

10:51

have created, right? To look at what are

10:53

the expected answers, and then it can

10:55

check against that. This is a sample

10:57

prompt on how these things work, but I'm

10:59

sure you've attended some of these

11:00

sessions where you've seen vendors doing

11:02

this automatically at scale. Uh for

11:05

example, in Databricks we in MLflow

11:06

you'll find automatic LLM as judge, uh

11:09

where you can create these custom LLM as

11:11

judges that run automatically on traces.

11:15

That's your second layer. The third

11:16

layer is behavioral, right? This is

11:18

where uh you think about a tool calls,

11:20

like is our agents calling the right

11:22

tool? Are they getting into loops? So,

11:25

for example, um you know, the first

11:28

layer you you can have a user ask a

11:29

question, "What is my account balance?"

11:31

And

11:32

you could go and check that, okay, this

11:34

there is no deterministic problem with

11:36

it. The seman- the agent answered right,

11:39

that, "Okay, your account balance is

11:40

this many dollars."

11:42

And that was right, and you can see this

11:43

is this is right, but when you go into

11:45

the behavioral checks, you will see that

11:46

the agent uh actually made three calls

11:49

to the database to find that answer.

11:52

Right? And that is because it was doing

11:54

making duplicate calls for whatever

11:55

reason.

11:56

Calls failed,

11:58

you know, calls did not work, it went

11:59

and retried and stuff like that. Now,

12:01

three API calls in demo environment is

12:04

fine, but in production, when you get

12:05

thousands of queries from users every

12:07

day, and there's like duplication in API

12:09

calls, that's an expensive operation.

12:11

And that's where you need to think about

12:12

behavioral evaluation. And this layer is

12:15

very, very important. I see a lot of

12:17

organizations, a lot of teams miss them

12:19

when when talking about this.

12:25

The second layer is observability,

12:27

right? Uh and in this pillar, what we're

12:30

talking about tracing, right? So, you

12:31

collect [snorts] all the decisions that

12:34

an agent is making. So, I want to

12:35

explain this with a scenario here,

12:37

right? And this is a scenario from an

12:39

actual project I worked on with a

12:41

banking uh retail retail banking

12:43

chatbot. Now, obviously, if you've seen

12:45

tracing data, it's not as beautiful as

12:47

this slide, right? So, I've simplified

12:49

it and made it beautiful for this slide.

12:51

But what this slide says is basically, a

12:52

user comes in and says, uh "You know, I

12:55

have been charged an overdraft fee, can

12:56

you waive it for me?" Because the user

12:58

thinks that the customer thinks that

13:00

that is not legitimate.

13:01

So, the agent does an intent

13:03

classification,

13:05

and you all you know about this because

13:06

you've enabled observability, you're

13:08

capturing traces, and you're actually

13:09

seeing what the agent is doing, right?

13:11

What AI is doing.

13:12

Intent classification, it is done, this

13:15

it took this many seconds, this was

13:16

confidence score. Then it goes and

13:18

connects to the customer's account,

13:19

maybe in a database, a customer

13:21

database, call calls an API, connects to

13:23

the customer database, gets the account

13:24

details.

13:25

It retrieves policy documents. It checks

13:28

from a rag vector database,

13:30

um what is uh what is the policy around

13:33

overdraft, right? Is what the customer

13:36

claiming is legitimate? So, it checks

13:38

for policy documents. Then it goes and

13:40

does a reasoning on what should be

13:43

uh

13:44

you know, responded to the customer, and

13:45

then it does some final guardrail

13:47

checks, and responds to the customer.

13:49

Now, if you did not set up a system that

13:52

helps you look visualize all of these

13:55

traces,

13:56

when the customer comes to you and

13:58

raises a dispute,

13:59

you have no way to check what the AI

14:01

did.

14:03

Right? You have nowhere to go, and you

14:04

end up saying that I don't have have any

14:06

idea. Let's Let's give the customer a

14:08

discount or something, and then make

14:09

them happy.

14:11

So, this is why you need this, and this

14:13

is why regulators are are are basically

14:15

mandating, because otherwise there's no

14:17

production system if you cannot do this

14:18

kind of stuff.

14:22

So, this is where um

14:24

you know, you you you detect this the

14:26

example that I gave around duplicate API

14:28

calls. This is where you start detecting

14:29

this stuff. So, when you when you enable

14:32

these traces, you can actually go and

14:33

see duplicate calls, and then take

14:35

relevant actions based on that. Not only

14:38

that, you can actually do that in online

14:39

monitoring. So, when it's happening in

14:41

production, in on you can set up online

14:43

monitoring, and at that point, if it is

14:45

doing duplicate calls, you can apply

14:47

fallback strategies. Or even if it is

14:49

doing a call that is failing, you can

14:51

actually go and apply a strategy where

14:53

it will say, "Okay, go and retry for

14:55

three times, not more than three times.

14:57

If it if it is more than three times,

14:58

then report somewhere, or pass it to a

15:00

human to take some action."

15:06

The third pillar is the most important

15:07

pillar, in my opinion, is the data data

15:09

data foundation, right? Uh

15:12

in my typical project projects, I spend

15:15

60% of my time, uh and I see I see a lot

15:18

of organizations spending a lot of time

15:20

here, because

15:21

no one expected agents to come suddenly

15:24

in the market and start querying data.

15:27

Data was always built for humans, and

15:29

humans are always forgiving.

15:31

You find the wrong data in a report, you

15:32

just go and ask someone to correct it.

15:34

Agents don't forgive you, right? Agents

15:36

will go, find it wrong, they'll give you

15:38

the wrong answer confidently.

15:40

Right? And you wouldn't know what's

15:42

happening.

15:43

And this is why data quality,

15:45

setting the right data strategy, has

15:47

become so important for enterprises now.

15:50

I divide it into two sections. One is

15:52

the question data, as I was explaining,

15:54

like data needed for actually serving

15:56

the AI's uh outcome.

15:59

And the other one is the tracking data.

16:01

This is the observability data, the

16:02

tracing data I was talking about

16:03

earlier.

16:04

You need a proper plan on how you

16:07

collect this tracing data, and how you

16:10

serve it to auditors, to regulators, to

16:13

do online monitoring, to run LLM as

16:15

judges on the tracing, and everything

16:17

else, right? So, there it needs a proper

16:19

strategy on how you structure the schema

16:20

and everything on the tracing data.

16:24

Um

16:27

On Databricks,

16:28

um we

16:31

create a robust data foundation for our

16:33

customers

16:34

using uh some of the technologies that

16:36

we provide. If you don't know

16:38

Databricks, Databricks has been built on

16:39

some open-source technologies like

16:40

Apache Spark, MLflow, and Delta Lake.

16:44

Uh we provide a bunch of capabilities on

16:46

top of it. So, the blue layer at the

16:48

bottom is basically your cloud storage.

16:50

Databricks works on the three major

16:52

clouds, Google, AWS, Azure.

16:57

Okay.

16:59

I thought it was for me.

17:01

So, so once you once you store raw data

17:03

on your cloud storage,

17:05

uh the data is then um

17:08

um we we we bring in a a layer called

17:10

the Delta Lake layer, which uh which

17:12

basically brings in database-like

17:14

properties on top of your raw data. So,

17:16

you have got images, text files, video

17:18

files, or whatever. We we help you

17:20

create this um

17:22

you know, uh table-like structure on top

17:24

of it using manifest files, right? And

17:27

and we help you to uh incrementally load

17:30

data, do all of those um data management

17:32

tasks in a structured way.

17:35

On top of that, we bring in Unity

17:36

Catalog, which is a data catalog. Uh

17:39

with Unity Catalog, you can centrally

17:40

apply permissions on top of the data.

17:43

You can

17:44

um you can uh share the data using uh

17:47

Delta Sharing, but also uh what happens

17:50

with Unity Catalog is uh you you can

17:53

enable discovery and um you know, um

17:56

ownership, metadata tagging capabilities

17:59

at the catalog level. What that means

18:01

is, when you apply table a description,

18:03

column description,

18:05

uh tag columns uh like PII columns with

18:07

metadata, it becomes really easy for AI

18:10

to then get that context when it queries

18:12

these tables on top of Unity Catalog.

18:15

So, everything is governed at one layer

18:17

through Unity Catalog, and on on top of

18:19

that we bring in different uh

18:21

applications. So, whether it's AI

18:23

through Mosaic AI, so to build LLM, tune

18:26

LLM, or even build AI applications, we

18:28

bring in uh data warehousing

18:29

capabilities, BI capabilities, and um

18:33

uh and some of the other text-to-SQL

18:35

capabilities. We have got Genie that uh

18:37

helps you write natural language to do

18:39

SQL querying, etc.

18:42

And one application of that in the

18:43

observability and tracking tracking

18:45

data, as I was showing, is is this. So,

18:48

basically, think about when I was

18:49

talking about the tracking data

18:50

strategy.

18:52

Organizations, especially enterprises,

18:54

will not be running AI in just one

18:56

framework. They'll be using different

18:58

frameworks,

18:59

CrewAI, LangChain, etc. etc. They'll be

19:02

using different cloud platforms.

19:04

And once they do that, you need a

19:06

centralized layer of collecting that

19:08

tracing data, so that you can serve sev-

19:11

several use cases on the right hand

19:13

side. So, whether it's for operational

19:14

dashboarding, for first line support, uh

19:18

a lot of these uh first line

19:20

um first line of defense teams need

19:22

health monitoring uh sort of dashboards,

19:25

right? These teams can also write SQL

19:27

using Databricks Genie to do

19:29

text-to-SQL. But they can also build

19:31

Databricks apps using coding agents uh

19:34

to create common workspaces or custom uh

19:38

UIs that customers might need for

19:39

different uh different use cases.

19:42

And then we've got Agent Bricks and

19:44

MLflow that serves you uh LLM out of the

19:46

box LLM as judges,

19:48

and uh proactively monitor a

19:51

The idea is,

19:52

no no no matter where your AI runs, you

19:55

can create this kind of strategy

19:57

bringing in data in one common place and

19:59

serving uh different teams from one

20:01

shared location.

20:05

The fourth pillar is multi-agent

20:07

orchestration patterns. As I said, one

20:09

agent is good, multiple agents increases

20:11

complexity. That's where you start

20:13

thinking about, okay, what pattern is

20:15

good for my use case. The first one I

20:17

describe here is the orchestrator worker

20:19

pattern. Where you have one orchestrator

20:22

which orchestrates all the work, which

20:24

controls all the work from a centralized

20:26

plane, and then distributes this work to

20:29

different agents based on their

20:30

specialized skills.

20:32

And then every request goes through the

20:34

orchestrator, so you have got central

20:36

control. If something goes wrong, you

20:38

can go to the orchestrator logs and look

20:40

into them and see what has happened.

20:42

Right? So, that's the orchestration data

20:44

uh pattern.

20:45

There is this choreography pattern where

20:48

each agent is independent, they're

20:50

autonomous, they don't depend on an

20:51

orchestrator. All of them talk to a

20:54

message bus

20:55

and they listen to the events that they

20:57

are interested in.

20:58

Right? So, think about agents that are

21:00

independent of each other, right? They

21:01

can run parallelly. So, they are not

21:03

sequential, like one agent is not

21:05

dependent on another. So, they run

21:07

parallelly, they listen to the message

21:08

bus for the

21:10

for the events that they are interested

21:11

in. Maybe it's a trigger for, let's say,

21:13

a mortgage application, and it says, uh

21:16

you know, uh the mortgage application

21:17

agent

21:19

uh one of the agent uh looks customer

21:21

details, right? The other agent looks at

21:23

approval details and everything else,

21:24

right? They can work in parallel, and

21:27

the advantage it brings you is the

21:28

latency is reduced because they are not

21:30

dependent on an orchestrator and sending

21:32

messages back and forth. Right? So, this

21:35

is the choreography pattern. And the

21:36

third one is human in the loop, which is

21:39

where when an agent crosses a threshold

21:42

or serves below threshold a confidence

21:44

threshold, then a human is called in the

21:46

workflow to look into the pattern uh so,

21:49

look into the looking into what the

21:50

agent has done and then take action

21:53

based on that.

21:55

I have

21:57

done a deep dive video on multi-agent

21:59

orchestration pattern

22:01

uh for the online track of this

22:02

conference. Uh it's already on YouTube,

22:04

so you can look into it. I talk about

22:07

the real implications of when you think

22:09

about multi-agent patterns.

22:11

One is uh state management, the other is

22:15

fault tolerance, like what happens when

22:16

things fail, like how do you manage

22:17

them? I talk about different patterns.

22:19

And then talk about how how you think

22:21

about scaling them in large scale on

22:22

enterprises.

22:25

Pillar five is governance, right? Now,

22:27

here I'm not talking about data

22:29

governance at all. That's given, we need

22:31

that, right? From AI perspective, what

22:33

what what are we thinking about?

22:34

Regulatory, right? Audit trails, have we

22:36

got the trail of every action, every

22:39

user connection, every request,

22:41

everything that happens in the system?

22:43

Are we capturing everything?

22:45

Are we doing pre-validation of personal

22:48

information? Are we using name entity

22:50

recognition? The the easy stuff, the

22:52

rejects and all of those things, right?

22:54

In our example, the work that I was

22:56

doing with the customer that I

22:57

mentioned, we already detected 47 PII

22:59

breaches during the testing phase by

23:02

applying this layer. So, that's that's

23:04

really important.

23:06

Fourth is um prompt versioning. You have

23:08

to treat prompt versioning as change

23:10

management in enterprise grade solution.

23:13

It cannot be just change to a prompt and

23:15

commit to get. It has to be is it has to

23:18

go through proper change management

23:19

processes as you do with code. So,

23:21

basically treating prompt as code.

23:24

Third is model change management. So, as

23:26

models change, the model providers

23:28

upgrade these models, you have to have a

23:30

system to understand whether that

23:31

upgraded model will be good for your use

23:33

case, for your data. Right? Model

23:36

providers you will put evaluation

23:38

benchmarks on three uh benchmark uh

23:41

boards,

23:42

but those are not really useful when you

23:44

put them in your context, in your

23:46

enterprise. So, that's where these

23:48

evaluation data sets come in handy,

23:49

where you try these different models on

23:51

this evaluation data set and try to

23:53

understand which one performs better.

23:55

And that management needs to be done

23:56

because

23:58

from a risk perspective, you cannot

23:59

really rely on one single model. You

24:01

have to have the flexibility to switch

24:03

to different models and also test them

24:06

on your own data.

24:07

That management needs to be done.

24:12

Uh in Databricks, uh we have taken all

24:14

of these these points, these pillars

24:16

that I've been talking about into Agent

24:18

Bricks. We are building Agent Bricks to

24:21

make uh all of these operations out of

24:24

the box for you,

24:25

uh so that it's easy to implement

24:27

production grade AI applications on uh

24:30

on in in your enterprises.

24:35

So, I wanted to quickly touch upon a

24:37

case study, just to give you a uh a

24:40

flavor of how these things go, right?

24:42

So, when I was working with this client

24:45

um

24:46

they were a retail banking they were

24:48

building a retail banking chatbot you

24:49

know, one and a half 18 months ago.

24:52

Uh

24:53

their their problem the the problem they

24:55

wanted to solve is they had got around

24:56

20,000 odd calls per month from

24:58

customers on their chatbot. They wanted

25:01

to deflect they they they saw that there

25:04

were like 60% of them were simple

25:05

queries, what is my account balance, you

25:08

know, what do I do with my overdraft and

25:09

all of those stuff, like that can be

25:11

answered simply. So, they wanted to the

25:13

reliance on human agents for those

25:14

answers. So,

25:16

they identified those queries and they

25:18

wanted to automate them. Right?

25:22

They spent around 85K in 6 months doing

25:24

a POC which did not succeed. When we got

25:26

involved, we found those insights that I

25:28

was talking like no one knew why things

25:30

were failing when it was in production

25:32

when when when when it was

25:34

in production. No one could actually

25:37

measure

25:39

why

25:40

why it's not succeeding and no one could

25:42

actually understand who is accountable

25:44

for what when things go wrong. Right?

25:47

So, the goal we set for them is AI agent

25:50

handles 60% of user queries, right?

25:53

Which were simple user queries and then

25:55

a way to identify and track them. The

25:57

key difference in this project that we

25:59

when we did is that we selected the

26:00

model in week seven, like in a eight

26:03

weeks POC.

26:04

Right? And this is how it turned out.

26:07

For the week one and two, we built the

26:09

evaluation layer. We collected 200

26:12

cases on their actual human agents

26:15

answering to their customers on simple

26:17

queries and understand how they are

26:19

responding to them.

26:21

We created that database.

26:23

Then we defined the success metrics.

26:24

What does success look like to you? So,

26:26

out of let's say 100 queries, you need

26:28

60 queries or the 60% of the queries

26:31

that are simple queries

26:32

to be

26:35

uh to be handled by the agent, right?

26:36

They needed some sort of accuracy. So,

26:38

85% it was around 85% accuracy target.

26:40

They needed latency, all of the

26:42

operational targets that you need. They

26:44

were there.

26:46

Then we created this automated

26:48

evaluation pipeline for them. And what

26:50

what I mean by that is an automated

26:51

system where you can capture a user's uh

26:54

a user's question and the AI agent's

26:56

response. You take that, compare that

26:59

against your evaluation data set.

27:02

You rate that, and if the rating is

27:04

below certain threshold, you get it

27:06

checked by a human. And you if if

27:08

something goes wrong, you make sure that

27:11

you find the solution. So, it could be a

27:12

change to the prompt, it could be change

27:14

to a tool calling system or something

27:15

else. Once you have done that, you add

27:18

that test case in the test data set. So,

27:20

that when it happens next time, the test

27:22

cases cases catch them. So, the the the

27:26

the summary of that story is that your

27:27

evaluation data set is a living system.

27:30

You start with 200, maybe there is no

27:32

correct number here, but once you start,

27:34

as you start building in production,

27:36

this is a living system. This will keep

27:39

growing. And the

27:40

and the bigger it grows, the better your

27:42

system will be.

27:45

In the second week, we talked we thought

27:47

about the foundational layer, right? So,

27:48

the the question data, we thought we

27:50

thought that, okay,

27:51

if you have to call the database, have

27:54

you got the API connections right? Have

27:55

you got a system in place that can trace

27:57

the API connections? Are those secure,

27:59

right? We were not talking about MCP at

28:01

that time.

28:02

Right? It was just direct API calls to

28:04

database to run queries.

28:08

Have you got the distributed storage?

28:09

Have you got the Have you Are you

28:11

collecting traces?

28:12

And this is where when when we started

28:14

testing after building these systems, we

28:16

could catch those duplicate API calls,

28:19

right? We could catch why customer

28:20

satisfaction was dropping and stuff like

28:22

that.

28:24

And then comes in week seven to eight,

28:26

we started talking about models. Now

28:28

that we had the evaluation data set, we

28:30

could run different models on that data

28:32

set to see the responses, compare them

28:34

against the expected responses, and

28:36

calculate a number on on on the

28:38

accuracy, right? That helped us to

28:41

decide which model to use.

28:44

Now, that decision didn't take long,

28:45

right? We in as I explained in the

28:48

introduction, like we spent weeks

28:49

debating on which model to use,

28:52

but when you took the other approach,

28:55

you can actually

28:56

uh do that in a very quick way.

29:00

So,

29:01

once once that's done, we we stitched

29:04

everything that I was talking about

29:05

around observability, evaluation, the

29:07

layers of evaluation. Once we had that

29:09

system that can make AI visible,

29:11

measurable, and accountable, that's when

29:14

we started launching it to production.

29:18

And that's when uh so, this is this is

29:19

the result uh six weeks post launch,

29:22

uh we we calculated the operational

29:24

metrics, of course, you know, the

29:25

accuracy, the deflection rate, the

29:28

response time, uh the customer CSAT. But

29:31

what's important here is

29:33

in few weeks time, when uh there was a

29:35

problem with uh so, one of the one of

29:38

the things that happened was that the

29:39

bank changed some uh

29:41

interest rate related policies. So, when

29:43

that when they changed the policy, they

29:45

actually sent emails to customers or

29:47

notifications in the application in the

29:49

in the in the mobile banking app about

29:51

the policy change.

29:52

But when the customers came and queried

29:55

on the chatbot for further questions,

29:57

they couldn't get the right answers and

29:59

they were like

30:00

putting thumbs down on the answers. So

30:02

they were getting this feedback, right?

30:03

So feedback decreased.

30:06

The problem with this kind of system, if

30:07

you did not have this measurement

30:08

system, is that you couldn't actually

30:10

know what's happening.

30:11

But because we had the measurement

30:13

system in place,

30:14

these

30:15

the drop in C set was detected, right?

30:18

Because we were getting negative

30:19

feedback from customers.

30:21

We could actually look into the tracing

30:22

decisions and see that the agent was

30:24

looking at a policy policy document that

30:27

was outdated. So the the new policy

30:29

document was not updated in the vector

30:31

database. The embeddings did not come

30:33

through.

30:34

Because it did not come through, it was

30:36

giving it stale answers. And that's when

30:39

we went and fixed that. But it it was

30:41

all possible because we built that those

30:43

systems

30:45

that

30:46

that led us

30:47

to to detect this.

30:53

Before you go, I I generally in these

30:55

sessions I share different

30:57

artifacts that you can take away. I have

31:00

a QR code at the end for you to download

31:02

and you will find multiple artifacts.

31:04

One of the important artifacts that I

31:06

want to talk about is the production

31:07

incident playbook. This is something

31:09

that a lot of us tend to miss when we

31:12

work in AI projects. And this playbook

31:15

is basically a definition of what needs

31:17

to happen when things fail in

31:19

production.

31:21

First, you detect using your eval

31:22

dashboard.

31:23

Then you diagnose using your tracing as

31:25

I explained.

31:26

Then you contain. So basically, you you

31:29

are versioning your prompts. Is there a

31:31

Is there is a If there is a problem with

31:32

the prompt, you you take that prompt

31:35

out, right? And start the changes,

31:37

deflect it to a human. Or in my

31:40

multi-agent orchestration video, I've

31:41

talked about multiple fault tolerance

31:44

failure recovery patterns

31:46

around saga pattern, compensation

31:48

pattern,

31:49

and circuit breaker pattern that you can

31:51

look into the video. I've explained them

31:52

in details on how you can handle them.

31:55

And then you use the test case library

31:56

to fix. So you look into LLM's judge

31:58

reports, you look into your evaluation

32:00

data set reports.

32:02

Then you fix your problem. Once you fix

32:05

your problem, you put those test cases

32:06

in your data set, right? And and create

32:09

that eval suite that is a living system

32:11

that will keep growing.

32:13

And

32:14

and and you you keep improving your AI

32:16

system based on that, right? But this

32:18

playbook needs to be in place. When it

32:20

runs in production, you will need to

32:22

integrate it with your ITSM system so

32:24

that it alerts the right person at the

32:26

right time. You know, a lot of these

32:27

organizations would have existing

32:30

ITSM systems, right? So which

32:32

which is used for alerting and you know,

32:35

making sure that the downstream systems

32:36

don't get affected, etc.

32:38

So

32:39

once once you have this in place, you

32:41

can go and stitch it together to other

32:43

systems.

32:46

So what can you do tomorrow, right?

32:49

Start with If If you have a project in

32:51

mind, start with defining success.

32:53

Success not from the technical sense,

32:55

from the business sense. What it means

32:56

means for the business, right? Come up

32:58

with a few examples of what good answers

33:01

look like.

33:02

And and create a data data set of that.

33:04

And then build that pipeline using

33:07

simple Python code. See if you can

33:09

automate that so that when you run AI

33:11

and get some response, it can become it

33:13

can go and compare You can go and

33:15

compare the answer against that data set

33:18

and then that can be delivered

33:20

to to the to the customer.

33:26

Now, these are three lessons that I have

33:28

learned while doing these things with my

33:31

you you know, my easily miss.

33:34

The test case library, as I explained,

33:35

is a growing system. It will grow over

33:37

time.

33:38

And because it grows over time, you need

33:40

some sort of governance around it. You

33:41

need a owner, right? You need to You

33:44

need to figure out which test cases

33:48

relate to what kind of problem. So that

33:50

whenever you go back to it, you can

33:53

you can relate your answers to those

33:54

sort of problems, right? If it is a

33:55

security If it is login, so you can say

33:57

that the agent did not ask for login

34:00

credentials when the customer asked the

34:01

answer. And all those kind of issues can

34:03

be put under a security category within

34:05

that data set. So to categorize the rows

34:07

in your data set so that you can pick up

34:10

what changed and compare it with them.

34:12

The second is prompt versioning. Now,

34:14

when you start versioning prompts using

34:16

Git, you know, we all know when you put

34:18

Git message commit messages

34:20

tend to be

34:21

simple commit messages. But you have to

34:24

put governance around what kind of

34:26

commit messages you are putting in when

34:28

you're changing these prompts because

34:29

you need to understand when a prompt was

34:31

changed, for exact what reason it was

34:33

changed, right? What was the failure

34:35

that caused this prompt to be changed?

34:38

What kind of failure would it address

34:40

and what would it correct, right? In the

34:42

next version. That needs to be

34:43

documented.

34:45

Otherwise, it becomes difficult because

34:46

when you go back and look into prompt

34:48

versioning and look at different

34:49

versions and you cannot trace why why

34:51

those changes were made, then it becomes

34:53

difficult to track what's happening.

34:55

The third, the layer three evals, right?

34:58

So the behavioral evals that I was

35:00

talking about around tool calls and

35:01

stuff like that, they can be really

35:03

expensive as you grow your eval data set

35:06

as well. So when you have a wrong tool

35:08

call, for example, and you want to

35:10

correct that system, when you correct it

35:12

and run it against the eval data set,

35:16

you have to basically run it against

35:17

let's say if you have got 300, 400, 500

35:20

rows in the data set, you have to run it

35:21

against them. And you do all the testing

35:23

again and again and again and again and

35:25

again,

35:26

that can cost you a lot of money. So you

35:27

have to put some governance around that.

35:29

So for example, when in your continuous

35:32

integration pipeline, when you do the

35:34

prompt

35:34

change,

35:36

you can actually put

35:39

some checks around

35:41

just just selecting a small subset of

35:43

the eval data set to do the testing. And

35:45

you only do the full test when you merge

35:47

to the main branch.

35:48

So you can put these kind of decisions

35:50

in place so that you can reduce cost

35:52

around

35:53

around

35:55

you know, expensive eval decision.

35:58

If you scan this QR code, it'll take you

36:00

to a Google Drive link where I have put

36:02

some examples on

36:04

some of these how these templates look

36:06

like, what evaluation checklist should

36:08

look like.

36:09

I've given you some guide on

36:11

set setting up tracing

36:13

with open source technologies

36:16

so that you can quickly set up some

36:17

tracing and start testing in the test

36:19

environment before you

36:20

decide on what kind of tools you want to

36:22

use.

36:25

Thank you very much for listening to me.

36:28

This is This QR code will take you to my

36:30

LinkedIn profile. So I share

36:33

I have a newsletter where I share this

36:35

kind of topics every week. So if you're

36:38

interested, you can join. It's free.

36:40

I basically share what I learn in in the

36:42

field working with customers, right? So

36:44

it might be useful for you.

36:46

Thank you very much.

36:47

>> [applause]

More transcripts

Explore other videos transcribed with YouTLDR.

Get the TLDR of any YouTube video

Transcribe, summarize, and repurpose videos in 125+ languages — free, no signup required.

Try YouTLDR Free