Full Transcript

·YouTLDR

The Self-Improving OSS Agent Stack — Marc Klingen, Langfuse

16:501,241 summary words · ~6 min readEnglishBy AI EngineerTranscribed Oct 8, 2026
Analyze another video with Pro30-day money-back guarantee
Summary

Effective self-improving agents require closing the loop between online production observability and offline evaluation, where human engineers govern dataset boundaries and evaluation criteria while automated agents hill-climb and patch implementations against those benchmarks.

Autonomous agent pipelines risk catastrophic drift and 'slop' without strict human-governed evaluation boundaries, but manual trace inspection creates an unsustainable engineering bottleneck as release velocity accelerates.

Section summaries

0:00-2:00

The Evolution from Static Prompts to Agent Loops

optional

Marc Klingen introduces the macro transition occurring across AI engineering from manual prompt crafting to closed-loop execution stacks. Reflecting on early Langfuse experiments in 2023, he contrasts early single-file HTML generation limitations with today's multi-file coding agents. This capability leap demands a reference architecture tailored to agent dynamics rather than static software.

  • Modern frontier models make continuous optimization loops viable where single-turn prompting failed.
  • Agentic software architectures fundamentally diverge from conventional software by requiring tightly coupled operational feedback.

Provides historical background and framing before diving into the concrete technical architecture.

2:00-4:00

Bridging Online Monitoring with Offline Evaluation

watch

The discussion dissects the traditional tooling divide between online observability systems like Datadog and offline experimentation suites like MLflow. Klingen argues that LLM applications break this boundary because offline datasets drift instantly without production feedback, while live telemetry cannot safely validate new changes. The resulting manual maintenance burden has forced teams to seek autonomous methods to drive the loop across ascending hierarchy levels.

  • Decoupling production APM from offline benchmarking results in evaluating against obsolete user distributions.
  • The hierarchy of loops has scaled from token completion (2023) to autonomous failure reasoning and repair proposals (2025+).

Establishes the core systems-level problem that the self-improving agent stack is designed to solve.

4:00-6:00

Where Humans Must Remain in the Loop

watch

This section isolates the precise boundaries where human engineers must retain authority versus what can be delegated. While autonomous agents are exceptionally effective at hill-climbing prompt variations and context strategies, leaving error classification to agents introduces severe overfitting to operational noise. Humans must explicitly approve data set additions and decide whether observed production quirks represent true scope defects or irrelevant outliers.

  • Autonomous hill-climbing is highly effective when benchmark boundaries and datasets are held constant.
  • Human governance is required to filter out non-critical production quirks and prevent models from overfitting to irrelevant edge cases.

Directly tackles the trade-offs of autonomous optimization versus human-in-the-loop validation.

6:00-8:00

Automating Evaluator and Dataset Synthesis

watch

Klingen explains an emerging workflow where meta-agents scan raw production traces to cluster recurring failure modes and draft corresponding assertions. Examples include detecting competitor mentions, unaligned output languages, or excessive verbosity in customer support systems. Because software requirements emerge incrementally as real users interact with an agent, this loop formalizes implicit requirements into repeatable eval suites under human review.

  • Agents can synthesize new evaluator rules and regression tests directly from clustered production trace failures.
  • AI application requirements are rarely known upfront; they are discovered iteratively by observing live operational errors.

Outlines the specific mechanism for automating evaluation generation from production traffic.

8:00-10:00

The Efficiency Frontier: Manual vs. Full Auto vs. Promised Land

optional

A conceptual framework is presented comparing high-investment manual trace curation against high-risk fully autonomous agent execution. The optimal operational target keeps human engineers anchored strictly at top-level goal-setting while removing them from tedious failure-hunting. By utilizing implicit user signals—such as rejection rates, agent edits, and conversational corrections—the system collects high-fidelity feedback with minimal human overhead.

  • Purely manual curation hits a hard endurance bottleneck, leaving high volumes of production errors uninspected.
  • Implicit telemetry (customer pushback, internal agent overrides, review diffs) provides supervisory signals without manual labeling.

Provides a conceptual matrix comparing engineering investment to agent quality.

10:00-12:00

Case Study: The Langfuse Internal Changelog Agent

watch

To demonstrate the loop, Klingen details Langfuse's internal bottleneck: engineering shipping velocity outpaced documentation and changelog generation. They constructed an autonomous changelog writer that ingests merged pull requests and files documentation pull requests on GitHub. Reviewers review the generated PRs by either approving them or leaving change requests, creating a clean feedback loop captured by their observability pipeline.

  • Automated documentation generation relieves engineering release bottlenecks caused by AI-assisted coding velocity.
  • Standard GitHub PR review workflows (approvals versus requested changes) function as high-signal evaluation checkpoints.

Presents a concrete, relatable production scenario demonstrating the end-to-end feedback architecture.

12:00-14:00

Diagnosing Jargon and Refining Skills

watch

A secondary coding agent inspects the execution traces of the changelog writer alongside human GitHub review comments. It identifies that while technical accuracy was high, end-user clarity was degraded by internal engineering jargon leaking from code comments into public copy. The agent proposes a new evaluator for user-domain language, updates the offline regression suite, and alters the changelog agent's skill prompt to eliminate internal pipeline terminology.

  • Meta-agents can parse human PR review commentary to diagnose subtle quality regressions such as leaking internal jargon.
  • Remediating agent behavior requires updating both the prompt skills and the corresponding automated evaluation assertions.

Shows the step-by-step diagnostic and remediation cycle applied to an agent implementation.

14:00-16:00

Regression Backtesting and Safe Promotion

watch

The candidate implementation (v2) is systematically benchmarked against the baseline implementation (v1) across formatting compliance, factual correctness, and user-facing clarity. Once a positive delta is confirmed without regressions on core assertions, the update is deployed via managed prompt endpoints. Klingen summarizes how top engineering teams run this loop continuously via scheduled cron jobs to process weekly batches of operational telemetry.

  • Candidate prompts must undergo regression backtesting against established baseline datasets before promotion.
  • Scheduled batch runs (e.g., weekly crons) offer an optimal balance between automated iteration and controlled engineering review.

Demonstrates the verification gate required before deploying auto-generated agent modifications.

16:00-16:00

Infrastructure Demands: Scale and Data Ownership

watch

Klingen concludes by examining the structural demands placed on underlying telemetry databases by agentic workloads. Observability is transforming from a write-heavy telemetry sink into a read-heavy querying engine as automated analysis agents repeatedly sweep historical traces. Consequently, teams must own an open-source, cost-effective storage layer that avoids aggressive sampling or short data-retention windows.

  • Self-improving loops flip telemetry workloads from write-heavy logging to read-heavy analytical sweeps.
  • Sampling or short time-to-live (TTL) trace retention breaks long-term agent self-improvement by destroying training and backtesting context.

Delivers crucial data engineering takeaways regarding storage architecture and retention strategies.

Key points

  • Unified Online-Offline Telemetry Loop — Bridging production APM tracing with offline experimentation datasets is mandatory; isolating them leads to either evaluating against stale data or deploying unverified changes directly to production.
  • Human-Bounded Meta Optimization — Engineers must maintain the outer loop by curating ground-truth datasets and failure definitions, leaving prompt tuning, model selection, and context aggregation tactics to automated inner hill-climbing loops.
  • Trace-Driven Eval and Dataset Synthesis — Agents can inspect production traces to extract recurring failure modes, draft synthetic test cases, and propose new domain-specific assertion evaluators for human review.
  • Shift from Write-Heavy Ingestion to Read-Heavy Agent Sweeps — Agent observability infrastructure is shifting from ingest-heavy logging to read-heavy query workloads as autonomous evaluators continuously scan historical traces to diagnose regressions.
“You need to bring the online and the offline together as as otherwise like you either benchmark on data that's inaccurate so the data sets they're not in loop and like uh not in sync with what's happening in production or you monitoring like production data but you don't really benchmark offline.” — Marc Klingen
“If you give this completely out of hand, you risk like that you just create more slop where we've used this in like a blog post of ours where I'd say the target of an AI application is always changing because you don't really know like usually you start with a very high level task.” — Marc Klingen

AI-generated from the transcript. May contain errors.

0:01

[music]

0:12

Everyone, super excited to be here. Ah,

0:15

okay. My laptop was alerting me that

0:17

this starts now. Uh, hi everyone. I'm

0:19

Mark, one of the founders of Langfuse.

0:21

super excited to chat about uh what we

0:23

have seen how how people like

0:24

self-improve agents now and what kind of

0:27

like open source reference stack uh we

0:29

see emerging. I mean obviously my view

0:31

is based on working with our community

0:33

of people that use language to trace and

0:34

evaluate their their agents but I think

0:37

I try to keep the talk mostly to to like

0:39

more generic takeaways that work with

0:41

whatever stacks you're using. Um but

0:43

yeah happy to chat about the different

0:44

pros and cons after the talk. high

0:47

level. I think this is the year uh where

0:50

uh teams talk about like upleveling the

0:53

the level of uh where where they operate

0:55

themselves where increasingly good good

0:57

models just abstract everything

0:58

downstream. I mean coming from I don't

1:00

write prompts I have loops u Peter

1:02

writing about it auto research becoming

1:04

really popular. So I think this is kind

1:06

of where things are going on the like

1:08

like producing um application side but

1:10

also on the how to actually um build

1:13

good good agent applications and um over

1:16

the years really models have have

1:18

expanded a lot in capabilities um back

1:20

when when we started working on length

1:21

fuse this was right when ch was released

1:24

so uh where like even the simplest

1:26

applications didn't really work so for

1:28

example we built like github issue to

1:30

pull request automation similar to what

1:32

devon or kurs are building but back then

1:34

we could only manipulate like single

1:35

page applications like a single HTML

1:37

file because multifile edits were too

1:39

complicated. Since then all of these

1:40

things really accelerated a lot and this

1:42

really made made loops not possible. At

1:44

the same time a more like reference tech

1:46

emerged of how great AI agents are built

1:49

because building agents is different

1:51

from building applications. It's like an

1:54

like a link of like tracing and

1:56

monitoring like online how users really

1:59

use an agent application and then like

2:01

an offline component of how to build

2:03

data sets how to experiment uh regarding

2:06

like making changes to these agents and

2:08

then evaluating offline whether things

2:09

work then deploying new to production

2:10

then learning again for real users use

2:13

the application uh so it's like this

2:14

this kind of loop where traditionally I

2:16

mean you have like obsibility and the

2:18

data docks here in the tracing

2:19

monitoring side and you have like the ML

2:21

ops toolings like MF flow weights and

2:23

biases more on the like offline side

2:25

where this new category emerged and

2:26

that's also why we built length fuse

2:28

because you need to bring the online and

2:29

the offline together as as otherwise

2:32

like you either benchmark on data that's

2:34

inaccurate so the data sets they're not

2:35

in loop and like uh not in sync with

2:37

what's happening in production or you

2:39

monitoring like production data but you

2:40

don't really benchmark offline so you

2:42

need to bring the the both together but

2:44

this has been like a lot of manual labor

2:46

because you need to look at traces

2:48

update these data sets think about like

2:50

new evaluators uh then make these

2:52

changes create new hypothesis like it's

2:54

it's like a lot of work going into this

2:56

process and since starting working on

2:57

language we try to educate teams how to

2:59

how to do this process well but it's

3:01

it's usually like the ask them like oh

3:03

models got better now how can we take

3:05

ourselves out of the loop to uh to

3:07

automate it because it's really tedious

3:08

uh to get things going and that's what I

3:10

want to talk about how we see how people

3:12

go from manually driving this loop to

3:14

using AI to to drive this loop

3:17

um we use the visualization um like of

3:20

for example like the loopcraft article

3:22

or like many many articles in the space

3:23

that like layer different loops on top

3:26

of each other. How I mean at the at the

3:28

lowest level you just have like this

3:29

token loop of what's the next what's the

3:31

next token that's produced and on the

3:32

very high level it's uh like like the

3:34

most abstract form of thinking where you

3:36

as a human are still involved. So what

3:38

we see how um we go through them one by

3:41

one high level I just want to highlight

3:43

these are again possible because models

3:46

improve. So uh when when we started

3:48

working on length in 2023 like I mean we

3:50

had github copilot so like the lowest

3:52

level or the uh level two u then level

3:55

three was possible like more in 2024 and

3:57

now2526 is more like the higher level

3:59

loops of actually agents reasoning about

4:01

what what to even fix uh how to propose

4:04

new fixes and how to how to verify

4:06

whether they actually work. Where humans

4:08

are currently still involved is one you

4:11

need to align data sets of example

4:13

questions of how the agents are used.

4:15

This is so like where where humans need

4:17

to review um like like what what gets

4:20

amended to the data set and how to then

4:22

evaluate whether the agents actually

4:23

work. These are the the two blue things

4:25

in in loop number three. Then on the

4:27

proposal of fixes because if um like

4:29

your agent proposes how to change an

4:31

agent, you still want to see what was

4:33

changed because you have the risk of

4:35

like overfitting it to um to like the

4:38

data set and the failure definition. Um,

4:40

so I mean usually agents find all sorts

4:42

of different uh failure patterns, but

4:45

maybe some of them don't really matter

4:46

that much and you're like, huh, this

4:48

kind of use case for the agent, it

4:49

doesn't really matter that much. Like

4:50

this is like out of what the agent

4:52

should be actually be doing. So we don't

4:53

need to over fit our application to this

4:55

random quirk that we found in

4:56

production, but we want to like be in

4:57

the loop to to monitor this. And now

5:00

coming from this loop pattern, we try to

5:02

dissect it again into more like a

5:03

workflow where where we currently see AI

5:07

used the most is in proposing fixes. So

5:10

if we have production traces and we have

5:12

like evil criteria and data sets of how

5:14

to reproduce issues that we find in

5:16

production, then this how do we hill

5:18

climb against the data sets? That's like

5:20

a really cool uh way to kind of like

5:22

loop your way to a success with with

5:23

agents because usually they're like 10

5:25

different things of what you could be

5:26

doing. I don't know, use a different

5:28

model like aggregate context in a new

5:31

different way. Try whatever you find on

5:32

X uh in a given week to to see whether

5:34

it can like uh make a dent into

5:37

improving the agent. But really like you

5:39

maintain the boundary of what you

5:41

optimize against. And you use AI mostly

5:43

for like proposing improvements to the

5:44

agent implementation.

5:46

Like over the last I'd say month, we see

5:48

more and more teams also using uh like

5:51

like AI and and their loops to maintain

5:53

evil criteria and maintain data sets. So

5:56

uh usually what goes in here is you want

5:58

to like align how a data set uh looks

6:02

like for an agent with what actual users

6:03

are doing. So for example if we have an

6:05

customer support application like

6:07

support agent then the data set should

6:10

be like common support questions but

6:12

like users do all sorts of things with a

6:13

support application. So you need to

6:15

continuously keep it in line with uh

6:17

with what users are doing and you can

6:18

use agents for that. And for EVAs as

6:21

well if you see usual error patterns

6:23

then you can reproduce the error

6:24

patterns and propose new evil criteria

6:26

on these offline data sets. So for

6:28

example I don't know I want I don't want

6:29

to name competitors. I want to be

6:31

concise. I want to answer in the same

6:32

language as the user actually requested.

6:34

Um like the question in my customer

6:36

support application. So you can also

6:37

propose new evaluators. This is I think

6:40

rather newish over the last months. um

6:43

and where we see like the AI more

6:45

autolooping. However, usually teams that

6:48

are still involved in reviewing these

6:49

changes because they create the new

6:51

boundary for how then other agents try

6:54

to auto improve against it. If you give

6:55

this completely out of hand, you risk

6:57

like that you just create more slop

6:59

where we've used this in like a blog

7:01

post of ours where I'd say the target of

7:03

an AI application is always changing

7:05

because you don't really know like

7:07

usually you start with a very high level

7:09

task of I stick with customer support.

7:11

You stick with someone says hey in our

7:14

company customers need to wait a long

7:16

time for support answers and we have a

7:18

lot of people employed to respond to

7:19

them. AI should be able to do this

7:21

automatically or at least eight in

7:22

customer support, but it's a very wake

7:24

target because people don't even know

7:25

what happens in customer support every

7:26

day. So, it's kind of like, oh, we want

7:28

to automate all of it. However, then on

7:30

the way of doing this, you figure things

7:31

out based on like error cases of what

7:33

you even need to do. And that's like

7:35

this kind of like map that you like you

7:36

assemble the plane while you're flying

7:37

it. Um and how we then look at the uh at

7:42

this kind of like looping graphic is you

7:44

really want to be like involved in

7:46

setting the setting like the the goals

7:49

and how you make changes to data sets

7:51

evaluators you want to be in that loop

7:52

involved because you set there by the

7:54

direction and then AI can automate

7:56

against it and then off of the newly

7:58

implemented changes you'll see new error

8:00

classes and you can then like change

8:02

cause again on the most highest level

8:04

but thereby you pull yourself out of a

8:06

lot of like the more manual and tedious

8:09

work on the lower levels.

8:11

Um, this is then the high level view of

8:13

how how I would conceptualize like the

8:15

different scenarios that you can find

8:16

yourself in. I'd say if you implement

8:19

this uh like AI workflow well of tracing

8:23

online uh how your agents are working

8:25

implementing evals on how they are

8:27

working um having data sets uh of

8:30

example questions and then benchmarking

8:32

against them you are in this era in this

8:34

class like very high time invest because

8:35

you need a team who really works with

8:37

this data every week but also quality of

8:39

the of the agents is pretty high I mean

8:41

that's what teams have been for example

8:43

using length use for for the last years

8:44

and this is how you can get to I'd say

8:46

success with your agent application. Um,

8:49

if you fully automate it, I would say

8:51

you don't really need to invest any time

8:52

because you're just like, I don't know,

8:53

codeex goal mode your way to success or

8:55

like you add higher level loops on top

8:57

of it, but you also risk that like it

9:00

produces a lot of tokens, but don't

9:01

really make sense. So, I think you want

9:02

to be in this kind of like promised land

9:04

of you don't really need to invest a lot

9:06

of time, but you're like in the loop at

9:08

the most important steps of like setting

9:11

the direction of where you want to go,

9:13

but you take yourself out of everything

9:14

that's tedious. Thereby you're way less

9:17

time invested than running it manually.

9:18

But also the quality can be even higher

9:20

because usually if you need to do

9:22

everything manually you're bottlenecked

9:24

by your own time your own like

9:25

perseverance endurance uh or like

9:28

patience to look through so much data

9:29

and AI can just look so through like so

9:31

much more error cases. This agent will

9:33

be higher performing and you need less

9:35

time invest. That's the I think that's

9:37

what we were going for and how this then

9:40

looks like is uh you involved at the top

9:42

but also I mean you want ideally your

9:45

application to feed in interesting

9:47

signal. So for example when we talk

9:49

about the like customer support

9:51

application use case usually I mean

9:53

customers can swear at the support agent

9:55

customers can be like oh no you don't

9:57

understand uh like this was not correct

9:59

or like you can propose messages to an

10:02

internal agent who then can accept them

10:03

or change them when sending them out. So

10:05

there's like a lot of like implicit

10:06

signal that can feed into your process

10:08

without you as the agent developer

10:10

needing to do this manually. Uh thereby

10:12

you have like an like an instream of of

10:14

signal that your loop can act on. Um so

10:17

it's either you your team or some signal

10:20

coming in um and the lower level loops

10:22

can be automated. That that's high level

10:24

how we how we think about it. And like

10:26

uh like we put together like a super

10:28

quick demo application um that can that

10:30

can show this. Um so for example like we

10:34

as the language team we have a big

10:36

problem now because our like engineering

10:38

team uses AI a whole lot thus they ship

10:40

a whole lot and somehow we need to

10:42

update our customers uh and user base uh

10:45

what is what is new every week because

10:47

uh like in recent weeks like a lot was

10:49

released uh and usually need to update

10:50

documentation update a change log post

10:53

about it on socials and historically

10:55

like engineers did this themselves

10:56

because they maybe released like a

10:58

bigger feature like once every month so

11:00

it's okay to spend a couple of hours.

11:01

But now if they take themselves out and

11:03

really ship a lot of like product, they

11:06

want to ideally uh like automate also

11:08

that kind of like release process

11:09

because otherwise we are bottlenecked by

11:10

how fast we can communicate about

11:12

changes. So um we thought about like

11:15

putting together like a change lock

11:16

writer that just uploads like like

11:18

updates a lot of like the public assets

11:19

based on what we've shipped. Um however

11:22

then the question is is like how we

11:24

communicate good because if we

11:25

communicate badly about what we've

11:26

released then I mean we we kind of like

11:29

sabotage our releases if we like

11:31

misrepresent for example what has even

11:33

been released. So um setup is we have a

11:36

merge PR, we have a change writer that

11:39

uh suggests a change to documentation,

11:41

but then there's like a just like a code

11:43

review step as like a GitHub PR where

11:45

someone can review and either like

11:47

approve on GitHub and get it merged or

11:50

like submit like change requests of like

11:52

what needs changes in this um in this in

11:56

this change in documentation and and the

11:57

change log. Then the change writer would

11:59

act on this again to update the draft

12:01

until it's finally released. And now

12:03

both of these kinds of like either like

12:05

approvals or uh requests for changes can

12:08

get feed into like the AI obserability

12:10

and eval to to serve as a basis for auto

12:13

improvement of that change writer agent.

12:16

So what we've now done is um just use a

12:20

coding agent pointed at the like length

12:23

use agent skills but generally I mean

12:24

this would also work with other systems

12:26

but yeah length is like well positioned

12:28

for this to to ask like okay what has

12:30

happened in this application how can we

12:32

improve it I'll I'll pause a couple

12:33

times because otherwise it it'll go

12:35

it'll go by very fast. So what we for

12:37

example here identified that the writer

12:40

uh so the chain lock rider is factually

12:44

uh like very correct in how it makes

12:46

changes to documentation but clarity is

12:50

low. the edit ratio of us humans making

12:52

ch like submitting change requests is

12:54

still high and that we leak internal

12:57

jargon to our customers because maybe

12:59

our code has like internal comments that

13:01

we leave there for ourselves for the

13:03

future but uh like our customers don't

13:05

think about our application as like I

13:07

don't know like an ingestion pipeline or

13:09

as like I don't know some kind of like

13:10

batched evil Q like they don't know this

13:12

is like an implementation detail we

13:14

should talk about it as like an user

13:15

domain language and that's like uh

13:17

something we have identified here we're

13:19

Now the agent suggests changes to the

13:24

data sets and evaluators to test for

13:27

leaking jargon and um also like

13:30

increasing clarity in customer domain

13:32

language basically. So um agent now

13:35

closes this gap suggest changes to the

13:39

data sets and evaluators and now um has

13:43

an idea of how to change the agent

13:45

implementation to basically fix this

13:47

problem because in the end it's kind of

13:48

like just like a a problem of how we

13:50

provide context how we like provide the

13:52

skill that then educates this um content

13:54

writer and now we can back test it

13:57

basically on the updated data set to see

13:58

whether the new implementation would do

14:00

better than the old implementation and

14:02

I'll stop here again where we have

14:05

basically the baseline. So um basically

14:08

we reproduce the issue of userf facing

14:11

language. So let's talk about features

14:13

in a way of how users would talk about

14:14

it. So we reproduce this kind of problem

14:17

with the v1 baseline. This basically

14:18

existing implementation of the agent and

14:20

agent came up with a new implementation

14:22

where we see the v2 candidate. We are

14:24

still doing on format compliance and

14:26

accuracy in the same way as we did

14:27

before but we improved on like userf

14:30

facing language. So we see a positive

14:32

delta which sounds like a good change

14:34

now. Um so uh agent reasons about it as

14:38

this. I mean this is a bit I mean

14:40

internal application it's a bit yolo. So

14:42

we are like okay we are good of of just

14:44

merging this and then acting again on

14:46

the next change because this agent never

14:48

publishes anything to actual production.

14:50

It just raises PS on a repo. So it's

14:53

okay if it kind of like auto improves

14:54

and then we just provide feedback again

14:56

to whatever the next increment is. So um

14:59

here we use our like prompt management

15:00

logic so that the agent can just

15:02

basically feature release this new um

15:05

implementation and uh now it just

15:08

summarizes basically um how how we acted

15:10

on this. So high level um this is what

15:14

we have seen like most of our best users

15:17

do already uh when they when they build

15:19

agents like collect production signals

15:21

uh track them alongside like detailed

15:23

execution traces and then run agents

15:25

like on a loop uh usually like a crown

15:27

job like every day every week whenever

15:29

you have like a batch of basically

15:31

interesting user data again to then act

15:32

on it. Um either it's just user feedback

15:35

how how it's happens in this case or

15:36

internal labeling from like a production

15:38

system but it could also be you ask

15:40

people to annotate data like in app for

15:41

example in langu we also allow like for

15:43

annotations to feed into that uh into

15:45

that queue and yeah that's very high

15:47

level what we've seen happy to give you

15:49

like a more deep dive um demo at our at

15:52

our booth or after this talk but just

15:53

wanted to to wrap it up very concisely

15:55

with how we see this this space evolves

15:58

and um I think what's what's interesting

16:00

here is it drives the need for like a

16:02

very scalable data system because you'll

16:04

want to have agents really loop on this

16:06

data and like produce lots of queries.

16:08

This we see how langu was historically

16:10

very right intensive to like um ingest

16:13

evils and traces. Now it's way more read

16:15

intensive because agents can like like

16:17

chew through so much more data.

16:19

interesting number one and two you want

16:21

to really like own the data layer

16:22

because like now the traces of like for

16:24

example a year ago are interesting

16:26

context and you don't want to want them

16:27

to be locked up in like like a more

16:29

commercial system or something where uh

16:31

you need to sample data or retain them

16:33

for a long short time period but we want

16:35

to like retain them for a long time

16:36

period and not sample them and yeah this

16:39

really drives the need for like a

16:40

scalable cheap solution and we've worked

16:42

with many customers on this so happy to

16:43

talk one about uh oneonone about this if

16:46

you're interested thank you so much for

16:47

your

Continue with YouTLDR

Analyze another video with Pro

Process a new video, search every timestamp, compare sources, and keep the result in your library.

Get Pro — $12/month30-day money-back guarantee

More transcripts

Explore other videos transcribed with YouTLDR.