# Build Agentic Workflows for Financial Applications

## Abstract

We'll walk you through how to turn financial services workflows into end‑to‑end AI agents, covering orchestration, tools, retrieval, and guardrails tailored to domain-specific applications in the financial services industry (FSI). Using NVIDIA’s NeMo Agent Toolkit and Blueprints, you'll get hands on to building multi‑agent systems that support use cases such as investment research and agentic commerce. Prerequisite(s): Participants should be comfortable with basic Python and familiar with core AI/ML or LLM concepts so they can follow the code and agent patterns efficiently. Experience working with financial services data or workflows (e.g., banking, insurance, payments) is recommended but not required; the material is designed for technical practitioners such as data scientists, ML engineers, and developers who want to build or own FSI agentic AI solutions. Explore training opportunities from the NVIDIA Deep Learning Institute &nbsp;(DLI). Access a wide range of self-paced courses &nbsp;or instructor-led virtual workshops &nbsp;designed to build key skills in AI, HPC, graphics &amp; simulation, and more. &nbsp; Validate your skills with NVIDIA certification &nbsp;and distinguish yourself in the industry.

## AI Summary

- The presentation covered the basics of building agentic workflows for financial services, including agent fundamentals, setup processes, and the use of NeMo Agent Toolkit.
- The course provided an overview of how large language models (LLMs) can be enhanced with tools to perform complex tasks, such as calculating mortgage payments and booking corporate travel.
- The importance of data privacy and security in financial services was emphasized, with solutions like Data Designer and Safe Synthesizer for generating and privatizing data.
- The React Agents design pattern was introduced, which involves a cyclical process of reasoning and tool usage to solve problems more accurately.
- The presentation included examples of multi-agent systems, where different agents work together to perform tasks like investment research and compliance checks.
- Observability tools like Phoenix were discussed to help track and debug the interactions and outputs of agentic workflows.

## Transcript

**[00:00:10 – 00:00:12]** Welcome, thank you so much for coming.

**[00:00:13 – 00:00:14]** Happy second day of GTC.

**[00:00:14 – 00:00:16]** My name is David Williams.

**[00:00:16 – 00:00:18]** I'm a financial services solution architect.

**[00:00:18 – 00:00:22]** This course is Building Agentic Workflows for Financial Services.

**[00:00:22 – 00:00:24]** So today we're gonna be learning exactly about that, how to

**[00:00:24 – 00:00:28]** take agents and make them work on banking, payments,

**[00:00:28 – 00:00:30]** trading use cases.

**[00:00:30 – 00:00:33]** So whether you're brand new to agents, have never

**[00:00:33 – 00:00:35]** built anything before,

**[00:00:35 – 00:00:38]** Whether you've done it a ton, working in the financial industry,

**[00:00:38 – 00:00:40]** hopefully there'll be something you can take away from today.

**[00:00:40 – 00:00:44]** We have a lovely little hour and 45 minutes here together.

**[00:00:44 – 00:00:48]** Alongside myself, I have many other of our wonderful solution architect

**[00:00:48 – 00:00:50]** TAs standing around the room.

**[00:00:50 – 00:00:53]** Anybody in this black polo shirt will be standing

**[00:00:53 – 00:00:55]** clustered through here.

**[00:00:55 – 00:00:57]** You are welcome to ask us anything, any time.

**[00:00:57 – 00:01:02]** The best way to do that so we all don't go crazy is using these

**[00:01:02 – 00:01:04]** cup systems here in front of you.

**[00:01:04 – 00:01:07]** So if you have an issue, rather than holding your hand up and

**[00:01:07 – 00:01:12]** your arm getting tired, please just flip it to the red cup underneath.

**[00:01:12 – 00:01:16]** So not for drinking, more for flagging us down, and

**[00:01:16 – 00:01:19]** somebody will then come quietly in the next couple of minutes

**[00:01:19 – 00:01:22]** and say hi and then flip it back to blue so we know that

**[00:01:22 – 00:01:24]** your question was addressed.

**[00:01:24 – 00:01:26]** Okay?

**[00:01:26 – 00:01:27]** This is how we're getting started.

**[00:01:27 – 00:01:29]** So please, if you haven't done it, go ahead and start

**[00:01:29 – 00:01:30]** doing this right now.

**[00:01:30 – 00:01:33]** You'll have a few minutes to get it sorted while I go through the

**[00:01:33 – 00:01:34]** first beginning lecture of the day.

**[00:01:34 – 00:01:39]** But we're going to go all to learn.nvidia.com

**[00:01:39 – 00:01:41]** slash DLI dash event.

**[00:01:42 – 00:01:44]** In the top right, there'll be a space where you can either

**[00:01:44 – 00:01:49]** create an NVIDIA developer account or log in with your existing one.

**[00:01:49 – 00:01:52]** Probably you have one, I imagine.

**[00:01:52 – 00:01:55]** Make sure you use, if you are creating, make sure you

**[00:01:55 – 00:01:59]** use an email that you can access right now because it's gonna

**[00:01:59 – 00:02:03]** do the two-factor and make you fill out a form and all these things.

**[00:02:03 – 00:02:05]** But whenever you're fully logged in, I've found the

**[00:02:05 – 00:02:09]** best thing is to just go back to that URL a second time,

**[00:02:09 – 00:02:12]** and now you'll see yourself logged in in the top right, and you can

**[00:02:12 – 00:02:18]** enter this lovely, easy-to-remember code, GTC, SJ, 26, T8, 18, 14, X.

**[00:02:19 – 00:02:23]** You can enter that code in the box in the middle, and that will give

**[00:02:23 – 00:02:27]** you access to the course materials for today, access to our hot labs.

**[00:02:27 – 00:02:31]** We're going to be running these off of GPU instances in the cloud

**[00:02:31 – 00:02:35]** that you're going to have access to just through your browser.

**[00:02:35 – 00:02:37]** And then it will also give you access to the material for

**[00:02:37 – 00:02:39]** the next several months as well.

**[00:02:39 – 00:02:42]** So you'll have the ability to, what we strongly encourage

**[00:02:42 – 00:02:45]** is you take away the high-level lessons in this moment.

**[00:02:45 – 00:02:48]** You really understand the core concepts of what we're going,

**[00:02:48 – 00:02:51]** and if you wanna go line by line, understand every little line of

**[00:02:51 – 00:02:55]** code, I love it, you probably will have more time to do that as you

**[00:02:55 – 00:02:57]** re-spin this course up later on.

**[00:02:57 – 00:02:59]** Okay?

**[00:02:59 – 00:03:01]** The only thing to be aware of is if you spin this course

**[00:03:01 – 00:03:03]** up later, it should actually be instantaneous today because

**[00:03:03 – 00:03:04]** we have them pre-spun up.

**[00:03:04 – 00:03:07]** If you spin it up later on, it's going to take about

**[00:03:07 – 00:03:10]** five, ten minutes.

**[00:03:10 – 00:03:15]** Okay, with that, this is our agenda day.

**[00:03:15 – 00:03:18]** I'm sorry I have to keep us moving on here.

**[00:03:18 – 00:03:21]** If you come in late and you need that event code, please

**[00:03:21 – 00:03:22]** just use your red cup.

**[00:03:22 – 00:03:24]** One of our lovely TAs will come bring a piece of paper

**[00:03:24 – 00:03:27]** that has the event code and sign-on instructions to you,

**[00:03:27 – 00:03:29]** and they'll get you started.

**[00:03:29 – 00:03:30]** Here's our agenda for the day.

**[00:03:30 – 00:03:33]** We got a lot to do.

**[00:03:33 – 00:03:35]** First thing, y'all are already killing it with setup.

**[00:03:35 – 00:03:38]** We're gonna talk about some agent fundamentals for folks who

**[00:03:38 – 00:03:42]** are newer to agentic workflows and all of this term, LLMs and agents.

**[00:03:42 – 00:03:43]** We're gonna just cover that.

**[00:03:43 – 00:03:46]** There'll be a slide, or a notebook, the work that you're

**[00:03:46 – 00:03:49]** actually doing, that you're logging into right now, is

**[00:03:49 – 00:03:54]** mostly Jupyter notebooks using our Agentic toolkits and workflows.

**[00:03:54 – 00:03:57]** So we'll just get comfortable with that for the first 30 minutes here.

**[00:03:57 – 00:04:00]** And then the next section will be covering the various

**[00:04:00 – 00:04:04]** aspects of Agentic workflows that we have realized as the

**[00:04:04 – 00:04:08]** US and Europe Financial Services Solution Architect team that we've

**[00:04:08 – 00:04:09]** been working on with customers.

**[00:04:10 – 00:04:12]** What are the extra things that financial institutions

**[00:04:13 – 00:04:16]** need to happen to make agentic workflows work for them?

**[00:04:16 – 00:04:21]** And then, I will be handing it off to Ben, over here in the corner,

**[00:04:21 – 00:04:24]** who will be talking us through an agentic commerce example.

**[00:04:24 – 00:04:27]** And then, Flora, there in the back, who will be taking

**[00:04:27 – 00:04:30]** us through an investment research agentic example for the second

**[00:04:30 – 00:04:33]** half of today's course.

**[00:04:33 – 00:04:36]** And probably, almost certainly, we will not have time to get

**[00:04:36 – 00:04:38]** to factor mining, but if you want to do something even more

**[00:04:39 – 00:04:44]** cool, that will be there for you to go back and take whenever you wish.

**[00:04:45 – 00:04:48]** Okay, let's get going.

**[00:04:48 – 00:04:51]** Again, folks who just came in, please flag down a TA.

**[00:04:51 – 00:04:53]** I see a red cup floating in there.

**[00:04:53 – 00:04:57]** Make sure that we're getting you guys the event code for it.

**[00:04:57 – 00:05:02]** Get you signed up while I talk for the next five, 10 minutes.

**[00:05:03 – 00:05:08]** So, I'll do my best Jensen keynote impression.

**[00:05:08 – 00:05:12]** This is the flow we've been seeing in the artificial intelligence

**[00:05:12 – 00:05:17]** space over the last, what now, decade and a half.

**[00:05:17 – 00:05:19]** The very beginning when we started with modern deep

**[00:05:19 – 00:05:24]** learning, modern AI, was ALEXNET, that was perceptive.

**[00:05:24 – 00:05:27]** First contact perceptive AI, right?

**[00:05:27 – 00:05:29]** Everything we were doing at that time was training deep

**[00:05:29 – 00:05:34]** learning models with large amounts of data to just make a single sort

**[00:05:34 – 00:05:40]** of decision, classification, image recognition sort of problem, right?

**[00:05:40 – 00:05:42]** It was take data, give a score.

**[00:05:42 – 00:05:44]** That's where we started.

**[00:05:44 – 00:05:47]** Well, then coming into 20...

**[00:05:47 – 00:05:52]** 18, 19, 20, 20, that was when the transformer architecture

**[00:05:52 – 00:05:55]** started to become very popular and brought us into this new

**[00:05:55 – 00:05:57]** wave of generative AI.

**[00:05:57 – 00:06:00]** All of the models starting, or the most powerful models

**[00:06:00 – 00:06:03]** starting around that time, were instead about going from a data

**[00:06:03 – 00:06:07]** to a label, they were going from data to new data that could just

**[00:06:07 – 00:06:10]** generate an entirely new thing.

**[00:06:10 – 00:06:13]** Now we find ourselves today square in the middle of the

**[00:06:13 – 00:06:17]** beginning of agentic AI, and this is all about taking those

**[00:06:17 – 00:06:20]** generative models and having them actually use that generative

**[00:06:20 – 00:06:26]** capability to do something for us, to apply a tool to solve a problem

**[00:06:26 – 00:06:30]** a little bit higher level, and this will naturally flow in really

**[00:06:30 – 00:06:33]** nicely into the next step, which is physical AI doing that same

**[00:06:33 – 00:06:39]** process, but in the real world, not being constrained to just digital.

**[00:06:39 – 00:06:43]** If you wanted us to define agentic AI or agents, this

**[00:06:43 – 00:06:44]** is how we would do it.

**[00:06:44 – 00:06:47]** Advanced AI systems designed to autonomously reason, plan,

**[00:06:48 – 00:06:50]** execute complex tasks.

**[00:06:50 – 00:06:56]** That's how we throw it on our lovely little glossary there.

**[00:06:57 – 00:07:01]** So again, generative AI, just recap, how does that work, right?

**[00:07:01 – 00:07:04]** What is a large language model at its most base form?

**[00:07:04 – 00:07:07]** Because that's still the underlying modeling technique of agents.

**[00:07:07 – 00:07:09]** We haven't left LLMs for now.

**[00:07:10 – 00:07:12]** It's the core piece of an agent.

**[00:07:12 – 00:07:16]** Well, a large language model most simply is this large deep learning

**[00:07:16 – 00:07:21]** model that we shove all of the internet's worth of text data into.

**[00:07:21 – 00:07:25]** We train it in an unsupervised manner on all of the human

**[00:07:25 – 00:07:27]** written text that we can find.

**[00:07:27 – 00:07:34]** That teaches the model to read and ultimately generate human text.

**[00:07:34 – 00:07:37]** What that then lets us do is the method of interaction

**[00:07:37 – 00:07:43]** with these generative models was an input and output string of words,

**[00:07:43 – 00:07:45]** or more specifically, tokens.

**[00:07:45 – 00:07:47]** Tokens are just words or part words.

**[00:07:47 – 00:07:50]** You know, we mathematically sometimes split words up to

**[00:07:50 – 00:07:51]** make it easier on us.

**[00:07:51 – 00:07:55]** But you put in words, you put in tokens, and then

**[00:07:55 – 00:07:59]** you get tokens back.

**[00:07:59 – 00:08:01]** New folks still coming in, thank you so much for being

**[00:08:01 – 00:08:04]** here, throw a red cup up for me, one of our TAs will get

**[00:08:04 – 00:08:07]** you set up with the material.

**[00:08:08 – 00:08:10]** So large language models, tokens in, tokens out.

**[00:08:10 – 00:08:12]** What's different about an agent?

**[00:08:12 – 00:08:15]** Well, it's that same large language model, and honestly, we're

**[00:08:15 – 00:08:20]** still putting tokens in and hoping to get tokens out, but instead,

**[00:08:20 – 00:08:23]** we have all of this other stuff.

**[00:08:23 – 00:08:27]** Agents typically have these sort of patterns, and in those

**[00:08:27 – 00:08:31]** patterns they're doing things like advanced high level things that

**[00:08:31 – 00:08:37]** let them do things more complex, like reasoning or tool usage.

**[00:08:38 – 00:08:41]** So you see, reasoning isn't just a fancy word for planning,

**[00:08:41 – 00:08:43]** whereas instead of just being tokens in, immediately throw

**[00:08:43 – 00:08:45]** tokens back to us.

**[00:08:45 – 00:08:48]** We define ways for these models to actually write tokens to themselves

**[00:08:48 – 00:08:51]** to plan out their That's their problem they're going to solve.

**[00:08:51 – 00:08:56]** And tool usage is, hey, some of those internal tokens can

**[00:08:56 – 00:09:02]** now be used to call other code on our behalf.

**[00:09:02 – 00:09:06]** So we're going to create these agentic loops or cycles or

**[00:09:06 – 00:09:09]** workflows that let them do these higher order tasks.

**[00:09:09 – 00:09:12]** So now it can be, hey agent, do a thing for me, and it

**[00:09:12 – 00:09:14]** has to go, okay, here's the parts I'm going to do, and

**[00:09:14 – 00:09:17]** here's the tools I'm going to use it, and go do it, and wait for them

**[00:09:17 – 00:09:20]** to get back, and did I get enough, and okay, yeah, now I can send

**[00:09:20 – 00:09:27]** the tokens back to the end user to say, I did it, here's the answer.

**[00:09:28 – 00:09:31]** And of course, we're really even going to already today

**[00:09:31 – 00:09:34]** be seeing multi-agent workflows where it's not just a single

**[00:09:34 – 00:09:38]** agent, those agents can interact, send tokens in between during that

**[00:09:38 – 00:09:39]** ongoing reasoning and tool usage.

**[00:09:39 – 00:09:44]** They can send tokens to and from other agents to collaborate

**[00:09:44 – 00:09:48]** or divide up problems, and you'll see an example of that today.

**[00:09:49 – 00:09:50]** So how do we do this?

**[00:09:50 – 00:09:52]** This doesn't just happen by magic.

**[00:09:52 – 00:09:56]** It's as much as fun as it is to just ask Cloud Code to just

**[00:09:56 – 00:09:57]** say like, hey, build me an agent.

**[00:09:57 – 00:10:00]** It is helpful to understand what it's actually doing.

**[00:10:00 – 00:10:04]** So we define new libraries, usually in Python, called

**[00:10:05 – 00:10:08]** agentic frameworks to organize all of these communication

**[00:10:08 – 00:10:10]** patterns we want to set up.

**[00:10:10 – 00:10:11]** Here's some of the most common ones.

**[00:10:11 – 00:10:14]** I'd say like by far and away, the best starting point here,

**[00:10:14 – 00:10:19]** like LangChain, LangGraph is probably half of the agentic

**[00:10:19 – 00:10:20]** work being done today.

**[00:10:20 – 00:10:21]** But there's several other of them.

**[00:10:21 – 00:10:24]** Some of them happen directly in the model services.

**[00:10:24 – 00:10:28]** Like if you're using OpenAI's large language models, they

**[00:10:28 – 00:10:31]** have agentic workflow things.

**[00:10:31 – 00:10:35]** We have one we'll be talking about today from NVIDIA called

**[00:10:35 – 00:10:37]** NeMo Agent Toolkit.

**[00:10:37 – 00:10:41]** So, you know how NVIDIA does, like, we want to have something

**[00:10:41 – 00:10:44]** that's well-optimized sitting in almost every space where we know

**[00:10:44 – 00:10:48]** that a GPU is ultimately going to be important, but we also want to

**[00:10:48 – 00:10:50]** bring up the ecosystem as we do it.

**[00:10:50 – 00:10:53]** So, NeMo Agent Toolkit is our agentic framework that

**[00:10:53 – 00:10:57]** we're going to use today that lets us, one, use the most optimized

**[00:10:57 – 00:11:01]** version of all of NVIDIA's models, our Nemotron models and things

**[00:11:01 – 00:11:03]** like that, super, super easily.

**[00:11:03 – 00:11:07]** But it also really plays very well with everything I just

**[00:11:07 – 00:11:09]** had up on the previous slide.

**[00:11:09 – 00:11:14]** In fact, LangChain and NeMo Agent Toolkit have direct tie-ins now.

**[00:11:14 – 00:11:16]** It can wrap around every single one of them.

**[00:11:16 – 00:11:20]** It's a really great tool that just can really fit in with any other

**[00:11:20 – 00:11:23]** agentics pipeline you might set up.

**[00:11:25 – 00:11:28]** NeMo Agent Toolkit gives us the ability to define these

**[00:11:28 – 00:11:32]** workflows through handy-dandy configuration files.

**[00:11:32 – 00:11:36]** So we don't have to, you know, go through a lot of hassle.

**[00:11:36 – 00:11:38]** We can really reuse code super easily.

**[00:11:38 – 00:11:42]** Ultimately, you're going to just see a few times, okay, I have

**[00:11:42 – 00:11:46]** these models and these tools and here's how they're all connected

**[00:11:46 – 00:11:49]** together, basically describing one of those workflow diagrams.

**[00:11:49 – 00:11:53]** That's sort of the coding approach of all of these agentic tools.

**[00:11:53 – 00:11:55]** We're going to define it in a YAML file, and then that

**[00:11:55 – 00:12:00]** YAML file can be ingested by the NeMo Agent Toolkit interface to

**[00:12:00 – 00:12:02]** ultimately do one of three things.

**[00:12:02 – 00:12:05]** It can either run it a single time, deploy it to run all the

**[00:12:05 – 00:12:11]** time, or do an evaluation test run.

**[00:12:14 – 00:12:17]** Here's the formalization of one of the most common ways

**[00:12:17 – 00:12:19]** we build those agentic loops.

**[00:12:19 – 00:12:22]** It's called React Agents.

**[00:12:23 – 00:12:27]** So this is both an agentic pipeline and also a way we

**[00:12:27 – 00:12:30]** have to train our large language models to do that reasoning

**[00:12:30 – 00:12:34]** and thinking looping pattern.

**[00:12:34 – 00:12:36]** The idea of React is we have trained these models and set

**[00:12:36 – 00:12:41]** themselves up to, when you send in tokens, first generate

**[00:12:41 – 00:12:45]** those internal reasoning tokens that say, this is my problem.

**[00:12:45 – 00:12:48]** Here are the steps to find that problem.

**[00:12:48 – 00:12:50]** Which tools or other interfaces do I have

**[00:12:50 – 00:12:52]** that solve one of those steps?

**[00:12:53 – 00:12:54]** And it goes and does that.

**[00:12:54 – 00:12:58]** It can make that tool call through an agentic framework call.

**[00:12:59 – 00:13:01]** And once it gets the data back,

**[00:13:01 – 00:13:03]** It doesn't immediately say, okay, I'm done.

**[00:13:03 – 00:13:04]** Go back and return the output.

**[00:13:05 – 00:13:07]** It says, now that I've gotten something back

**[00:13:07 – 00:13:09]** from a tool, what is next?

**[00:13:09 – 00:13:12]** Do I have enough information to answer my problem, or do

**[00:13:12 – 00:13:14]** I need to do another loop through?

**[00:13:14 – 00:13:18]** That's this sort of cyclical nature that's called React,

**[00:13:18 – 00:13:21]** reasoning and acting.

**[00:13:22 – 00:13:28]** This is one of many design patterns of Agentic AI.

**[00:13:28 – 00:13:30]** There's a really great link here.

**[00:13:30 – 00:13:32]** You should have access to these slides, by the way, on the main

**[00:13:32 – 00:13:35]** page, not in the Jupyter Notebook.

**[00:13:35 – 00:13:38]** Though, actually, we may not have put this one in it.

**[00:13:38 – 00:13:39]** We'll update it later.

**[00:13:39 – 00:13:43]** But you can also just look for Google, did a great blog

**[00:13:43 – 00:13:46]** on design patterns where we got these images from.

**[00:13:46 – 00:13:48]** And these are just some of the many, many ways that you

**[00:13:48 – 00:13:50]** can design these agentic systems.

**[00:13:50 – 00:13:52]** When you start talking about one agent calling tools in

**[00:13:52 – 00:13:54]** a bunch of different ways, do we wanna just loop through it?

**[00:13:54 – 00:13:57]** Do we wanna branching path and try a bunch of things

**[00:13:57 – 00:13:59]** in parallel and then come back?

**[00:13:59 – 00:14:03]** Do we wanna have multiple agents working on subparts of our problem?

**[00:14:03 – 00:14:06]** I wouldn't recommend you go, oh my gosh, this is a hard

**[00:14:06 – 00:14:09]** and fast rule, but just use this as inspiration to think

**[00:14:09 – 00:14:15]** about how can these large language models work with call tools, reason

**[00:14:15 – 00:14:19]** through the steps of a problem.

**[00:14:20 – 00:14:22]** And with that,

**[00:14:22 – 00:14:24]** We are going to go ahead and get started over into our

**[00:14:24 – 00:14:25]** Jupyter Notebooks.

**[00:14:26 – 00:14:28]** There's gonna be two first notebooks for y'all to

**[00:14:28 – 00:14:29]** get going on today.

**[00:14:29 – 00:14:34]** We have an Introduction Notebook and an Agent Fundamentals Notebook.

**[00:14:34 – 00:14:37]** So I haven't even logged in myself yet, so you can follow

**[00:14:37 – 00:14:39]** along with me if you want.

**[00:14:39 – 00:14:41]** But for those of y'all who are logged in, please go

**[00:14:41 – 00:14:42]** ahead and get started.

**[00:14:42 – 00:14:44]** Jupyter Notebook, super easy.

**[00:14:44 – 00:14:46]** You can follow through the Introduction Notebook if you're

**[00:14:46 – 00:14:49]** a little, you know, if you can't remember how to execute

**[00:14:49 – 00:14:51]** cells or anything like that.

**[00:14:51 – 00:14:55]** But then ultimately, I would like over the next 10, 15 minutes

**[00:14:55 – 00:14:59]** for most of us to get through this Agent Fundamentals Notebook, just

**[00:14:59 – 00:15:02]** start to get familiar with NeMo Agent Toolkit and how it works.

**[00:15:02 – 00:15:04]** Okay? Everybody get started.

**[00:15:04 – 00:15:05]** Follow along with me up here.

**[00:15:05 – 00:15:09]** I'll pick the mic back up in 5 or 10 more minutes.

**[00:15:09 – 00:15:12]** Okay.

**[00:15:12 – 00:15:16]** Hopefully folks are getting settled in.

**[00:15:17 – 00:15:21]** I'm going to start casually making our way through this notebook,

**[00:15:21 – 00:15:24]** and then we can start learning about the financial aspects of

**[00:15:24 – 00:15:27]** it, things we're going to change.

**[00:15:33 – 00:15:36]** So again, as I mentioned, if you like reading every

**[00:15:36 – 00:15:38]** line of code, it's there for you.

**[00:15:39 – 00:15:43]** Some things we hid away just to make it not look hideous.

**[00:15:47 – 00:15:50]** OpenAI really set a great standard for what the actual,

**[00:15:50 – 00:15:55]** like, JSON objects or, you know, data packet objects that we send

**[00:15:55 – 00:15:59]** in to large language models, that has become sort of the standard.

**[00:15:59 – 00:16:04]** So even though we in this course today are using a hosted

**[00:16:04 – 00:16:09]** model service by NVIDIA, using our NIM containers, they accept

**[00:16:09 – 00:16:12]** the exact same data packets, and that's broadly true across

**[00:16:12 – 00:16:15]** all models, mostly now.

**[00:16:16 – 00:16:21]** So we can just send in and set up ourselves an endpoint

**[00:16:21 – 00:16:24]** and ask it some questions and get answers back.

**[00:16:24 – 00:16:26]** This is just a standard LLM, right?

**[00:16:26 – 00:16:29]** We haven't done any agentic stuff yet.

**[00:16:29 – 00:16:33]** We just put in tokens and it gave us tokens back based

**[00:16:33 – 00:16:36]** on whatever it had been trained on.

**[00:16:41 – 00:16:45]** LangChain, and in fact, this NVIDIA AI Endpoints extension

**[00:16:45 – 00:16:49]** of LangChain, is where we've integrated those NeMo Agent

**[00:16:49 – 00:16:51]** Toolkit components.

**[00:16:51 – 00:16:54]** So even though we kind of use both today, if you're

**[00:16:54 – 00:16:58]** interested in using LangChain rather than NAT, totally fine.

**[00:16:58 – 00:17:00]** You know, but I would maybe recommend looking at that

**[00:17:00 – 00:17:04]** just to see if there's any features that would be helpful.

**[00:17:10 – 00:17:13]** And you see here that really what LangChain and all agentic

**[00:17:13 – 00:17:17]** frameworks do is they're controlling our prompts

**[00:17:17 – 00:17:20]** and they're controlling our flows.

**[00:17:20 – 00:17:23]** That's kind of what they more or less come down to, is they have

**[00:17:23 – 00:17:27]** a bunch of prompts that say, okay, if I just type in a word, well,

**[00:17:27 – 00:17:29]** here's all the other instructions that we know the language

**[00:17:29 – 00:17:33]** model needs, so we're going to stitch that on on top before

**[00:17:33 – 00:17:36]** handing it to a language model.

**[00:17:37 – 00:17:41]** That's how honestly all of the tool calling and things like that work

**[00:17:41 – 00:17:45]** are just really complex, smartly written prompts that say, here

**[00:17:45 – 00:17:48]** is your tools, here's what they do, here's how you can call them.

**[00:17:48 – 00:17:52]** You just add that as instructions a lot of times.

**[00:17:53 – 00:17:56]** I mean, train a model as well.

**[00:17:59 – 00:18:03]** And so it just really helps us to go to production with

**[00:18:03 – 00:18:06]** this higher level tool where instead of having to like

**[00:18:07 – 00:18:11]** write out prompts and everything every time I can actually

**[00:18:11 – 00:18:16]** build this chain of saying I'm going to do a specific prompt, then

**[00:18:16 – 00:18:21]** the large language model call, then some additional post processing.

**[00:18:21 – 00:18:24]** Now I have a single chain I can call over and over again.

**[00:18:24 – 00:18:28]** You can see how that is a functionality we would, of

**[00:18:28 – 00:18:31]** course, need to start building these agentic loops.

**[00:18:31 – 00:18:35]** We need some sort of pipelining tool.

**[00:18:38 – 00:18:41]** And so here's NAT, NIMO Agent Toolkit.

**[00:18:41 – 00:18:43]** They don't love it when we call it NAT, but I don't know

**[00:18:43 – 00:18:46]** what they were expecting.

**[00:18:49 – 00:18:54]** So NAT, very similarly, lets you build workflows.

**[00:18:54 – 00:18:56]** But again, instead of doing it sort of programmatically

**[00:18:56 – 00:19:00]** in Python, like the default LangChain way is, we like using

**[00:19:00 – 00:19:06]** config files, that makes it super easy to copy and reuse this code.

**[00:19:06 – 00:19:10]** So this one, very sort of similarly, right, is just saying,

**[00:19:10 – 00:19:15]** okay, here's the model name I want to set up, that replaces the sort

**[00:19:15 – 00:19:19]** of section one, where we set up a model, section one, section two.

**[00:19:19 – 00:19:21]** And here's the workflow.

**[00:19:21 – 00:19:24]** This one is just simple, but you could easily, just as easily

**[00:19:24 – 00:19:28]** with a couple more lines, build a chain like we did in section three.

**[00:19:29 – 00:19:32]** So by creating that config,

**[00:19:32 – 00:19:36]** we then can just say, NIMBO Agent Toolkit, run!

**[00:19:36 – 00:19:40]** And it will send something through.

**[00:19:40 – 00:19:43]** Instead, we could also just as easily change one word, change

**[00:19:43 – 00:19:46]** it to deploy, and it will spin it up as a server that's sitting

**[00:19:46 – 00:19:52]** there waiting for us to submit more data in an ongoing fashion.

**[00:19:52 – 00:19:56]** So our question we ask here, we haven't added any tools

**[00:19:56 – 00:20:01]** yet, is just what is the monthly payment on something?

**[00:20:01 – 00:20:04]** On a mortgage?

**[00:20:08 – 00:20:11]** And you can scroll through this if you want, or you can

**[00:20:11 – 00:20:16]** look at our more formatted version.

**[00:20:16 – 00:20:18]** And it does a pretty good job. I mean, honestly, like,

**[00:20:18 – 00:20:21]** it still kind of blows my mind when I think about how this

**[00:20:21 – 00:20:26]** technology is maximum five or six years old, and that already we have

**[00:20:26 – 00:20:30]** trained it with enough knowledge from the internet that it can,

**[00:20:30 – 00:20:31]** it gets what you're talking about.

**[00:20:31 – 00:20:33]** Like, it knows.

**[00:20:34 – 00:20:35]** But is it up to date?

**[00:20:35 – 00:20:36]** Is it accurate?

**[00:20:36 – 00:20:38]** Is it really the best version of this?

**[00:20:38 – 00:20:41]** I don't think that anything that's been pre-trained on historical

**[00:20:41 – 00:20:44]** data and has been probably sitting alone for a month and doesn't pull

**[00:20:44 – 00:20:49]** any of that data in at real time can ever really be truly accurate.

**[00:20:49 – 00:20:53]** So as you see the problems we identified is, okay, it assumed

**[00:20:53 – 00:20:56]** a fixed rate mortgage, assumed no additional fees, insurance.

**[00:20:56 – 00:20:58]** You need a more precise number.

**[00:20:58 – 00:21:01]** You know, these are the things it was telling us

**[00:21:01 – 00:21:02]** about its assumptions.

**[00:21:03 – 00:21:05]** And in fact, of course, there's also an actual bit

**[00:21:05 – 00:21:07]** of math that gets wrong.

**[00:21:07 – 00:21:10]** You can imagine math is very tricky for large language

**[00:21:10 – 00:21:12]** models out of the box.

**[00:21:12 – 00:21:18]** Because that's a very different

**[00:21:18 – 00:21:21]** problem solving than here's a bunch of words, predict

**[00:21:21 – 00:21:24]** the The next word in the sequence.

**[00:21:24 – 00:21:27]** Can you imagine numbers predicting the next numbers in the sequence?

**[00:21:27 – 00:21:30]** There's some underlying logic that we need it to learn.

**[00:21:30 – 00:21:34]** And in this case, yeah, there are ways to train large language

**[00:21:34 – 00:21:36]** models to train them on a bunch of math problems and they

**[00:21:36 – 00:21:40]** can maybe get a little bit better at total prediction for math.

**[00:21:40 – 00:21:43]** But wouldn't we rather just plug this hole with a little

**[00:21:43 – 00:21:46]** calculator and say, hey, here's a thing you can use for the

**[00:21:46 – 00:21:47]** things you're not very good at.

**[00:21:47 – 00:21:50]** That's the idea of adding in tools.

**[00:21:52 – 00:21:55]** And so our last step here I'm gonna walk us through

**[00:21:55 – 00:22:00]** as we can see, where did we go?

**[00:22:03 – 00:22:09]** Simply adding into our configuration file definitions

**[00:22:09 – 00:22:10]** of these mortgage calculators.

**[00:22:10 – 00:22:12]** You can go look through and see how they're defined in

**[00:22:13 – 00:22:16]** code if you want, but ultimately they boil down to we write

**[00:22:16 – 00:22:19]** a bit of code, we write a description of that code, and then

**[00:22:19 – 00:22:23]** we tell Nat where to read that code and where to read that description.

**[00:22:26 – 00:22:29]** And then we add it to our workflow.

**[00:22:29 – 00:22:32]** We wrote some code, we wrote a description of the code,

**[00:22:32 – 00:22:36]** and we added it to the workflow.

**[00:22:36 – 00:22:40]** And this time, we will get a very different result, or

**[00:22:40 – 00:22:43]** a much more accurate result.

**[00:22:43 – 00:22:47]** Where it will first decide and say, hey, I'm doing some math, and I

**[00:22:47 – 00:22:49]** have a thing that does math for me.

**[00:22:49 – 00:22:50]** I should use that instead.

**[00:22:50 – 00:22:54]** You can see here, I need to calculate the monthly

**[00:22:54 – 00:22:56]** mortgage payment and the total interest paid.

**[00:22:56 – 00:23:00]** To do this, I will use the mortgage calculator tool and just

**[00:23:00 – 00:23:03]** give it the details that I extract.

**[00:23:03 – 00:23:06]** That's something now a large language model is really good

**[00:23:06 – 00:23:08]** at, extracting data and just handing it over to something

**[00:23:08 – 00:23:11]** that does it more accurately.

**[00:23:14 – 00:23:19]** Okay, the rest of this, I'm gonna just skip through here.

**[00:23:19 – 00:23:24]** Again, you can set these containers up, or like I said, just very

**[00:23:24 – 00:23:28]** easily say, run this as a service, serve it as an API,

**[00:23:28 – 00:23:31]** single command, NAT serve.

**[00:23:32 – 00:23:34]** You might have to wait a second for it to spin up, but once

**[00:23:34 – 00:23:37]** it's spinning up, you can just submit things to it all day long.

**[00:23:37 – 00:23:40]** But for now, I'm gonna keep us moving.

**[00:23:40 – 00:23:44]** I'm only four minutes behind now.

**[00:23:44 – 00:23:47]** Not too shabby.

**[00:23:48 – 00:23:54]** Let's talk about how to make this work for financial services.

**[00:23:56 – 00:23:58]** So I mentioned that we are all the solution architect

**[00:23:58 – 00:24:02]** team for North America and Europe for financial services.

**[00:24:02 – 00:24:05]** These are some of our customers that we work with and have

**[00:24:05 – 00:24:10]** been public about their adoption of Agentic pipelines and applications.

**[00:24:11 – 00:24:13]** Really across so many different things, these are what our

**[00:24:13 – 00:24:16]** customers are asking us about.

**[00:24:16 – 00:24:19]** Of course, the easiest one to wrap your mind around is

**[00:24:19 – 00:24:20]** customer service.

**[00:24:21 – 00:24:23]** Now, of course, some of this is, yeah, like talk to a chat

**[00:24:23 – 00:24:25]** bot to solve your problem.

**[00:24:25 – 00:24:30]** It is helpful to, you know, I think long-term for human society.

**[00:24:30 – 00:24:32]** We don't necessarily want somebody to just sit there

**[00:24:32 – 00:24:36]** and log into your checking account to check your balance for you.

**[00:24:37 – 00:24:40]** But this is also just improving customer service, giving those

**[00:24:40 – 00:24:43]** customer service agents who have to deal with us calling

**[00:24:43 – 00:24:47]** in all the time access to data, access to our tools a lot

**[00:24:47 – 00:24:50]** faster through agents of their own.

**[00:24:50 – 00:24:52]** And so Capital One is a great example and has been public

**[00:24:52 – 00:24:55]** about the work they're doing there.

**[00:24:56 – 00:25:00]** There's a massive amount of financial documents that dictate

**[00:25:00 – 00:25:03]** financial use cases we work with.

**[00:25:03 – 00:25:07]** So BLACKROCK is a great example where they're using our document

**[00:25:07 – 00:25:13]** processing agentic pipelines to take all of the financial services

**[00:25:13 – 00:25:18]** documents, 10-Ks, 10-Qs, SEC filings, all these things and make

**[00:25:18 – 00:25:22]** them available to their analysts.

**[00:25:22 – 00:25:24]** To all the people who are working inside of the Aladdin

**[00:25:24 – 00:25:30]** platform and the various financial analyst space, equity research,

**[00:25:30 – 00:25:32]** all these things.

**[00:25:32 – 00:25:35]** That extends as well to RBC, who's doing that for equity

**[00:25:35 – 00:25:38]** research specifically, and also building out these multi-agent

**[00:25:38 – 00:25:42]** systems that you'll see today, very similar to what we're doing today.

**[00:25:42 – 00:25:45]** Where we have, hey, this agent go do this sort of analysis, this

**[00:25:45 – 00:25:47]** agent go do this sort of analysis.

**[00:25:47 – 00:25:49]** Bring the results back together, much like a

**[00:25:49 – 00:25:52]** team would do with humans.

**[00:25:52 – 00:25:54]** And then we're also talking about agentic commerce.

**[00:25:54 – 00:25:59]** How can we interact with the payment systems, with all of

**[00:25:59 – 00:26:03]** the decision making that goes into making a purchase so much faster?

**[00:26:03 – 00:26:05]** Does anybody here like?

**[00:26:05 – 00:26:09]** Wish you didn't have to spend an hour booking your flights

**[00:26:09 – 00:26:11]** and figuring out what's the right balance and waiting

**[00:26:11 – 00:26:13]** to see if you get the best price?

**[00:26:13 – 00:26:18]** That's the idea of agentic commerce PayPal is working on.

**[00:26:18 – 00:26:23]** But, we absolutely, thank goodness our team is, I guess,

**[00:26:23 – 00:26:26]** employed, because if this was so simple that they could just use

**[00:26:26 – 00:26:30]** everything, the financial services team wouldn't really need to exist.

**[00:26:30 – 00:26:33]** As soon as all of this stuff got started two or three years

**[00:26:33 – 00:26:36]** ago, we were bombarded as a team with all these questions

**[00:26:37 – 00:26:41]** from those same customers and more about, well, yeah, yeah, this all

**[00:26:41 – 00:26:47]** sounds great in theory, I'd love to apply these things, but this.

**[00:26:47 – 00:26:49]** What about this?

**[00:26:49 – 00:26:53]** We can't pass our data all around the world and go give

**[00:26:53 – 00:26:55]** it to a model service that's running who knows where.

**[00:26:56 – 00:27:00]** We can't pass in data that has PII that hasn't been

**[00:27:00 – 00:27:04]** anonymized properly.

**[00:27:04 – 00:27:08]** We can't have a hallucination that's like funny when I ask

**[00:27:08 – 00:27:10]** it to tell me a joke and it tells me something nonsensical,

**[00:27:10 – 00:27:13]** not funny when it messes up something that gets them I'm

**[00:27:13 – 00:27:15]** in trouble with regulators.

**[00:27:15 – 00:27:19]** We can't certainly have any of our customer data be leaked or be

**[00:27:20 – 00:27:25]** exposed to malicious prompts from outside users and we need to have

**[00:27:25 – 00:27:29]** high, high levels of observability throughout the entire flow so that

**[00:27:30 – 00:27:35]** we know what's happening at every step for that regulatory purpose.

**[00:27:36 – 00:27:37]** So there we go.

**[00:27:37 – 00:27:41]** It's not, you know, this is an evolving field.

**[00:27:41 – 00:27:44]** Don't let me present this as like we have solved all these problems

**[00:27:44 – 00:27:48]** forever, but these are where we are starting and where our customers

**[00:27:48 – 00:27:54]** are modifying this agentic technology to make it work for them

**[00:27:54 – 00:27:57]** in this highly regulated space.

**[00:27:57 – 00:27:59]** So this is exactly what we're gonna go through in notebook number

**[00:27:59 – 00:28:01]** two, is each one of these points, you're gonna learn how to do

**[00:28:01 – 00:28:04]** these things at a very high level.

**[00:28:05 – 00:28:11]** To start with, NIM, this is the solve, this is one of

**[00:28:11 – 00:28:13]** the tools that can help with data sovereignty problems.

**[00:28:13 – 00:28:16]** My compute and models have to come live where my data

**[00:28:16 – 00:28:19]** is, not the other way around.

**[00:28:19 – 00:28:22]** NIM stands for NVIDIA Inference Microservice.

**[00:28:22 – 00:28:28]** This is one of NVIDIA's largest large language model efforts,

**[00:28:28 – 00:28:31]** where what we do is we take these large language models that

**[00:28:31 – 00:28:34]** are popular in the open source,

**[00:28:34 – 00:28:39]** We optimize them for inference and build an inference server

**[00:28:39 – 00:28:42]** because that's just repeated work that you downloading

**[00:28:42 – 00:28:45]** Mistral and you downloading Mistral, like you're both gonna

**[00:28:45 – 00:28:47]** wanna do the same things to run it through an inference toolkit,

**[00:28:47 – 00:28:49]** build a server, build an API.

**[00:28:49 – 00:28:52]** Why don't we just do it once for everybody involved?

**[00:28:52 – 00:28:57]** And then we also deliver that as a container that can be run anywhere.

**[00:28:58 – 00:29:02]** It can be run on any GPUs that are sitting in your on-prem

**[00:29:02 – 00:29:06]** environment, in your own private cloud environment, but it's

**[00:29:06 – 00:29:09]** where you then probably have your data regulated and allowed

**[00:29:09 – 00:29:14]** to be, rather than this floating API service you have to go

**[00:29:14 – 00:29:19]** sign a new contract with to make it meet your regulatory demands.

**[00:29:19 – 00:29:23]** That's what a lot of our customers ask us for with NIM.

**[00:29:23 – 00:29:28]** So that just simply is a way to download and you just point

**[00:29:28 – 00:29:33]** all of your agentic calls to this NIM instead of to a service.

**[00:29:34 – 00:29:36]** NIM is a part of a broader set of tools you've hopefully

**[00:29:36 – 00:29:40]** heard a little bit about now, maybe from Jensen yesterday, called NeMo.

**[00:29:40 – 00:29:45]** NeMo is our end-to-end frameworks and examples for large language

**[00:29:45 – 00:29:48]** models and agentic applications.

**[00:29:48 – 00:29:49]** There's a lot of stuff in NeMo.

**[00:29:49 – 00:29:53]** The main thing we're talking about today is this first box,

**[00:29:53 – 00:29:57]** and actually two of those boxes there too, so like half of it.

**[00:29:58 – 00:30:01]** The starting point is data privacy.

**[00:30:01 – 00:30:04]** And before we can do all this cool stuff we want to do, where

**[00:30:04 – 00:30:07]** we customize at large language models, evaluate them to see

**[00:30:07 – 00:30:10]** if they're doing well and all that, it doesn't matter if you can't get

**[00:30:10 – 00:30:14]** data into the model to train it.

**[00:30:14 – 00:30:19]** Data Designer and Safe Synthesizer are two microservices that

**[00:30:19 – 00:30:22]** we built in collaboration with this company that joined

**[00:30:22 – 00:30:27]** NVIDIA called Gretl, that lets you either privatize or generate

**[00:30:27 – 00:30:29]** entirely new private datasets.

**[00:30:29 – 00:30:33]** So it uses large language models themselves to appropriately

**[00:30:33 – 00:30:37]** clean up data or make fresh new clean data that we can actually

**[00:30:37 – 00:30:40]** use for these agentic pipelines.

**[00:30:41 – 00:30:43]** Specifically, Data Designer is a complete scratch generator.

**[00:30:43 – 00:30:46]** That's what we're using today, where we're just

**[00:30:46 – 00:30:49]** gonna generate a bunch of 10ks.

**[00:30:49 – 00:30:52]** But then probably more of our customers are more interested

**[00:30:52 – 00:30:55]** in the safe synthesizer approach where they want to say, hey,

**[00:30:55 – 00:30:59]** just look through my data and locally, wherever this

**[00:30:59 – 00:31:03]** data is residing, go ahead and turn it into PII safe data.

**[00:31:03 – 00:31:10]** Right, remove all the social security numbers and things.

**[00:31:11 – 00:31:15]** Step number three was hallucinations.

**[00:31:15 – 00:31:19]** The most singular point to avoiding hallucinations is giving

**[00:31:19 – 00:31:24]** large language models and agentic systems access to real-time data.

**[00:31:24 – 00:31:25]** They are just so much better.

**[00:31:25 – 00:31:28]** I mean, you know, me too, quite frankly.

**[00:31:28 – 00:31:31]** It's like, I was so much better in school at an open-book test

**[00:31:31 – 00:31:34]** than a closed-book test, you know?

**[00:31:34 – 00:31:36]** That's when it becomes more about your understanding of

**[00:31:36 – 00:31:40]** what's in front of you, less so your vague recollection

**[00:31:40 – 00:31:44]** of what you learned when you were pre-training a year ago.

**[00:31:44 – 00:31:49]** So NeMo Retriever is a set of microservices and models

**[00:31:49 – 00:31:52]** and pipelinings that does this process called RAG.

**[00:31:52 – 00:31:55]** We've hopefully heard about it a little bit.

**[00:31:55 – 00:31:57]** For folks new to RAG, two stages.

**[00:31:57 – 00:31:58]** First is an ingestion.

**[00:31:58 – 00:32:01]** We're going to take all of our Enterprise documents.

**[00:32:01 – 00:32:04]** We're gonna use some sets of models to chunk them

**[00:32:04 – 00:32:09]** up into small, think like paragraph-sized amounts of text.

**[00:32:09 – 00:32:12]** Now, of course, our documents are not just text.

**[00:32:12 – 00:32:16]** We have additional split up here for pictures, for tables of

**[00:32:16 – 00:32:18]** data, or like text, but not really.

**[00:32:18 – 00:32:20]** You know, we're gonna handle those separately.

**[00:32:20 – 00:32:23]** But either way, we're gonna end up with chunks describing every

**[00:32:23 – 00:32:25]** important thing in a document.

**[00:32:25 – 00:32:29]** And then we're gonna pipe those all into a bunch of vectors.

**[00:32:29 – 00:32:34]** A bunch of learned numbers with another deep learning

**[00:32:34 – 00:32:38]** neural network, large language model, called this Nemotron Embed.

**[00:32:38 – 00:32:42]** And we're going to store those vectors in a vector database.

**[00:32:42 – 00:32:44]** And what's really great about vectors is it's very easy

**[00:32:44 – 00:32:48]** for us to then take our incoming question from a large language

**[00:32:48 – 00:32:54]** model, pass it through the same vectorization process,

**[00:32:54 – 00:32:57]** And we can compare that incoming vector to all the vectors

**[00:32:57 – 00:33:01]** of our database, all the vectors of our enterprise documents.

**[00:33:01 – 00:33:03]** And that helps us give a easy mathematical comparison to

**[00:33:03 – 00:33:07]** say, yes, this is the incoming question, and here's the things

**[00:33:07 – 00:33:12]** in my database that are relevant to that question.

**[00:33:12 – 00:33:16]** Therefore, we can extract the underlying actual chunk, the real

**[00:33:17 – 00:33:20]** chunk, the non-vectorized chunk of data, and give it to our large

**[00:33:20 – 00:33:25]** language model at runtime to say, hey, here's a hint, here's the data

**[00:33:25 – 00:33:28]** that is actually needed to answer this question, and it will do

**[00:33:28 – 00:33:31]** much, much better at answering it.

**[00:33:33 – 00:33:35]** Now we do have, it's unfortunately called

**[00:33:35 – 00:33:37]** safety there, but guardrails here.

**[00:33:38 – 00:33:41]** Fourth step, right, was malicious data leakage,

**[00:33:41 – 00:33:44]** malicious prompting, just all the things we don't want to happen.

**[00:33:45 – 00:33:50]** guardrails is a toolkit in microservice for putting

**[00:33:50 – 00:33:54]** checks before and after large language models.

**[00:33:55 – 00:33:59]** This can check an incoming prompt and say, is this prompt on topic?

**[00:33:59 – 00:34:01]** Is somebody asking something malicious?

**[00:34:02 – 00:34:05]** Is somebody using words we don't want them to use?

**[00:34:05 – 00:34:08]** It sets up the ability to say, hey, do some quick checks,

**[00:34:08 – 00:34:10]** some of them deep learning based, some of them rules

**[00:34:10 – 00:34:12]** engine based, whatever you want.

**[00:34:12 – 00:34:14]** You can set up these checks to check your incoming

**[00:34:14 – 00:34:18]** prompts before it ever hits a large language model.

**[00:34:18 – 00:34:21]** And, likewise, on the way out, we can set up additional

**[00:34:22 – 00:34:25]** checks to check what the language model is sending back to a

**[00:34:25 – 00:34:29]** customer and jump in and cut it off before it gets there.

**[00:34:29 – 00:34:32]** So if we see, hey, it starts to say, even though we tried our

**[00:34:32 – 00:34:35]** best and we said it shouldn't have access to PII data, we're going to

**[00:34:35 – 00:34:38]** do another check to be absolutely sure it's not going to happen.

**[00:34:38 – 00:34:43]** So guardrails is this sort of intermittent safety container

**[00:34:43 – 00:34:46]** that will check everything coming in and out.

**[00:34:47 – 00:34:50]** And finally, observability is solved by just having these

**[00:34:50 – 00:34:52]** hooks into observability tools.

**[00:34:53 – 00:34:56]** You will use one today called Phoenix, but I know like OTel

**[00:34:56 – 00:34:59]** is very popular, Weights and Biases is another one.

**[00:34:59 – 00:35:02]** The idea here is just have something grabbing all of

**[00:35:02 – 00:35:05]** the inputs and outputs, not just to our language models,

**[00:35:05 – 00:35:08]** but also to even the tools we're using, the outside tools.

**[00:35:08 – 00:35:10]** Maybe we have multiple agents.

**[00:35:10 – 00:35:13]** We want to document all of the inputs and outputs as

**[00:35:13 – 00:35:15]** we work through these flows.

**[00:35:15 – 00:35:20]** That way we can say, oh, something did go wrong, here's what happened.

**[00:35:20 – 00:35:21]** Here's where it's going wrong.

**[00:35:22 – 00:35:25]** And you will use this today to learn.

**[00:35:25 – 00:35:29]** With that, it is time for notebook number two.

**[00:35:29 – 00:35:33]** So everybody go ahead and flip back over.

**[00:35:33 – 00:35:37]** Pull-up number two. I'm gonna let y'all process that internally

**[00:35:37 – 00:35:38]** for a couple of minutes.

**[00:35:38 – 00:35:40]** It's very similar to what I just went over in a high level.

**[00:35:40 – 00:35:44]** It just kind of gives you a first example use of each of those tools.

**[00:35:44 – 00:35:47]** So hopefully we'll work through this step in the next 10 or

**[00:35:47 – 00:35:51]** so minutes, and then it'll be time to actually build your

**[00:35:51 – 00:35:55]** own agentic commerce agent next.

**[00:35:57 – 00:36:01]** Okay, last minute here too before I hand over to Ben.

**[00:36:01 – 00:36:03]** So hopefully folks are kinda working through this and you

**[00:36:03 – 00:36:10]** just got to kinda see individual components of it.

**[00:36:10 – 00:36:12]** Rather than walk you through all the same concepts again,

**[00:36:12 – 00:36:15]** the only thing I really wanted to highlight here is Phoenix.

**[00:36:15 – 00:36:18]** This is, I think, some of the coolest stuff and it just

**[00:36:18 – 00:36:21]** gives you that a little bit more picture version of all of this

**[00:36:21 – 00:36:24]** material we're running through.

**[00:36:26 – 00:36:28]** And we'll actually leave this up and come back to it.

**[00:36:28 – 00:36:31]** I would say just leave the tab up when you're done because

**[00:36:31 – 00:36:35]** you can reload this page as we build more agents and you'll see

**[00:36:35 – 00:36:38]** them pop up throughout the course.

**[00:36:38 – 00:36:41]** But you can see here, like one of the ones I was doing,

**[00:36:41 – 00:36:46]** I look over here and I can see like for this, whichever workflow this

**[00:36:46 – 00:36:49]** was, this was, I guess, the very last one I was doing, which was

**[00:36:49 – 00:36:54]** the PII NeMo guardrails agent, just testing to see that that works.

**[00:36:55 – 00:37:00]** We can see that there's React Agent, is then ultimately

**[00:37:00 – 00:37:03]** calling which model it calls.

**[00:37:03 – 00:37:05]** You see, there's the agent.

**[00:37:05 – 00:37:08]** Which model it calls.

**[00:37:08 – 00:37:11]** From that model, that model decided, and you can see even

**[00:37:11 – 00:37:14]** in the output, okay, well I guess this is the overall

**[00:37:14 – 00:37:18]** output, you can see at some point, hey, I know I needed to

**[00:37:18 – 00:37:24]** call this tool of FSI document rag.

**[00:37:25 – 00:37:27]** And then handing that back to the model.

**[00:37:27 – 00:37:29]** I think I didn't get far enough in the guardrails, but I think

**[00:37:29 – 00:37:34]** you're supposed to see it say where it caught it as well.

**[00:37:34 – 00:37:36]** Any questions?

**[00:37:36 – 00:37:40]** Rakesh, do you have a material question or just about, please?

**[00:37:40 – 00:37:45]** I wanted to know how did you get to that URL?

**[00:37:45 – 00:37:47]** So yeah, it's easy to miss.

**[00:37:47 – 00:37:50]** All the way down here when it says traceability with

**[00:37:50 – 00:37:56]** Phoenix, running this cell generates a little URL that

**[00:37:56 – 00:37:58]** says, click here to open Phoenix.

**[00:37:58 – 00:38:00]** That's how you should have that. And so please leave

**[00:38:00 – 00:38:02]** that up for the rest of the time.

**[00:38:02 – 00:38:04]** Because you'll get to see when we are doing much more

**[00:38:04 – 00:38:07]** interesting looking agents.

**[00:38:08 – 00:38:10]** I think, yeah, most of these are, here's sort of our intro

**[00:38:10 – 00:38:13]** one, most of these are just very basic RAG, but this will actually

**[00:38:13 – 00:38:16]** populate with all of the different tools we called, and all of the

**[00:38:16 – 00:38:20]** inputs and outputs to each of those tools, and that is what really

**[00:38:20 – 00:38:23]** helped me understand and learn this for the first time, was really

**[00:38:23 – 00:38:28]** seeing that flowchart in my head of tokens come in, reasoning model

**[00:38:28 – 00:38:32]** thinks this, it calls tool, tool gets this, it sends it back, that's

**[00:38:32 – 00:38:38]** what you can see with these sorts of observability tools very nicely.

**[00:38:39 – 00:38:40]** Team, great job.

**[00:38:40 – 00:38:42]** We are right about halfway here.

**[00:38:42 – 00:38:46]** I will hand it over to Benjamin, who will take us through some

**[00:38:46 – 00:38:49]** agentic commerce examples.

**[00:38:53 – 00:38:55]** Thank you, David.

**[00:38:55 – 00:38:59]** I guess before I jump into the examples,

**[00:38:59 – 00:39:01]** It's so weird hearing your own voice.

**[00:39:01 – 00:39:05]** But any questions on the content so far?

**[00:39:05 – 00:39:11]** So we're going to show examples of kind of semi-realistic

**[00:39:11 – 00:39:16]** financial services workflows, but we'll be using a lot of the tools

**[00:39:16 – 00:39:19]** that we had learned about just now.

**[00:39:19 – 00:39:23]** So if you have any questions on, you know, if like RAG or

**[00:39:23 – 00:39:27]** guardrails or anything like that, now would be a good time to ask or.

**[00:39:27 – 00:39:31]** Just in general, like questions on financial

**[00:39:31 – 00:39:34]** services, agentic workflows.

**[00:39:34 – 00:39:44]** I think there's a microphone, oh, oh no, no, there was no hand.

**[00:39:46 – 00:39:50]** So question on the documents, those were retrieved by the

**[00:39:50 – 00:39:55]** RAG, why are those documents not available in the LLM?

**[00:39:57 – 00:40:03]** Oh, yeah, so this example is using, like, we generated

**[00:40:03 – 00:40:07]** these private financial documents that would live only kind

**[00:40:07 – 00:40:12]** of within the secure institution, and so the LLM would not have seen

**[00:40:12 – 00:40:14]** them or trained on them before.

**[00:40:14 – 00:40:19]** Yeah, and so this is something a lot of financial services

**[00:40:19 – 00:40:26]** institutes would want to keep private from public networks.

**[00:40:31 – 00:40:35]** I work with financial people who are interested in using AI

**[00:40:35 – 00:40:40]** and with the hallucinations and the inference, I see, you know, in the

**[00:40:40 – 00:40:44]** example you generated, financial

**[00:40:44 – 00:40:47]** Information highlights, how can I be sure that that's

**[00:40:48 – 00:40:51]** not a hallucination, because it's an inference, and you

**[00:40:51 – 00:40:56]** did spin up a financial calculator, an FSI calculator in the beginning,

**[00:40:56 – 00:40:59]** but financial people are looking at this and saying, how can

**[00:40:59 – 00:41:02]** you triple check this answer, because he even said it's not

**[00:41:02 – 00:41:06]** a calculator, so how can I trust it with $100 billion, much less

**[00:41:06 – 00:41:11]** a dollar, if it can't calculate, if it's going to inference and guess.

**[00:41:12 – 00:41:14]** Even with the RAG, I mean, I know you want to ground it

**[00:41:14 – 00:41:20]** in, you know, in data, but how can we make sure that these answers are

**[00:41:20 – 00:41:22]** calculated correctly every time?

**[00:41:22 – 00:41:24]** Triple-check your answer.

**[00:41:24 – 00:41:27]** I tried to triple-check my answer and hallucinated three

**[00:41:27 – 00:41:30]** times because the inference was the same all three times, right?

**[00:41:30 – 00:41:34]** So, my question is, how can you make sure this is grounded

**[00:41:34 – 00:41:38]** in reality and doesn't have an inference hallucination

**[00:41:38 – 00:41:42]** in a financial calculation?

**[00:41:42 – 00:41:46]** Yeah, no, that's a really good question, and I think, um,

**[00:41:46 – 00:41:50]** oh, there's microphones too, so, um, yeah, I think hallucinations

**[00:41:50 – 00:41:54]** is definitely one of the things that we want to, you know, deal

**[00:41:54 – 00:41:59]** with in something like as important as our finances, and so RAG is

**[00:41:59 – 00:42:04]** one of the, um, solutions to, like, ground, and then it might not be

**[00:42:04 – 00:42:09]** perfect all the time, uh, there's also, we'll show kind of these, um,

**[00:42:09 – 00:42:13]** You know, different workflows where if it knows it should

**[00:42:13 – 00:42:16]** be, you know, grabbing an actual like deterministic tool,

**[00:42:16 – 00:42:21]** for example, it should do that for like the calculator instead of just

**[00:42:21 – 00:42:26]** guessing, for example, and then I think we didn't go over in this.

**[00:42:26 – 00:42:30]** Of course, but there is, like, ways to evaluate, so if you

**[00:42:30 – 00:42:33]** want to test how well it does, you know, your pipeline is

**[00:42:33 – 00:42:37]** doing, there's evaluation data sets that can be, like, vetted, and you

**[00:42:37 – 00:42:43]** know what the correct answer should be, and that's probably in several

**[00:42:43 – 00:42:46]** other, like, agentic courses, we couldn't get to it in the

**[00:42:46 – 00:42:50]** financial services one, but that's also one, and then another thing,

**[00:42:50 – 00:42:55]** you can even have, like, agentic

**[00:42:55 – 00:42:56]** What is it called?

**[00:42:56 – 00:43:00]** The Agent as a Judge, where you have a much more powerful

**[00:43:00 – 00:43:04]** model that automatically kind of judges if you have a more

**[00:43:04 – 00:43:05]** streamlined model.

**[00:43:05 – 00:43:10]** It's also not perfect, but usually, if you're evaluating

**[00:43:10 – 00:43:13]** spending that many more tokens on a much more powerful model,

**[00:43:13 – 00:43:18]** can catch some of the mistakes that a smaller model could do.

**[00:43:18 – 00:43:21]** So those are just a few out of many, yeah.

**[00:43:21 – 00:43:25]** Oh, there's a, is there a microphone, or?

**[00:43:36 – 00:43:41]** My question is around the synthetic data generator,

**[00:43:41 – 00:43:45]** what kind of data can it generate, I know in this case you're

**[00:43:45 – 00:43:49]** generating some financial documents right, like 10K or something like

**[00:43:49 – 00:43:55]** that, can it generate different types of images, what kind of

**[00:43:55 – 00:43:58]** different modality can it support?

**[00:43:58 – 00:44:00]** Question.

**[00:44:00 – 00:44:03]** Generally right now it's like the particular version of the tool we

**[00:44:03 – 00:44:05]** use as a textual tabular generator.

**[00:44:05 – 00:44:08]** It's been built and trained on how to generate tables of data.

**[00:44:08 – 00:44:11]** You give it the design of what you want that table to look like.

**[00:44:11 – 00:44:13]** You can go back and look at like the config we sent

**[00:44:13 – 00:44:15]** into it if you want.

**[00:44:15 – 00:44:18]** There are plenty of other generative methods that are

**[00:44:18 – 00:44:22]** designed around generating, yeah, generate me a bunch of images,

**[00:44:22 – 00:44:25]** generate me a bunch of videos, generate me a bunch of audio clips.

**[00:44:25 – 00:44:29]** I don't think data designer itself does that.

**[00:44:31 – 00:44:33]** How do you select the embedding model?

**[00:44:34 – 00:44:36]** You know, there are a lot of different choices based

**[00:44:36 – 00:44:44]** on the task and all of NVIDIA's stack here for the most part.

**[00:44:44 – 00:44:48]** In a real-world use case, what is the best way to evaluate the right

**[00:44:48 – 00:44:50]** embedding model for a given task?

**[00:44:50 – 00:44:51]** Yes.

**[00:44:51 – 00:44:55]** Honestly, actually, similar applies to the gentleman in

**[00:44:55 – 00:44:57]** the green hat's question.

**[00:44:57 – 00:44:59]** All of these models are benchmarked.

**[00:44:59 – 00:45:03]** And you go through rigorous evaluation to make sure, like,

**[00:45:03 – 00:45:07]** hey, not just in the general sense, benchmarking helps me give an

**[00:45:07 – 00:45:09]** idea of overall model performance.

**[00:45:10 – 00:45:12]** Is this model better or worse and how are they improving?

**[00:45:12 – 00:45:16]** And then we apply it to a specific use case and evaluate

**[00:45:16 – 00:45:18]** it on test data as well.

**[00:45:18 – 00:45:21]** So we say, okay, we know that this percentage of the time

**[00:45:21 – 00:45:25]** it gives me the right answer, or in the embedding case that

**[00:45:25 – 00:45:29]** it creates a cluster that gives me downstream accuracy.

**[00:45:29 – 00:45:33]** Ours we're using today is NV embed, one NVIDIA trained, that's

**[00:45:33 – 00:45:35]** very high up on the benchmarks.

**[00:45:35 – 00:45:39]** So, thank you. Last question here and we'll keep moving.

**[00:45:39 – 00:45:40]** Yeah.

**[00:45:40 – 00:45:43]** Real quick, Hippie.

**[00:45:43 – 00:45:47]** My question is that do you have a part of the framework

**[00:45:47 – 00:45:53]** where the agent will spin out not a sub-agent but a sub-human that

**[00:45:53 – 00:45:57]** sometimes like a human-in-the-loop part of the framework that

**[00:45:57 – 00:46:03]** would be another part of the pipeline, not just agent to agent?

**[00:46:15 – 00:46:21]** Of course, yeah, but certainly the hooks could be there for that.

**[00:46:23 – 00:46:25]** You absolutely can.

**[00:46:25 – 00:46:29]** The question was why do we have to manually choose which

**[00:46:29 – 00:46:30]** tools are available to it?

**[00:46:30 – 00:46:31]** Sometimes there's access.

**[00:46:31 – 00:46:35]** You don't want an agent to be able to do something until

**[00:46:35 – 00:46:37]** you tell it to do it.

**[00:46:37 – 00:46:41]** Thank you very much.

**[00:46:53 – 00:46:58]** It's fair, and there's really not a huge downside if you

**[00:46:58 – 00:47:01]** had a hundred tools and just you and your institution wanted

**[00:47:01 – 00:47:04]** to have every tool available to every agent, you could do that.

**[00:47:05 – 00:47:07]** Eventually, there might be a little degradation where

**[00:47:07 – 00:47:11]** the more tools you have, the LLM has to decide which tool it's

**[00:47:11 – 00:47:16]** using, but they're so powerful now, I really don't think, unless you

**[00:47:16 – 00:47:19]** had two tools that were very, very similar, I don't think the quantity

**[00:47:19 – 00:47:22]** of tools would be a problem.

**[00:47:22 – 00:47:27]** How is it going to be aware of what the tool actually does?

**[00:47:27 – 00:47:31]** You write the tool code, and then you write in human terms

**[00:47:31 – 00:47:32]** what the tool does.

**[00:47:32 – 00:47:36]** Those are the instructions the LLM will read when it says,

**[00:47:36 – 00:47:38]** hey, is this the tool I should use?

**[00:47:38 – 00:47:42]** So if you write two descriptions when you're setting up that

**[00:47:42 – 00:47:45]** tool that look the same, it won't be able to distinguish

**[00:47:45 – 00:47:45]** what the code does.

**[00:47:45 – 00:47:49]** It's not like it goes and it reads all the tool code, it reads

**[00:47:49 – 00:47:52]** the description we set for it.

**[00:47:53 – 00:47:57]** Similar tools, similarly defined tools are confusing.

**[00:47:57 – 00:47:59]** Thank you for the questions. I really appreciate it.

**[00:47:59 – 00:48:02]** Ben, thanks for letting me jump back in.

**[00:48:02 – 00:48:07]** Yeah, no, so, yeah, if you look at some of the code for

**[00:48:07 – 00:48:13]** the workflows and the tools, you can see the human descriptions of

**[00:48:13 – 00:48:17]** what the tool does, and, yeah, it is kind of mind-blowing, like just

**[00:48:17 – 00:48:22]** describing it in English will allow an agent to, like, know when to use

**[00:48:22 – 00:48:26]** this tool, so that's really cool.

**[00:48:26 – 00:48:31]** And to the question on having a human in the loop, that

**[00:48:31 – 00:48:34]** brings us to, let's see.

**[00:48:37 – 00:48:40]** The next section.

**[00:48:46 – 00:48:50]** The example I'll show is one of these design patterns, which

**[00:48:50 – 00:48:56]** is a single agent plus React, and I'll be applying it to an agentic

**[00:48:56 – 00:49:00]** commerce kind of example use cases.

**[00:49:00 – 00:49:04]** Up here, it's only like five out of many infinite kind

**[00:49:04 – 00:49:06]** of design patterns you can have.

**[00:49:06 – 00:49:11]** And so on this website, there's actually a version of human

**[00:49:11 – 00:49:12]** in the loop as well.

**[00:49:12 – 00:49:15]** So if you need, if there's like a kind of mission critical kind

**[00:49:15 – 00:49:22]** of task that you need a human to check off, you can also have that

**[00:49:22 – 00:49:24]** as part of your workflow as well.

**[00:49:24 – 00:49:31]** So very useful, you know, as these things are quickly evolving.

**[00:49:31 – 00:49:34]** Alright, and so this is the example.

**[00:49:34 – 00:49:40]** How many of you flew into GTC?

**[00:49:40 – 00:49:43]** Okay, maybe like half or so.

**[00:49:43 – 00:49:46]** How many of you booked hotels?

**[00:49:46 – 00:49:52]** Okay, and then how many of you enjoyed that process?

**[00:49:53 – 00:49:58]** Oh, there's one person who enjoyed finding flights and booking hotels.

**[00:49:58 – 00:50:03]** Well, so this example is kind of, wouldn't it be great if you

**[00:50:03 – 00:50:07]** kind of had that automated, this agent kind of can book your hotels

**[00:50:08 – 00:50:12]** for you, book your flights for you, even check your company's travel

**[00:50:12 – 00:50:16]** policy to make sure that you're not breaking any of the rules.

**[00:50:16 – 00:50:18]** And so this is like...

**[00:50:18 – 00:50:21]** Obviously just like a toy example but applied to that use case

**[00:50:21 – 00:50:27]** and you can see how useful like an agent can be for this use case

**[00:50:27 – 00:50:31]** and you can think of so many more cases where you can kind of use a

**[00:50:31 – 00:50:38]** similar design pattern and so this one again is like one React agent,

**[00:50:38 – 00:50:42]** give it a lot of different tools and functions it has access to.

**[00:50:42 – 00:50:48]** And so, this example is like, okay, book a flight from JFK

**[00:50:48 – 00:50:55]** to GTC San Jose, and then we'll use our NeMoTron-powered React

**[00:50:55 – 00:50:59]** Agent, which, again, goes through this thought process, action

**[00:50:59 – 00:51:04]** process, and then observing the outcomes from your tools, and then

**[00:51:04 – 00:51:09]** it can loop through that and figure out when it has completed the task.

**[00:51:09 – 00:51:13]** And so some of the tools in our example that we have

**[00:51:13 – 00:51:17]** given it access to is like a search flight tool.

**[00:51:17 – 00:51:20]** There'll be some fake flight itineraries it can search

**[00:51:20 – 00:51:22]** through, search hotels.

**[00:51:22 – 00:51:26]** It can check your budget, if your company limits how much you

**[00:51:26 – 00:51:28]** can spend on this kind of thing.

**[00:51:28 – 00:51:32]** And then it'll also check your travel policy.

**[00:51:32 – 00:51:35]** And so for checking the travel policy, obviously that's not

**[00:51:36 – 00:51:40]** something an LLM would have been trained on, and you know,

**[00:51:40 – 00:51:44]** company by company, maybe even team by team, it's slightly different.

**[00:51:44 – 00:51:49]** And so we have generated like a fake travel policy that

**[00:51:49 – 00:51:53]** we use through RAG to teach the LLM like what rules, you

**[00:51:53 – 00:51:55]** know, it needs to follow.

**[00:51:55 – 00:51:59]** And so given all of that, then it will like check all of these And

**[00:51:59 – 00:52:02]** then proceed with the book travel.

**[00:52:03 – 00:52:07]** And then either it, you know, everything's like happy path

**[00:52:07 – 00:52:10]** and you get your flights booked and then, you know, you save

**[00:52:10 – 00:52:13]** a few hours of like clicking through all the travel kind of

**[00:52:13 – 00:52:18]** things yourself, or it will reject if you kind of violate some sort

**[00:52:18 – 00:52:21]** of policy and it will let you know.

**[00:52:21 – 00:52:25]** So that's. That's kind of the whole workflow that we're going to be

**[00:52:25 – 00:52:27]** diving into in the next notebook.

**[00:52:27 – 00:52:30]** And then again, we also use like the Phoenix Tracing so

**[00:52:30 – 00:52:34]** that we can follow along the agents like thought process,

**[00:52:34 – 00:52:36]** the inputs and outputs there.

**[00:52:36 – 00:52:41]** Any questions before we go to the notebooks?

**[00:52:45 – 00:52:49]** So I may or may not have built this also because of selfish

**[00:52:49 – 00:52:54]** desires for my own travel.

**[00:52:54 – 00:52:56]** Let's see.

**[00:52:57 – 00:53:02]** Let's get up here.

**[00:53:02 – 00:53:05]** So this is notebook three.

**[00:53:05 – 00:53:08]** And I will give you all a few minutes to go through it, and then

**[00:53:08 – 00:53:09]** we can go through it all together.

**[00:53:10 – 00:53:14]** Thank you so much for watching, and I'll see you in the next episode!

**[00:53:14 – 00:53:18]** Again, this is kind of the diagram of the workflow that

**[00:53:18 – 00:53:21]** we go through in the notebook.

**[00:53:21 – 00:53:26]** And so going through here, we can see that, again, the scenario is

**[00:53:26 – 00:53:28]** you want to book round trip travel.

**[00:53:28 – 00:53:32]** So this includes your flights and hotels.

**[00:53:32 – 00:53:37]** And this is generally the type of, you know, the actions

**[00:53:37 – 00:53:40]** that the React Agent will take.

**[00:53:40 – 00:53:45]** And then the final result will be, if everything is

**[00:53:45 – 00:53:48]** approved and compliant, it will book it automatically

**[00:53:49 – 00:53:52]** using your corporate card.

**[00:53:53 – 00:53:57]** And so stepping through here, one of the interesting

**[00:53:57 – 00:54:01]** documents here is a generated corporate travel policy.

**[00:54:02 – 00:54:05]** So you don't need to read through every single line.

**[00:54:06 – 00:54:10]** But this is something that, you know, if you have a different

**[00:54:10 – 00:54:14]** policy, you can imagine just loading that instead into

**[00:54:14 – 00:54:18]** the RAG, but here are a couple of the highlights.

**[00:54:18 – 00:54:23]** So a few rules are like a maximum

**[00:54:23 – 00:54:28]** cost limit for your domestic flights, maximum hotel nightly

**[00:54:28 – 00:54:31]** limit, unless there's conference.

**[00:54:31 – 00:54:34]** Another one is if you want to fly a business class, you

**[00:54:34 – 00:54:40]** need VP approval, and then you have some meals, also trips

**[00:54:40 – 00:54:46]** above a certain price, and then it'll have some preferred airlines.

**[00:54:46 – 00:54:51]** And then similar to the previous examples, this is how you kind

**[00:54:51 – 00:54:55]** of build your project structure within NeMo Agent Toolkit.

**[00:54:55 – 00:54:59]** You have your basically travel tools and the travel policy here.

**[00:54:59 – 00:55:02]** And then you can.

**[00:55:02 – 00:55:06]** Build it as a package that it can show.

**[00:55:06 – 00:55:12]** So this is just the kind of fake airline and hotel databases,

**[00:55:12 – 00:55:17]** and then hotel, sorry, the employee information.

**[00:55:17 – 00:55:20]** And then you also have the different functions, like

**[00:55:20 – 00:55:24]** checking the budget, extracting the policy, and then booking

**[00:55:24 – 00:55:27]** the travel, et cetera.

**[00:55:28 – 00:55:32]** And then the RAG, again, is it ingests your company travel

**[00:55:32 – 00:55:36]** policy into the vector kind of representation.

**[00:55:37 – 00:55:41]** And then that can be used to help with the generation

**[00:55:41 – 00:55:45]** of much more grounded answers.

**[00:55:46 – 00:55:48]** This is how you register.

**[00:55:48 – 00:55:51]** And then, okay, so for the fun part...

**[00:55:51 – 00:55:54]** This is like the workflow configuration just in a

**[00:55:54 – 00:55:56]** simple YAML file.

**[00:55:56 – 00:55:58]** So we're tracing here with Phoenix.

**[00:55:58 – 00:56:02]** We're using the Nemotron 3 Nano model.

**[00:56:02 – 00:56:05]** We use this embed model for the RAG.

**[00:56:06 – 00:56:09]** And then here are the tools that it has access to.

**[00:56:09 – 00:56:12]** And then this is kind of the general flow that you want

**[00:56:12 – 00:56:15]** the agent to have.

**[00:56:16 – 00:56:20]** And so this first example is, you know, I need to book round-trip

**[00:56:20 – 00:56:26]** travel, give the days, nights, this is your employee number,

**[00:56:26 – 00:56:29]** and looks like everything works.

**[00:56:29 – 00:56:33]** We can look into Phoenix for that, so if you go back into

**[00:56:33 – 00:56:36]** the Phoenix page, you might need to kind of go, I think

**[00:56:36 – 00:56:42]** previously we were here, so if you just go to like the projects page,

**[00:56:42 – 00:56:44]** and you might need to refresh.

**[00:56:44 – 00:56:50]** But your travel booking should be right here.

**[00:56:52 – 00:56:58]** So the first one should be at the bottom.

**[00:57:02 – 00:57:05]** And so, yeah, you can see that the agent, you know,

**[00:57:05 – 00:57:09]** first searches the flights.

**[00:57:11 – 00:57:16]** Searches hotels, checks the budget, checks the travel

**[00:57:16 – 00:57:22]** policy, and then the final output, you know, is this compliant and,

**[00:57:22 – 00:57:25]** you know, something that matches within the budget, matches all the

**[00:57:26 – 00:57:32]** rules, if you got the, these are, you know, made up, but, you know,

**[00:57:32 – 00:57:37]** this one is the San Jose Marriott, and there's some information there.

**[00:57:38 – 00:57:40]** And so this is what it did.

**[00:57:40 – 00:57:43]** And then here's an example where you actually kind of

**[00:57:43 – 00:57:46]** violate the travel policy.

**[00:57:46 – 00:57:51]** So this employee wants a business class flight because we all

**[00:57:51 – 00:57:54]** want to be comfortable at GTC.

**[00:57:55 – 00:58:01]** But let's see how the agent, let's go back here.

**[00:58:01 – 00:58:04]** And so this one.

**[00:58:04 – 00:58:09]** So you can also see like a similar workflow, all the different

**[00:58:09 – 00:58:14]** steps and the tools that the model looked through and then the

**[00:58:14 – 00:58:18]** kind of the tokens that it used.

**[00:58:19 – 00:58:22]** But the final output shows that the requested business

**[00:58:22 – 00:58:27]** flight cannot be booked without VP level approval, so it is

**[00:58:27 – 00:58:31]** aware of this travel policy and it is aware that this

**[00:58:31 – 00:58:35]** request did not have VP approval.

**[00:58:39 – 00:58:43]** Right, and so it comes back with something that is compliant

**[00:58:43 – 00:58:46]** and then asks the user if they want to proceed with that instead.

**[00:58:46 – 00:58:51]** And so that's kind of the workflow there.

**[00:58:53 – 00:58:56]** And then finally, there's one where, you know, someone's

**[00:58:56 – 00:58:59]** not trying to break the rules, but they just didn't have enough

**[00:58:59 – 00:59:02]** money left in their travel budget.

**[00:59:02 – 00:59:07]** So, this person only has like 900 remaining, and it searches through

**[00:59:07 – 00:59:12]** all the various combinations of flights and hotels, but

**[00:59:12 – 00:59:18]** was not able to find this one.

**[00:59:18 – 00:59:22]** The like a compliant option that is under the budget.

**[00:59:22 – 00:59:24]** And so that was the part that was violated.

**[00:59:25 – 00:59:29]** And so it gives like all of the, like an explanation where

**[00:59:29 – 00:59:31]** all of them exceed the budget.

**[00:59:31 – 00:59:35]** And so, yeah, you might need to adjust your dates, reduce

**[00:59:35 – 00:59:38]** the nights, et cetera.

**[00:59:38 – 00:59:41]** So pretty helpful information there.

**[00:59:41 – 00:59:45]** And that kind of wraps up this example, but you can

**[00:59:45 – 00:59:49]** see that, you know, in this very hopefully useful and relatable

**[00:59:49 – 00:59:55]** example of like an agentic commerce for your corporate travel, this

**[00:59:55 – 01:00:01]** React Agent and the proper tools and the RAG are all very useful.

**[01:00:01 – 01:00:02]** And so.

**[01:00:02 – 01:00:05]** Yeah, I think, you know, feel free, I think you can apply

**[01:00:05 – 01:00:09]** this to many other problems, but yeah, I'll wrap here.

**[01:00:09 – 01:00:14]** Any final questions before I hand it off to Flora?

**[01:00:14 – 01:00:15]** Thank you, Ben.

**[01:00:15 – 01:00:17]** And if no other questions, then we're going to move on

**[01:00:17 – 01:00:19]** to our second example.

**[01:00:19 – 01:00:23]** That is an investment research assistant.

**[01:00:23 – 01:00:27]** And this is actually a very broad implementation within

**[01:00:28 – 01:00:30]** the financial services.

**[01:00:30 – 01:00:33]** In this street that our team collaborates with a lot of

**[01:00:33 – 01:00:38]** trading firms that works on building this agentic workflow

**[01:00:38 – 01:00:45]** to mimic their investment research, how they make these decisions.

**[01:00:45 – 01:00:48]** And I kind of want to take a step back because we get

**[01:00:48 – 01:00:52]** these questions a lot after yesterday Jensen's keynote that

**[01:00:52 – 01:00:57]** he highlighted a lot of the work we do in the algo trading space.

**[01:00:58 – 01:01:01]** And you've heard during keynote as well as a lot of mention

**[01:01:01 – 01:01:03]** about this AI factory.

**[01:01:03 – 01:01:08]** When this is applied into the financial services trading space,

**[01:01:08 – 01:01:14]** we have been building AI trading factories with multiple firms.

**[01:01:14 – 01:01:17]** This follows the whole reference architecture of how we built

**[01:01:17 – 01:01:18]** from the full stack bottom I'm going to show you how

**[01:01:18 – 01:01:19]** to create a custom layer.

**[01:01:19 – 01:01:23]** I'm going to show you how to create a custom layer.

**[01:01:23 – 01:01:27]** You probably also heard the five-layer cake from Jensen,

**[01:01:27 – 01:01:32]** all the way from electricity to feeding it into our compute,

**[01:01:32 – 01:01:37]** accelerated networking, storage, to the software layer application

**[01:01:37 – 01:01:41]** models on top of it, till you get the final intelligence

**[01:01:41 – 01:01:46]** tokens and business outcome, in this case, your revenue.

**[01:01:46 – 01:01:51]** AI factory concept is also applied in the trading space, and within

**[01:01:51 – 01:01:55]** Algo Trading, our team actually works with a lot of firms to

**[01:01:55 – 01:01:57]** accelerate the end-to-end workflow.

**[01:01:57 – 01:02:01]** Some of the things you've heard during the keynote yesterday

**[01:02:01 – 01:02:05]** on the accelerated data, to when you have alpha research,

**[01:02:05 – 01:02:09]** how do you do factor search, how do you extract signals,

**[01:02:09 – 01:02:11]** how do you make predictions.

**[01:02:11 – 01:02:15]** To actually deploy it in execution for portfolio optimization, our

**[01:02:15 – 01:02:21]** team also has a Kufolio booth at our show floor that you can go and

**[01:02:21 – 01:02:26]** see our example and our accelerated library for portfolio optimization.

**[01:02:26 – 01:02:30]** This is the AI trading factory and the end-to-end investment

**[01:02:30 – 01:02:34]** workflow that we have been partnering with firms to build

**[01:02:34 – 01:02:36]** out within their enterprise.

**[01:02:36 – 01:02:39]** For today's session, since we're focusing on the agentic

**[01:02:39 – 01:02:44]** part, we're just focusing on how do you construct an

**[01:02:44 – 01:02:49]** investment research team while substituting each current

**[01:02:49 – 01:02:55]** function and role with sub-agents and how these agents work together.

**[01:02:55 – 01:02:59]** In a hierarchical order, so these are, you know, going

**[01:02:59 – 01:03:04]** back to design pattern, kind of like following a parallel agent,

**[01:03:04 – 01:03:10]** but in our actual implementation case, we'll deploy a hierarchical

**[01:03:10 – 01:03:15]** multi-agent system and allows like the thinking reasoning components

**[01:03:15 – 01:03:17]** within with some React agents.

**[01:03:17 – 01:03:20]** So thinking about an investment team, I did hear from the

**[01:03:20 – 01:03:23]** earlier questions that some of you work within the investment

**[01:03:23 – 01:03:27]** space, so you might be familiar with this already, but typically

**[01:03:28 – 01:03:32]** when you think about what we call like a desk, a pod, it typically

**[01:03:32 – 01:03:35]** consists of several different roles that you will have a portfolio

**[01:03:35 – 01:03:41]** manager that owns a portfolio and works with several analysts that

**[01:03:41 – 01:03:46]** look at different type of financial documents, look at different...

**[01:03:46 – 01:03:50]** Signal metrics, some of them focus on a company fundamental,

**[01:03:50 – 01:03:56]** some focuses on more on the quant side or analyzing risk, as

**[01:03:56 – 01:03:58]** well as some technical indicators.

**[01:03:58 – 01:04:03]** There might be specializations within the team too on the

**[01:04:03 – 01:04:07]** sector that they cover or the type of specific information

**[01:04:07 – 01:04:09]** that they look into.

**[01:04:09 – 01:04:14]** And we can kind of like map these roles into each individual

**[01:04:14 – 01:04:19]** AI agent that has the function or tools to call different,

**[01:04:19 – 01:04:24]** for example, if it needs to analyze a certain company's technical

**[01:04:24 – 01:04:30]** indicators here or analyze its most recent new sentiment analysis

**[01:04:30 – 01:04:37]** comments or what is the current consensus of a certain ticker or...

**[01:04:37 – 01:04:41]** Within the portfolio, what might be some risks to watch out for?

**[01:04:41 – 01:04:45]** These different AI agents can work together, and in

**[01:04:45 – 01:04:48]** this example they'll run through, the portfolio manager will

**[01:04:49 – 01:04:53]** act as the orchestrator to then delegate tax to these sub-agents

**[01:04:53 – 01:04:58]** and work together to build an investment report, a synthesized

**[01:04:58 – 01:05:01]** analysis on a certain ticker.

**[01:05:01 – 01:05:05]** So this is the architecture of the example that will walk

**[01:05:05 – 01:05:10]** together in the notebook, but this kind of combines a hierarchical

**[01:05:10 – 01:05:18]** task decomposition pattern plus the reason and act component of it.

**[01:05:18 – 01:05:23]** So when a user asks the question about a ticker, it will feed

**[01:05:23 – 01:05:29]** into the agentic system that then it will call the coordination agent

**[01:05:29 – 01:05:31]** that acts as a portfolio manager.

**[01:05:31 – 01:05:35]** The portfolio manager will reason through what is the

**[01:05:35 – 01:05:39]** ask here, what is what his

**[01:05:40 – 01:05:43]** These are the sub-agents that you can see, and then it will

**[01:05:43 – 01:05:48]** send a query then to each sub-agent to retrieve, use

**[01:05:48 – 01:05:52]** their tool, use their expertise to retrieve the right information, to

**[01:05:52 – 01:05:57]** perform certain analysis, and once all of these are performed, it will

**[01:05:57 – 01:06:00]** save the relevant information back.

**[01:06:00 – 01:06:03]** To the orchestrator agent for it to look through.

**[01:06:03 – 01:06:06]** So these are the technical signals that I get.

**[01:06:06 – 01:06:11]** These are some of the risks that are being flagged.

**[01:06:11 – 01:06:15]** What would be the final recommendation or analysis on,

**[01:06:15 – 01:06:20]** in this example, a ticker that we generated synthetically and then we

**[01:06:20 – 01:06:25]** also feed it into another language model to generate the final

**[01:06:25 – 01:06:31]** report with an executive summary as well as listing now the key

**[01:06:31 – 01:06:34]** Factors that it was able to find and highlighting

**[01:06:34 – 01:06:38]** some key considerations when making the investment.

**[01:06:38 – 01:06:41]** So this is the kind of like the architecture workflow

**[01:06:41 – 01:06:43]** that we're going to run through.

**[01:06:43 – 01:06:48]** But if you're interested in also using this, what NVIDIA does is we

**[01:06:48 – 01:06:51]** publish our reference architecture, what we call Blueprint.

**[01:06:52 – 01:06:55]** And we have an IQ Blueprint that's building an AI agent

**[01:06:55 – 01:06:58]** system for enterprise research.

**[01:06:58 – 01:07:02]** So, a lot of functionality that you can find in this Blueprint is

**[01:07:02 – 01:07:06]** how it connects different agents, leverage some of the reasoning

**[01:07:06 – 01:07:11]** model, being able to connect with the RAG, like what Ben just walked

**[01:07:11 – 01:07:13]** through in the previous notebook.

**[01:07:13 – 01:07:18]** We have this Blueprint deployed on build.nvidia.com, so that is

**[01:07:18 – 01:07:23]** the website we recommend all of you to check out, because we have all

**[01:07:23 – 01:07:25]** of our Blueprints, Models, Library.

**[01:07:25 – 01:07:29]** It's a central hub of what you can leverage from NVIDIA.

**[01:07:29 – 01:07:33]** And we have a dedicated financial services page there too of

**[01:07:33 – 01:07:37]** a collection of available blueprints that you can leverage

**[01:07:37 – 01:07:41]** if you're interested in going deeper after this workshop.

**[01:07:41 – 01:07:48]** So let's move to the notebook four and go into it together.

**[01:07:48 – 01:07:54]** So I am also just running through this notebook.

**[01:07:54 – 01:08:00]** Most of the things about how you define the config, the workflow,

**[01:08:00 – 01:08:06]** I think in previous notebooks you have already went through it.

**[01:08:06 – 01:08:09]** Some of the things about how do you register function as

**[01:08:09 – 01:08:13]** tools, how do you configure it through a YAML file.

**[01:08:13 – 01:08:19]** I'll show you the multi-agent, the hierarchical config, but these

**[01:08:19 – 01:08:24]** will be the things that you'll get to go through within this notebook.

**[01:08:24 – 01:08:29]** Let me show you the configs.

**[01:08:30 – 01:08:37]** Yeah, so this is the orchestrator agent that acts as the portfolio

**[01:08:37 – 01:08:43]** manager, it's config, so it has the telemetry tracing that

**[01:08:43 – 01:08:49]** I'll also show, it's a nice way to visualize through what the agentic

**[01:08:49 – 01:08:53]** system actually run through the tools that it called, then like the

**[01:08:53 – 01:08:59]** language, the LLMs that it called here, the function that it has, so

**[01:08:59 – 01:09:04]** There are several sub-functions of how you can analyze company

**[01:09:04 – 01:09:08]** profile, price level, sentiment.

**[01:09:08 – 01:09:11]** It defines the specialist sub-agents, so you can

**[01:09:11 – 01:09:16]** see here each fundamental agent, technical agent.

**[01:09:16 – 01:09:19]** These are what the portfolio manager, the orchestrator

**[01:09:19 – 01:09:22]** agent have access to.

**[01:09:22 – 01:09:27]** And you can define the full workflow here of...

**[01:09:27 – 01:09:29]** There was a question about what is defined here.

**[01:09:29 – 01:09:35]** This is letting the agentic system know what available tools it has.

**[01:09:35 – 01:09:38]** And there are several following configs.

**[01:09:38 – 01:09:42]** If you go into these individual configs, then, for example,

**[01:09:42 – 01:09:46]** the fundamental agent would also know what type of tools

**[01:09:46 – 01:09:51]** it has as well, and the prompt of what I'm asking it to analyze

**[01:09:51 – 01:09:55]** through a particular ticker.

**[01:09:56 – 01:10:01]** So if you're running through this, this is just setting up, loading

**[01:10:01 – 01:10:04]** the synthetic financial data.

**[01:10:04 – 01:10:09]** These are all companies that we synthetically generated.

**[01:10:09 – 01:10:13]** So hopefully there's not actual company.

**[01:10:13 – 01:10:16]** If we did, it's just coincidence.

**[01:10:16 – 01:10:18]** And defining the agent's data structure.

**[01:10:19 – 01:10:26]** So if you define an agent, there are several ways for you to define

**[01:10:26 – 01:10:29]** for a certain output, because that makes it easier when you have

**[01:10:29 – 01:10:34]** so many sub-agents below to have a consistent way of formatting.

**[01:10:34 – 01:10:39]** So the orchestrator agent will be able to understand

**[01:10:39 – 01:10:43]** following the same structure of knowing What field to fetch

**[01:10:43 – 01:10:46]** or use for its final report.

**[01:10:46 – 01:10:52]** And registering the tools here, so each sub-agent also have

**[01:10:52 – 01:10:55]** the tools that they have access to.

**[01:10:55 – 01:10:59]** Configure the full workflow, I just showed that is the

**[01:10:59 – 01:11:01]** comprehensive agent here.

**[01:11:01 – 01:11:07]** I did some testing on the tools, so randomly just picking one symbol.

**[01:11:07 – 01:11:11]** Just making sure that each of the agents are working fine.

**[01:11:11 – 01:11:14]** It was able to retrieve the information.

**[01:11:15 – 01:11:18]** And finally, on the hierarchical multi-agent workflow, and

**[01:11:18 – 01:11:24]** this is showing you the YAML file that I just shared.

**[01:11:24 – 01:11:26]** Then you can execute the workflow.

**[01:11:26 – 01:11:30]** You'll see the different steps that this model went through,

**[01:11:31 – 01:11:36]** but it's more challenging to read it from the logs because it also

**[01:11:36 – 01:11:40]** contains a lot of the reasoning steps the model went through.

**[01:11:40 – 01:11:46]** I'm going to show you the Phoenix part, so it's a much nicer

**[01:11:46 – 01:11:49]** way to see what actually happened.

**[01:11:49 – 01:11:51]** So, lemme...

**[01:11:51 – 01:11:57]** If you go to the following cell, you can go back to Phoenix,

**[01:11:57 – 01:12:03]** and now you'll be able to see the investment analysis here.

**[01:12:03 – 01:12:07]** Some of the things that I would note for, you know, how to

**[01:12:07 – 01:12:11]** best leverage this is when you...

**[01:12:11 – 01:12:17]** As your agentic system grows and gets more complicated, you will

**[01:12:17 – 01:12:22]** definitely face issues with latency because each agent will be doing

**[01:12:22 – 01:12:24]** its own reasoning, tool calling.

**[01:12:24 – 01:12:30]** What the observability of Phenix helps is to look into not just the

**[01:12:30 – 01:12:34]** latency of each, to find out what might be the bottlenecks within

**[01:12:35 – 01:12:37]** the system, But also like the

**[01:12:38 – 01:12:41]** Different tokens as well as the reasoning trace of it.

**[01:12:41 – 01:12:45]** And that will help in reducing and addressing some of the

**[01:12:45 – 01:12:47]** issues with hallucinations.

**[01:12:47 – 01:12:49]** For some parts, you might want to build a guardrail

**[01:12:49 – 01:12:51]** that's attached to this agent.

**[01:12:51 – 01:12:54]** For some parts, like right in this example, we're all

**[01:12:55 – 01:12:56]** using a reasoning model.

**[01:12:56 – 01:13:00]** But there might be tags that you don't actually need reasoning.

**[01:13:00 – 01:13:03]** So a lot of the optimization work following up from this

**[01:13:03 – 01:13:05]** you can leverage.

**[01:13:05 – 01:13:07]** But here you can see

**[01:13:07 – 01:13:12]** That it called through different sub-agents, it went to the

**[01:13:13 – 01:13:16]** fundamental agent, the fundamental agent looked into the company

**[01:13:16 – 01:13:21]** profile, the financial metrics, it was able to fetch back

**[01:13:21 – 01:13:26]** these numbers, the technical agent also went to its tool

**[01:13:26 – 01:13:32]** to look for technical indicators, the sentiment agent analyzed...

**[01:13:33 – 01:13:39]** The news and was able to pull back like the news and did some analysis

**[01:13:39 – 01:13:43]** on what was the news sentiment

**[01:13:43 – 01:13:47]** and some factor score risk metrics and that compiled into the

**[01:13:48 – 01:13:53]** final agent's response, let me see.

**[01:13:53 – 01:13:58]** Here of what was the input and what was the output of

**[01:13:58 – 01:14:02]** the final analysis for this ticker.

**[01:14:02 – 01:14:05]** There are some reasoning part that the model went through.

**[01:14:05 – 01:14:08]** For example, you can see the reasoning content

**[01:14:08 – 01:14:09]** that you think about.

**[01:14:09 – 01:14:13]** We need to call each tool exactly once in order, passing

**[01:14:13 – 01:14:15]** the stock symbol.

**[01:14:15 – 01:14:19]** So there's more to dive deep into when you're really building

**[01:14:20 – 01:14:22]** and optimizing your agentic system.

**[01:14:22 – 01:14:28]** But the Phoenix tool is a nice way, because sometimes the log might be

**[01:14:28 – 01:14:32]** Hard to see, read through all of the very extensive logs.

**[01:14:32 – 01:14:37]** Then finally, the additional piece that we have is just writing these,

**[01:14:37 – 01:14:43]** sending it back to an LLM to write an investment report, because these

**[01:14:43 – 01:14:48]** might be the final deliverables that you'll actually send to your

**[01:14:48 – 01:14:50]** end user, your profile manager, or.

**[01:14:50 – 01:14:55]** In any type of use case, being able to have a final analysis,

**[01:14:55 – 01:15:01]** compiling the key findings for the ticker and what was analyzed

**[01:15:01 – 01:15:03]** and the final recommendation.

**[01:15:03 – 01:15:08]** So this is kind of like the workflow and now hopefully

**[01:15:08 – 01:15:12]** everyone was able to run through the notebook.

**[01:15:12 – 01:15:15]** I can give you some time, but also as people were running

**[01:15:15 – 01:15:20]** through the notebook, take any questions or thoughts people have.

**[01:15:26 – 01:15:29]** Thank you.

**[01:15:30 – 01:15:37]** If an agent is calling another agent inside it, how is that

**[01:15:37 – 01:15:40]** inside agent configured?

**[01:15:45 – 01:15:47]** So that's a great question.

**[01:15:47 – 01:15:51]** And the nice thing about the NeMo Agent Toolkit is

**[01:15:51 – 01:15:54]** everything is configured within the YAML config file.

**[01:15:54 – 01:15:57]** So you can go into the config.

**[01:15:57 – 01:16:01]** This config file, you can see exactly you follow the

**[01:16:01 – 01:16:05]** same way of defining the LLM, the functions, that's the tools

**[01:16:05 – 01:16:11]** it has access to, as well as the workflow for this specific agent.

**[01:16:11 – 01:16:14]** So all of these are defined in the same way.

**[01:16:14 – 01:16:17]** It's more so that when you define the hierarchical

**[01:16:18 – 01:16:23]** workflow, you define how different agents call each other.

**[01:16:23 – 01:16:26]** But when you define or build a specific agent, you can

**[01:16:26 – 01:16:30]** use the same YAML, that's the YAML config file and follows

**[01:16:31 – 01:16:35]** the same way of how you build agent, regardless of where

**[01:16:35 – 01:16:39]** they are within the agentic system.

**[01:16:46 – 01:16:50]** Right, yeah, in this example, no, but within the NeMo Agent

**[01:16:50 – 01:16:53]** Toolkit, we do support A2A.

**[01:16:54 – 01:16:57]** It's not a true multi-agent system.

**[01:16:57 – 01:16:59]** If there's no agent-to-agent communication, you're just

**[01:16:59 – 01:17:01]** running a task graph.

**[01:17:01 – 01:17:04]** You have a list of tasks, agents are running the tasks,

**[01:17:04 – 01:17:06]** they're not really communicating, they're just finishing the task.

**[01:17:06 – 01:17:10]** So it might be a bit of a misnomer here, because it's

**[01:17:10 – 01:17:13]** not a true multi-agent system based on the kind of stuff

**[01:17:13 – 01:17:15]** that some of us have come.

**[01:17:15 – 01:17:18]** Number two, you mentioned about

**[01:17:18 – 01:17:19]** hallucinations and errors.

**[01:17:20 – 01:17:21]** How are you checking that?

**[01:17:21 – 01:17:23]** How are you checking that you're getting the right information?

**[01:17:24 – 01:17:27]** Because in finance, you make one mistake in the latest information,

**[01:17:27 – 01:17:28]** you get wrong decisions.

**[01:17:28 – 01:17:32]** So there's no way to check anything here, right?

**[01:17:33 – 01:17:34]** But that doesn't make any sense.

**[01:17:34 – 01:17:37]** If you have human in the loop, then this whole thing is lost, right?

**[01:17:37 – 01:17:39]** It's not really a multi-agent system.

**[01:17:40 – 01:17:42]** Neither is it doing any validation.

**[01:17:42 – 01:17:45]** So these are the two important things that I came to learn,

**[01:17:45 – 01:17:46]** but I don't see it in here.

**[01:17:46 – 01:17:52]** So you need to verify that, because I feel like the title is not right.

**[01:17:52 – 01:17:55]** That's a really good feedback, and thank you for flagging it.

**[01:17:55 – 01:17:58]** I'll take it, you know, both your questions.

**[01:17:58 – 01:18:02]** So in this example, we're defining like a hierarchical agentic

**[01:18:02 – 01:18:06]** structure, whereas you have an orchestrator agent, and then calls

**[01:18:06 – 01:18:09]** the sub-agent to perform certain tasks, and then feed it back

**[01:18:09 – 01:18:12]** to the agent for the final output.

**[01:18:12 – 01:18:14]** So this is kind of like the multi-agent layer, a

**[01:18:14 – 01:18:16]** hierarchical structure.

**[01:18:16 – 01:18:19]** While you were mentioning about the A2A, that's definitely things that

**[01:18:19 – 01:18:22]** you can incorporate, unfortunately, and thank you for flagging that

**[01:18:22 – 01:18:26]** we don't have it in this workshop, but I think there will be like, you

**[01:18:26 – 01:18:29]** know, perhaps like other agentic workshops as well as sessions that

**[01:18:29 – 01:18:32]** can cover and go more in depth.

**[01:18:32 – 01:18:37]** I will still say, you know, the agentic structure of how you have

**[01:18:37 – 01:18:41]** different sub-agents, and I think there was a slide about the agentic

**[01:18:41 – 01:18:45]** pattern, how you can, like, define each sub-agent to work together,

**[01:18:45 – 01:18:47]** will still be, you know, a...

**[01:18:47 – 01:18:50]** Starting and even like a medium phase of how you scale out

**[01:18:50 – 01:18:51]** an agentic system.

**[01:18:51 – 01:18:55]** And A2A, as well as incorporating with MCP and all the different

**[01:18:56 – 01:18:59]** tools, there are some recent developments in not just the

**[01:18:59 – 01:19:02]** tool calling, but actually defining the agent skills.

**[01:19:02 – 01:19:05]** All those are like additional advanced topics that

**[01:19:05 – 01:19:07]** you can add to it.

**[01:19:07 – 01:19:08]** And, you know, I...

**[01:19:08 – 01:19:13]** See that you attended this workshop for it, and we didn't have this

**[01:19:13 – 01:19:17]** in the material, but those are also things that the NeMo Agent Toolkit

**[01:19:17 – 01:19:21]** support, and we have examples on our GitHub that you can reference

**[01:19:21 – 01:19:25]** when you want to build out a full-scale A2A multi-agent system.

**[01:19:25 – 01:19:30]** And on the second part with the guardrail, in this example,

**[01:19:30 – 01:19:33]** it was mainly to showcase the hierarchical agentic structure.

**[01:19:33 – 01:19:36]** But I think there were some guardrail components that

**[01:19:36 – 01:19:39]** were highlighted in other examples, as well as to your

**[01:19:39 – 01:19:41]** question, how do we make sure?

**[01:19:41 – 01:19:45]** So a lot of the guardrail to check the hallucination, there was

**[01:19:45 – 01:19:49]** a simple just checking the function output, but you don't want to

**[01:19:49 – 01:19:51]** do it as like a very manual labor.

**[01:19:51 – 01:19:55]** Within the NeMo guardrail, there are several ways

**[01:19:55 – 01:19:56]** that you can check.

**[01:19:56 – 01:19:59]** So one of the example is if it is retrieving any

**[01:19:59 – 01:20:03]** financial metric, you can check it against the RAG system.

**[01:20:03 – 01:20:07]** So you can build a RAG guardrail that checks every information

**[01:20:07 – 01:20:11]** that this model is using, as well as the final output.

**[01:20:11 – 01:20:14]** So, for the NeMo guardrail, you're able to check not just the

**[01:20:14 – 01:20:18]** input, but also the output, as well as all of the following data that

**[01:20:18 – 01:20:21]** was extracted or fed into a model.

**[01:20:21 – 01:20:25]** So, those are NeMo guardrail that if you're interested in, you

**[01:20:25 – 01:20:28]** know, how to prevent hallucination, because that's definitely a very

**[01:20:28 – 01:20:32]** big key, and could compound when you have, you know, multi-agents,

**[01:20:32 – 01:20:36]** like a larger agentic system, I think NeMo guardrail and the

**[01:20:36 – 01:20:38]** concept of how to integrate that.

**[01:20:38 – 01:20:42]** Would be a good topic to look forward and to look further

**[01:20:42 – 01:20:46]** into, but yeah, in this example, it was mainly to showcase, you

**[01:20:46 – 01:20:50]** know, how to build a hierarchical structure as well as taking

**[01:20:50 – 01:20:54]** all of these retrieved information analysis and build it into

**[01:20:54 – 01:20:56]** a synthesized investment report.

**[01:20:57 – 01:21:00]** Hope that answers your question.

**[01:21:03 – 01:21:06]** I would love to sit here and answer every question, but

**[01:21:06 – 01:21:08]** I'm going to go ahead and wrap us first, and then I think

**[01:21:08 – 01:21:12]** several of us can absolutely stick around for a few more minutes.

**[01:21:12 – 01:21:13]** At least I certainly can.

**[01:21:13 – 01:21:16]** Thank you, Flora and Ben.

**[01:21:17 – 01:21:21]** Truthfully, give them the round of applause.

**[01:21:21 – 01:21:23]** Ben and Flora not only came here to present to you today,

**[01:21:23 – 01:21:25]** but they've been the ones working very hard to actually

**[01:21:26 – 01:21:28]** build all of this material.

**[01:21:28 – 01:21:31]** Make sure, you know, pulling all that together, it's been

**[01:21:31 – 01:21:33]** a hectic couple of weeks.

**[01:21:33 – 01:21:36]** And thank you all for making the time to be here with us today.

**[01:21:36 – 01:21:39]** Thank you tech team in the back for handling our mic shuffling.

**[01:21:39 – 01:21:44]** I have one last thing for you.

**[01:21:45 – 01:21:48]** Let's see here, so y'all did a lot today.

**[01:21:48 – 01:21:51]** It was just probably the very beginning of your agentic journey.

**[01:21:51 – 01:21:54]** There's lots of things still to keep learning and adding.

**[01:21:54 – 01:21:56]** Like I said, feel free to come ask us questions for

**[01:21:56 – 01:21:59]** those of us who can stick around for a few more minutes.

**[01:21:59 – 01:22:03]** One slide you don't have in here, we wanted to highlight.

**[01:22:03 – 01:22:09]** Then you wanna, since we're talking open-claw.

**[01:22:10 – 01:22:14]** We would be remiss to talk about agents in the last month or

**[01:22:14 – 01:22:16]** so without talking about OpenCLAW.

**[01:22:16 – 01:22:20]** It's this massive, amazing, again, sort of an agentic

**[01:22:20 – 01:22:23]** framework tool that lets you...

**[01:22:23 – 01:22:26]** Their tag phrase, basically, is like, agents that

**[01:22:26 – 01:22:27]** actually do stuff.

**[01:22:27 – 01:22:30]** And it lets you have these hooks built into, instead

**[01:22:30 – 01:22:33]** of into financial documents and things like that, they provided

**[01:22:33 – 01:22:37]** a really easy interface for you to use your chat with your agent

**[01:22:37 – 01:22:41]** on your phone and provide hooks into, like, your digital life,

**[01:22:41 – 01:22:45]** your text messages, your email, your calendar, all these things.

**[01:22:45 – 01:22:50]** Really really cool. Basically the uptick in usage of it has been

**[01:22:50 – 01:22:52]** Because people really are excited for these agents to

**[01:22:52 – 01:22:55]** do things for them in that space.

**[01:22:56 – 01:23:01]** Of course, it also is very easy to find stories on the

**[01:23:01 – 01:23:03]** internet of OpenClaw deleted everything on my calendar,

**[01:23:03 – 01:23:05]** OpenClaw deleted all my emails.

**[01:23:05 – 01:23:09]** You know, again, this really just highlights the importance

**[01:23:09 – 01:23:11]** of taking our time when we start exposing this to really

**[01:23:11 – 01:23:13]** important pieces of data.

**[01:23:14 – 01:23:17]** So NVIDIA announced this week.

**[01:23:17 – 01:23:21]** So OpenShell, which is more or less just a sandbox environment

**[01:23:21 – 01:23:22]** to run OpenClaw in.

**[01:23:22 – 01:23:26]** So it starts by everything that you might give it access to

**[01:23:26 – 01:23:30]** is a default no, and has additional configuration that's needed

**[01:23:30 – 01:23:32]** to expose it to different things.

**[01:23:33 – 01:23:38]** And then the example of how to run all of that is called NemoClaw.

**[01:23:38 – 01:23:42]** So this link, if you get started, will take you to Brev, our

**[01:23:42 – 01:23:44]** online sort of cloud portal website, and you can see an

**[01:23:44 – 01:23:45]** example of how to get started.

**[01:23:45 – 01:23:47]** You can also go to GTC Park.

**[01:23:47 – 01:23:49]** Where I think all day in the next few days, there are folks

**[01:23:49 – 01:23:51]** helping people get this started and set up if you wanted to

**[01:23:51 – 01:23:53]** even give it a try.

**[01:23:54 – 01:23:58]** There's certainly plenty more blueprints and things.

**[01:23:58 – 01:24:01]** The AIQ one is a really great one for research agents, is

**[01:24:01 – 01:24:03]** a good one to get started.

**[01:24:03 – 01:24:07]** And you can scan this code as well to see whatever else is happening

**[01:24:07 – 01:24:10]** in financial services this week.

**[01:24:10 – 01:24:12]** The last task I'll make of you is please continue to

**[01:24:12 – 01:24:13]** give us feedback.

**[01:24:13 – 01:24:17]** If you go back into the Learn page And you click.

**[01:24:17 – 01:24:20]** Next or something? There's a survey link in there that would be really

**[01:24:20 – 01:24:23]** helpful for you to tell us what worked for you, what you would like

**[01:24:23 – 01:24:26]** to see more of, so we can continue to build these courses for exactly

**[01:24:26 – 01:24:28]** what y'all are looking to learn.

**[01:24:28 – 01:24:30]** With that, thank you so much.

**[01:24:30 – 01:24:31]** Thanks for starting your day with us, and go have

**[01:24:31 – 01:24:32]** an amazing rest of GTC.

