1. Introduction to AI Agents#

A client wants Deere's revenue growth last quarter, compared with Caterpillar, by noon. Build the agent that answers it, design the tools it is allowed to call, and read every step it took.

▶ Open in Colab
/// Toward the goal

Watch the lecture#

The session 1 recording, in two parts. Watch them in order, before or alongside the sections below; the cells on this page are the ones the lecture walks through.

Part 1 · Introduction to AI Agents · 19 min · Open on YouTube
Part 2 · Introduction to AI Agents · 47 min · Open on YouTube · Slides (PDF)

Learning objectives

  • Say what an agent is and tell it apart from a single model call and from a workflow.

  • Trace the model → tool → model cycle on a two-company comparison and name exactly what enters each call.

  • Design a tool schema that a model cannot misuse, and log every step so an analyst can check the memo before it goes to the client.

1.1 What an agent is#

It is Monday morning at Champaign Capital Research and the first item in your inbox is from a portfolio manager at a pension-fund client: Deere’s revenue growth last quarter, compared with Caterpillar, by noon. If you have not read the firm page, do that first; this chapter, and every chapter after it, is one task from that firm’s plan.

Priya’s team answers about thirty requests like this a week. Each one is two lookups in filings the analyst has already read, then a paragraph saying who grew faster and by how much. It takes forty minutes, and the forty minutes is not the hard part of the job. This chapter builds the thing that does the lookups and drafts the paragraph, and leaves the sending to Priya.

An agent is a language model that decides what to do next, does it through a tool, looks at what came back, and repeats until the job is done or a limit stops it.

Every word in that sentence is doing work. Decides: the model, not your code, picks the next step. Through a tool: the model never touches data or systems directly; it asks for a named tool with arguments, and your code runs it. Looks at what came back: the tool’s result goes into the conversation, so the next decision is made with more information than the last. Until a limit stops it: an agent without a budget is a bug, not a feature.

It helps to put the agent next to the two things it is most often confused with.

Who decides the steps

What the model does

Example

Single model call

nobody; there is one step

reads the input, writes the output

“Summarize this earnings call transcript.”

Workflow

your code, in advance

fills in a slot at each fixed step

Extract the vendor, then classify the invoice, then draft an email. Three calls, fixed order.

Agent

the model, as it goes

chooses which tool to call, reads the result, chooses again

“Compare Deere’s revenue growth with Caterpillar’s.” The model decides it needs Deere’s figures, then Caterpillar’s, then writes the comparison.

A workflow is a recipe. An agent is a cook. Most business problems are recipes, and a recipe is cheaper, faster, and easier to test. You reach for an agent when you cannot write the recipe in advance because the steps depend on what the data turns out to say.

Here is the smallest possible demonstration. Ask the model a question it cannot answer from the question alone, and look at what it returns:

from mock_model import model
from agent import explain_reply

reply = model([{"role": "user", "content": "What was Deere's revenue growth last quarter?"}])
print(explain_reply(reply))

That is not an answer. It is a request to run a tool and hand back the result. A single model call stops here. An agent is what happens when something picks that request up, runs it, and asks the model again. That something is the loop, and you will write it in section 1.4.

1.2 The three ingredients#

Every agent, from a forty-line script to a commercial product, is made of the same three parts.

The model. It reads the conversation so far and produces either a tool request or a final answer. It has no memory between calls and no access to anything outside the text it is given. In this book the model is a mock for the in-page cells and a real one in Colab and on the campus copy of this site; the other two ingredients do not change.

The tools. Named functions with a typed contract: a name, a description of when to use them, and a schema for their arguments. The model sees only the contract. Your code owns the implementation, and with it every decision about what the agent is allowed to reach. A read-only get_financials is a very different risk from a send_to_client, and the difference lives entirely in this layer.

The loop. The code that carries messages to the model, executes the tool it asks for, appends the result, and goes again. It is also where the guardrails live: the step budget, the timeout, the log, and the human approval gate for anything that writes.

You already own a version of all three. Hover a part on either figure:

/// Same three parts, two bodies
An equity analyst
HABIT · notice → act → check → again BRAIN decides next move HANDS & SENSES reach, look, read
The agent
LOOP · ask → run tool → append → again MODEL decides next move TOOLS get_financials search_docs
Model ↔ BrainReads what it has so far and decides what to do next. Never touches anything directly.
Tools ↔ Hands & sensesThe only way either one reaches the world. What you give it decides what it can do.
Loop ↔ HabitAct, look at what came back, decide again — until done, or until a limit says stop.
A brain with no hands can't act. Hands with no brain can't decide. Either one without the habit of checking the result never learns it got something wrong.

The split matters because it tells you where to look when something goes wrong. A wrong answer with the right tool calls is a model or prompt problem. A tool called with nonsense arguments is a contract problem. A run that never ends is a loop problem. Chapters 2, 5, and this one map onto those three.

from tools import TOOLS, TOOL_SCHEMAS

print("tools the agent can call:", list(TOOLS))
print()
print("what the model sees for one of them:")
print(TOOL_SCHEMAS[0]["name"], "-", TOOL_SCHEMAS[0]["description"])
print("arguments:", ", ".join(TOOL_SCHEMAS[0]["input_schema"]["properties"]))

Notice what is not in the list: nothing that sends an email, clears a trade, publishes a note, or changes a record. That is a design choice you will make on purpose in 1.9, and chapter 6 is about how to relax it safely.

1.3 When an agent is the right tool#

Because an agent decides its own steps, it costs more per task than a workflow, it is slower, and it is harder to test. Those are real costs, so an agent has to earn its place. Four questions settle it:

  1. Can you write the steps down in advance? If yes, write a workflow. Invoice matching, monthly report generation, and document classification are recipes.

  2. Do the steps depend on what the data says? A client question that might need a filing lookup, a data-vendor pull, a policy check, or none of them, depending on what the client wrote, is agent territory.

  3. What does a mistake cost, and will anyone see it? An agent that recommends and a human who approves is a cheap mistake. An agent that acts on a live system is not. Start with the first.

  4. Is the task worth the latency and the tokens? Three model calls to compare two companies for a client is fine. Three model calls per row of a million-row price table is not.

Chapter 7 turns these four questions into a full decision framework, with the “neither” answer treated seriously. For now the rule of thumb is: workflow by default, agent when the recipe cannot be written, and never let an agent hold a pen until its log has earned your trust.

Checkpoint. One question before moving on.

Every month, Champaign Capital's data team needs each data-vendor invoice extracted into fields, matched to the contract, and flagged if the amounts differ. Which shape fits?

1.4 The loop, in four lines#

Now the third ingredient. Strip away the frameworks and an agent is a loop. You send the conversation to the model. If the reply contains a tool call, you run the tool, append the result to the conversation, and send it again. You stop when the model answers in plain text, or when a step budget runs out.

That is the whole thing. Every agent product you will see this year is this loop plus opinions about what goes into the conversation and which tools exist. Watch one run:

/// Interactive · watch the loop run
Model Tool call? Run tool Append result STEP 0 / 6 0 model calls · 0 tool calls
Press Play. The dot is where the loop is right now.
msgs — what the model sees on the next call

Three things to notice. The model never runs code; it only asks for a tool by name. The loop is the only place with access to real data. And the conversation grows on every step, so whatever the model sees on step three includes everything from steps one and two.

1.5 Try it: the comparison#

This cell runs in your browser against a mock model: a small function that behaves like an LLM for the questions in this book, so no API key is needed. The Colab notebook swaps in a real model (glm-5.3-flash on Lumen) with the same loop.

from agent import agent
from tools import TOOLS

answer, log = agent("Compare Deere's revenue growth with Caterpillar's last quarter", TOOLS)
print()
print("ANSWER:", answer)
print("STEPS :", len(log))

Three steps: look up Deere, look up Caterpillar, write the memo. The model could not write the memo from the question alone, and it could not skip either lookup, because the memo needs both figures. In the Watch view you see the analyst open Firm Records twice, then type the memo and file it.

Now swap Caterpillar for NVIDIA and run again. The model needs a different second lookup, and the loop does not care; it only checks whether each reply is a tool call or text. Then try "What is the current price of NVDA?": a different tool, one step, same loop.

Here is the loop itself, shortened. Read it once slowly; the rest of Part I builds on it.

def agent(question, tools, model, max_steps=6):
    msgs = [{"role": "user", "content": question}]
    for step in range(max_steps):
        reply = model(msgs)                       # 1. ask the model
        if reply["type"] == "text":
            return reply["text"]                  # 2. plain text → done
        result = tools[reply["tool"]](**reply["args"])   # 3. run the tool
        msgs.append({"role": "tool", "content": result}) # 4. append, go again
    return "Stopped: step budget exhausted."

1.6 What enters each call#

The model is stateless. It has no memory of the previous step except what you put in msgs. That makes the messages list the single most important object in the system: it is the model’s entire world.

Print it and look:

from mock_model import model
from tools import TOOLS
from agent import explain_reply, _describe_result

msgs = [{"role": "user", "content": "What was Deere's revenue growth last quarter?"}]
reply = model(msgs)
print("The model's first move:", explain_reply(reply))

result = TOOLS[reply["tool"]](**reply["args"])
msgs.append({"role": "assistant", "content": f"[{reply['tool']}]"})
msgs.append({"role": "tool", "name": reply["tool"], "content": result})

print("\nWhat the model sees on the NEXT call, one line per message:")
for m in msgs:
    text = _describe_result(m["content"]) if m["role"] == "tool" else str(m["content"])
    print(" ", m["role"].ljust(9), text)

This is what “transparent” means in this chapter. A transparent agent is one where you can print the messages list at any step and every line is something a human put there or a tool returned. Nothing is hidden inside a framework object.

The agent() function in this book returns a log alongside the answer: one dict per step with the tool name, arguments, result, and elapsed time. In a business setting that log is the audit trail. When the client asks “where did that figure come from?”, you open the log, not the model.

from agent import agent, narrate

answer, log = agent("Compare Deere's revenue growth with Caterpillar's last quarter", verbose=False)
print(answer)
print()
for line in narrate(log):
    print(line)

1.7 Designing the tool API#

The loop is simple. The hard part is the contract the model sees: the tool’s name, description, and input schema. That contract is the only thing that stands between a well-behaved model and one that calls search_docs("") forty times.

Open the schemas this book ships with:

from tools import TOOL_SCHEMAS, describe_schema
describe_schema(TOOL_SCHEMAS[0])

Four rules, each of which fixes a failure you will otherwise see in week one:

Enums over free text. ticker is an enum of five symbols, not a string. A string invites "Deere", "DE Corp", and "deere.com". An enum makes the invalid call impossible instead of merely unlikely.

Required means required. period is required even though it has one value today. The day a second quarter is added, the model must choose rather than default silently.

Describe when not to call. The get_financials description says “Use only when the user asks about revenue, growth, or earnings.” Models over-call tools that sound useful. The negative instruction is usually more valuable than the positive one.

Close the schema. additionalProperties: false rejects arguments you did not define. Without it, a model that invents {"ticker": "DE", "currency": "EUR"} gets a silent success and a wrong answer.

Try breaking one. The cell below validates a call against the schema with a tiny checker (the real API does this for you when you set strict: true, which you will use in the notebook):

from tools import TOOL_SCHEMAS

def validate(schema, args):
    props, req = schema["input_schema"]["properties"], schema["input_schema"]["required"]
    problems = [f"missing required '{r}'" for r in req if r not in args]
    problems += [f"unexpected '{k}'" for k in args if k not in props]
    for k, v in args.items():
        if k in props and "enum" in props[k] and v not in props[k]["enum"]:
            problems.append(f"'{k}'={v!r} not in {props[k]['enum']}")
        if k in props and "minLength" in props[k] and len(v) < props[k]["minLength"]:
            problems.append(f"'{k}' shorter than {props[k]['minLength']}")
    return problems or ["ok"]

fin = TOOL_SCHEMAS[0]; search = TOOL_SCHEMAS[3]
print(", ".join(validate(fin, {"ticker": "DE", "period": "Q2-2026"})))
print(", ".join(validate(fin, {"ticker": "Deere"})))
print(", ".join(validate(fin, {"ticker": "DE", "period": "Q2-2026", "currency": "EUR"})))
print(", ".join(validate(search, {"query": ""})))

Checkpoint.

The model keeps calling get_financials with ticker="Deere Corp" and getting errors back. Where is the fix most likely to live?

1.8 Two failure modes and their guards#

The runaway loop. A model that keeps calling tools never returns text. max_steps is the guard. Six is a good default for a single question; set it from the task, not from hope. When the budget is exhausted the loop returns a message that says so, and the log records it. Never let the loop end silently.

The bad call. The model asks for a tool with arguments that crash it. The loop in this book catches the exception and returns {"error": ...} as the tool result. The model then sees the error on the next step and can correct. A crash tells the model nothing; an error message is data.

from agent import agent
from tools import TOOLS

# a model that never stops calling tools
def stubborn(msgs):
    return {"type": "tool_call", "tool": "get_price", "args": {"ticker": "DE"}}

answer, log = agent("anything", TOOLS, model=stubborn, max_steps=3)
print()
print(answer)

1.9 A live business example: the client comparison, by noon#

Everything so far used one question because the numbers are easy to check. Here is the same loop doing the job the firm actually wants automated: the comparison memo a client is waiting for.

The scenario. Thirty times a week someone asks Priya’s team how one company’s last quarter compares with another’s. The figures are in filings the analyst has already read. Finding them, checking them, and writing the paragraph takes forty minutes, and two analysts answering the same request have been known to lead with different numbers. The firm’s research-process rule says every figure in anything that leaves the building must cite its source, so the paragraph also has to say where the numbers came from.

The agent gets one tool, get_financials, which returns one company’s quarter. It does not get a send_to_client tool. It drafts; Priya reads the memo and the log, then sends. That split is the single most important design decision in the example, and it is a tool-API decision, not a prompt decision.

from agent import agent

answer, log = agent("Compare Deere's revenue growth with Caterpillar's last quarter")
print()
print(answer)

Read the steps. Two lookups, in the order the question named the companies, then a memo with both figures, the gap, and a source tag. The model could not have written the memo from the question, and it could not have skipped a lookup, because the memo needs both numbers.

Now change the pair. Each one in the fixture data makes a different point:

Pair

What is different

What the memo should say

Deere vs Caterpillar

same sector; Deere is smaller and growing faster

the growth gap, the size gap, both figures

NVIDIA vs Apple

NVIDIA is half the size and growing ten times faster

lead with growth, note the base

Microsoft vs Apple

Apple is larger; Microsoft is growing three times faster

size and growth point different ways

Deere vs Tesla

the firm holds no data for Tesla

the tool returns an error; the agent must say so, not guess

from agent import agent

for pair in ["Deere and Caterpillar", "NVIDIA and Apple", "Microsoft and Apple"]:
    answer, log = agent(f"Compare revenue growth for {pair} last quarter", verbose=False)
    print(f"{pair}: {answer}\n")

Every memo ends with a [source: …] tag. Chapter 3 makes that mandatory and makes the tag point at a document. For now, notice what it buys you: a client who questions a number can be shown where it came from, not argued with.

The fourth pair is the one that matters most. The client asks about a company the firm does not cover:

from agent import agent

answer, log = agent("Compare Deere's revenue growth with Tesla's")
print()
print(answer)

The tool returned an error instead of a number, the loop passed the error to the model as data (section 1.8), and the model said it could not complete the comparison. A model without the tool would have written a confident paragraph with a made-up figure in it. The error is the feature.

The log is what Priya reads before she sends. This is what you would store per request, and what you would show a client who asks how the firm produces its numbers:

from agent import agent, narrate

answer, log = agent("Compare Deere's revenue growth with Caterpillar's last quarter", verbose=False)
for line in narrate(log):
    print(line)

What the firm gets from this loop, compared with an analyst doing it by hand: the same figures pulled the same way every time, a memo with a source on every number, a log that can be reviewed, and an analyst who now reads thirty drafts instead of researching thirty requests. What it does not get is an agent that talks to clients. That stays behind a human click until the log has earned trust, which is the subject of chapter 6.

The Colab notebook runs this exact comparison against the real model. Compare its tool sequence with the mock’s; a well-designed schema should make them match.

1.10 Run it against a real model#

The mock model answered every question so far. This section sends the same loop, the same tool schemas, and the same client comparison to a GPT deployment on Illinois Azure.

Your browser never sees a key. The page calls /api/chat on this site, a small proxy that holds the key, checks that you are signed in with your campus account, and forwards the request. That is the same “no write tools, a human approves” idea from 1.9 applied to the key itself: the model is reachable, the credential is not.

Model cells need the campus copy of this book: open it there and sign in with your @illinois.edu account. Everything else on this page works here.
from agent import agent
from llm import azure_model      # real model, via /api/chat on this site

answer, log = agent("Compare Deere's revenue growth with Caterpillar's last quarter", model=azure_model)
print()
print(answer)

Compare with the mock’s run in 1.9. Both lookups should be there, because the schema only returns one company at a time. The wording will differ; the two figures should not.

Now the question the mock could never handle, because it only knows the scripts in this book:

from agent import agent
from llm import azure_model

answer, log = agent("Which of Deere, Caterpillar and NVIDIA grew revenue fastest last quarter, and is the fastest also the largest? Two sentences for a client.", model=azure_model)
print()
print(answer)

Read the log. Did the model look up all three, and nothing else? Did the memo cite where the figures came from? Every extra tool call costs tokens and time, so a good loop is not the one that calls the most tools, it is the one that calls exactly the ones the question needs. Your token budget for the day is in the header of this page.

1.11 Exercise#

Open the Colab notebook. It contains the same loop, wired to glm-5.3-flash on Lumen (see Setup for the key) with strict: true schemas.

  1. Classify three tasks from your own work or internship as single call, workflow, or agent, using the four questions in 1.3. One sentence each.

  2. Run the Deere-vs-Caterpillar comparison against the real model and compare its tool sequence and memo with the mock’s. Did it ask for both companies in one reply or one at a time?

  3. Add a second data tool, get_price(ticker), with a closed schema, and run “Is Deere’s revenue growth better than Caterpillar’s, and what are both trading at?” It should show four tool calls and one memo.

  4. Ask for Deere compared with Tesla. The schema’s enum does not allow TSLA. Write two sentences on what the model did instead, and whether a client could tell.

  5. Extend the step log with the total elapsed time of the whole run.

  6. In three sentences, explain one tool call the model made that you would not have made, and what schema change would prevent it.

Further reading#