nup_logo

Building AI Agents

From LLM to AI agents


Alex Avdiushenko
October 8, 2026

AI Agent Applications

Knowledge & Productivity

  • Personal assistants for research and decisions
  • ChatGPT, Perplexity, Gemini

Action & Execution

  • Browser agents that navigate, click, fill forms
  • Developer agents such as Claude Code and Codex

Business & Interaction

  • Customer support across channels
  • Operations automation and follow-up

Interactive Worlds

  • Game NPCs with adaptive behavior
  • Context-aware dialogue and quests

The Evolution of Human–LLM Interaction

Interaction Description Pattern example
Prompting LLM output has no impact on program flow process_llm_output(llm_response)
Router LLM determines the basic control flow if llm_decision(): path_a() else: path_b()
Tool Calling LLM selects the function and arguments run_function(llm_chosen_tool, llm_chosen_args)
Multi-step Agent (Autonomous Agent) LLM output controls iterations and continuation of program execution while llm_should_continue(): execute_next_step()
Multi-Agent System One agent process can trigger another agent process if llm_trigger(): execute_agent()

What Is an AI Agent?

OpenAI Agents SDK presentation:

An agent is an AI application consisting of

  • a model equipped with
  • instructions that guide its behavior,
  • access to tools that extend its capabilities,
  • encapsulated in a runtime with a dynamic lifecycle.

Lilian Weng (rejoin OpenAI in July 2026 after about 2 years in Thinking Machines Lab) , LLM Powered Autonomous Agents [1]:

Agent = LLM + memory + planning skills + tool use

[1] Weng, Lilian. “LLM-powered Autonomous Agents”. Lil’Log (Jun 2023).
[2] Weng, Lilian. “Harness Engineering for Self-Improvement”. Lil’Log (Jul 2026).

What Is an AI Agent?

Agent = runtime( AI model + prompts + tools + memory + guardrails + planning skills )

Agent runtime
Short-term
Context
Long-term
Memory
Guardrails
Calendar()
Search()
... more
Tools
AI model
Prompts
System
User
Assistant
Planning
Reflection
Reasoning
Subgoal decomposition

LLM as an Operating System

Peripheral devices I/O
Multimodal input and output
Vvideo
Aaudio
Ttext
Software 1.0 tools
Classical computer tools
123Calculator
PyPython interpreter
shTerminal
⇄
CPU
LLM
Interprets goals, selects actions, and coordinates tools
reasoning planning policy
⇄
Network
External systems
Webbrowser
APIservices
LLMother models
Disk
Persistent state
FSfile system
Embembeddings
DBvector store
⇄
RAM
Context window Working memory for the current task
limited size

Windows macOS Linux
Traditional OS analogy
GPT Claude Gemini
AI model ecosystem
Source: Andrej Karpathy, Intro to Large Language Models
youtube.com/watch?v=zjkBMFhNj_g

What Is an LLM

Large Language Model (LLM) is a neural network model trained on a large corpora of text data that predicts the next token based on the context

Chunk of the internet,
~10TB of text
< > llama-2-70b
parameters
.c
run.c
140GB
~500 lines
of C code

What is an LLM: Next Token Prediction

Large Language Model (LLM) is a neural network model trained on a large corpora of text data that predicts the next token based on the context


Start with words
cat
sat
on
a
mat (97%)

Tokens: Where Do They Come From

How to encode "I was reading an interesting book in New York"

Character encoding

  • I=1, _=2, w=3, a=4, s=5, r=6,
    e=7, d=8, i=9, n=10, g=11, t=12,
    b=13, o=14, k=15, N=16, Y=17
  • 1 2 3 4 5 2 6 7 4 8 9 10 11 2 4
    10 2 9 10 12 7 6 7 5 12 9 10 11
    2 13 14 14 15 2 9 10 2 16 7 3 2
    17 14 6 15
! Long sequences Every character becomes a step for the model.

Word encoding

  • I=1, was=2, reading=3, an=4,
    interesting=5, book=6, in=7,
    New=8, York=9
  • 1 2 3 4 5 6 7 8 9
? Huge vocabulary New words, names, and typos need new ids.

Tokenization

  • I=1, was=2, read=3, ing=4,
    an=5, interest=6, book=7, in=8,
    New_York=9
  • 1 2 3 4 5 6 4 7 8 9
OK Balanced units Common chunks stay compact, rare words can split.

Autoregression

Autoregression Schema

Source: Hugging Face Agents Course

Tokens: Special Tokens

  • Reserved tokens with predefined functions
  • Do not represent normal text
  • Control model behavior and structure inputs/outputs

Examples

  • EOS tokens: <eos>, <|endoftext|> — stop generation
  • Chat roles: <|user|>, <|assistant|>, <|system|>
  • Control tokens: <translate>, <summarize>, <formal>

Softmax Temperature

Softmax Logits
Temperature Effects Temperature Scaling GIF

Source: Medium: Softmax Temperature by Harshit Sharma

How LLMs Are Trained

Training
Steps
Self-supervised
Pre-training
Supervised
Instruction
Fine-tuning
Reinforcement learning
from human feedback
(RLHF)
Models
Foundational
LLM Model
Next →
SFT LLM Model
Next →
RLHF LLM Model
Train ↑
Train ↑
Train ↑
Datasets
Wiki, Articles, Books,
Web sites…
Human instructions
& responses
Human annotated
scores of LLM outputs
Reward Model
(RM)

LLMs: Key Takeaways

LLMs operate on tokens (including special ones), not words or meaning

  • Billing and context limits
  • Prompt structure and context formatting matter

LLMs predict the next token based on probabilities

  • temperature
  • hallucinations
    • We automate everything that can be automated
    • And we AI-ify everything that can't

Strengths:

  • Language tasks (summarization, paraphrasing, text/code generation)
  • Structured reasoning patterns, tool use, file/table processing

Limitations:

  • Precise algorithmic/symbolic tasks
  • Up-to-date knowledge (post-training events)
  • Long multi-step planning and persistent memory
➡

Many limitations can be mitigated in agents using tools and memory.

Spoilers Everywhere Meme

A few tasks to solve (15-20 min)

  1. Iterator Warm-up
  2. Profiler Decorator

What is "agent"?

  • An "intelligent" system that interacts with some "environment"
    • Physical environments: robot, autonomous car, ...
    • Digital environments: DQN for Atari, Siri, AlphaGo, ...
    • Humans as environments: chatbot
  • Define "agent" by defining "intelligent" and "environment"
    • It changes over time!
    • Exercise question: how would you define "intelligent"?

What is "LLM agent"?


Agents Diagram
  • Level 1: Text agent
    • Uses text action and observation
    • Examples: ELIZA, LSTM-DQN
  • Level 2: LLM agent
    • Uses LLM to act
    • Examples: SayCan, Language Planner
  • Level 3: Reasoning agent
    • Uses LLM to reason and act
    • Examples: ReAct, AutoGPT
    • The key focus of the field and the talk

Source: LLM Agents MOOC, Fall 2024, Slide 6

A brief history of LLM agents

  • Reasoning
    • CoT (Chain of Thought)
    • Zero-shot CoT
    • Self-consistency
    • ...
  • Acting (Grounding, tool use, etc.)
    • Game
    • Robotics
    • RAG (Retrieval-Augmented Generation)
    • ...
  • $\to$ LLM agent (but not reasoning agent) evolves into ReAct (Reasoning + Acting)
    • New apps/tasks/benchmarks:
      • Web browsing
      • Software engineering (Junie)
      • Scientific discovery (Deep Research)
      • ...
    • New methods:
      • Memory, learning, planning, multi-agent...

RAGs

RAGs_scheme

Reasoning OR Acting

CoT

Reasoning Traces $\rightleftarrows$ LM

  • Flexible and general to augment test-time compute
  • Lack of external knowledge and tools

RAG/Retrieval/Code/Tool use

Actions $\rightleftarrows$ LM $\rightleftarrows$ Env

  • Retrieval
  • Search engine
  • Calculator
  • Weather API
  • Python
  • ...
  • Lack of reasoning
  • Flexible and general to augment knowledge, computation, feedback, etc.
ReAct

Source: LLM Agents MOOC, Fall 2024

ReActSpace

Source: LLM Agents MOOC, Fall 2024

ReAct is general and effective

(NLP tasks) (RL tasks)
PaLM-540B HotpotQA
(QA)
FEVER
(fact check)
ALFWorld
(Text game)
Reason 29.4 56.3 N/A
Act 25.7 58.9 45
ReAct 35.1 64.6 71

Let's talk about..

Fish Memory Comic

So we need a long-term memory

  • Read and write
  • Stores experience, knowledge, skills, ...
  • Persists over new experience

Code-based controller

Instruction: ...

Thought: ...

Action: ...

Observation: ...

Thought: ...

...

Programming Task Reflection

  • Task: Given two strings of parentheses, determine if they are balanced
  • Trajectory: def match_parens(lst): ...
  • Evaluation:
    • Internal or external evaluation: Self-generated unit tests fail: the function only checks if the counts of parentheses are equal but ignores their order
  • Reflection: The function is wrong because it only counts the parentheses, but the order matters for valid balancing


GPT-4-reflexion

  • $\to$ Next Trajectory
reflexion_learning

Source: LLM Agents MOOC, Fall 2024

Semantic memory of (reflective) knowledge

semantic_memory

Cognitive Architectures for Language Agents (CoALA)

  • Memory
  • Action space
  • Decision making

Exercise Questions

  • What distinguishes external environment VS internal memory?
  • What distinguishes long VS short term memory?

Observation $\to$ (what "language"?) $\to$ Action

Symbolic AI Agent: 1960 - 1990

sp {blocks-world*opsub*proposal*clear}
(state <s> ^name blocks-world ^desired <d*1> ^clear <dobject> <ontop2>)
^bottom-block <object>

(Deep) RL Agent: 1990 - 2020

[-0.3432, 2.444, 0.34342, ...
0.4545, 0.443, 3.34234]

timeline.png

LLM Agent: 2020 - now

Let’s think step by step…
The room is dark, so I need a
lamp, the lamp is in
bedroom, so I should ...

Symbolic state or neural embedding

  • Intensive efforts to design or train
  • Task-specific, hard to generalize


Open-ended natural language

  • Rich priors from LLMs
  • Inference-time scalable
  • General and generalizable

How to evaluate LLM agents?

WebShop (2022)

WebShop Interface
  • Large-scale complex environment based on 1.16M Amazon products
  • Automatic reward based on instruction and product attribute matching
  • Challenges language and visual understanding, and decision-making

Web Arena (2023)

WebArena.png

SWE-Bench (2023)

SWE-Bench

ChemCrow: ReAct enables discovery of a novel chromophore

ChemCrow.png

What's Next?

Training

FireAct: Toward Language Agent Fine-tuning

Interface

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Robustness

Human

$\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Benchmark

FireAct: Training LLM for agents



Establish model-agent synergy:

  • Improve "agent capabilities" like planning, self-evaluation, calibration
  • Open-source agent backbone model
  • Next trillion tokens for model training
ReActFT.png

SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

LLMs and Humans are Different, So Should Their Interfaces

  • e.g., humans have smaller short-term memory, so they have to trade off time for space


ACI Design Can Help Us

  • Better solve tasks (without changing the agent)
  • Better understand agents (vs. humans)
SummarizedSearch.png SWEresults.png

Human in the loop needs robustness

  • You need to solve Riemann hypothesis once (one time out of millions)
  • But you always need to buy tickets correctly
t-bench.png

TAU-bench: A Benchmark for Tool-Agent-User

job-automation-ladder

Language Agents in the Digital World

Practice

Notebooks for this intensive:

Summary

  1. LLM Agents Overview
    • LLM agents are goal-oriented systems using language models for reasoning and acting
    • ReAct combines reasoning and action, improving performance across multiple tasks
    • Challenges: Tool integration and improving autonomous planning
  2. CoT vs RAG
    • Chain of Thought (CoT) focuses on internal reasoning traces, lacking external tools
    • RAG augments LLMs by connecting them with external knowledge, feedback, and computation tools
  1. Reflection & Long-term Memory
    • Reflection and self-evaluation improve task-solving abilities over iterations
    • LLM agents need long-term memory to store and update experience persistently
  2. Future of LLM Agents
    • Training improvements via models like FireAct
    • Interfaces tailored for human-agent collaboration (SWE-agent)
    • Robustness and evaluation benchmarks such as TAU-bench and WebShop