| Interaction | Description | Pattern example |
|---|---|---|
| Prompting | LLM output has no impact on program flow | process_llm_output(llm_response) |
| Router | LLM determines the basic control flow | if llm_decision(): path_a() else: path_b() |
| Tool Calling | LLM selects the function and arguments | run_function(llm_chosen_tool, llm_chosen_args) |
| Multi-step Agent (Autonomous Agent) | LLM output controls iterations and continuation of program execution | while llm_should_continue(): execute_next_step() |
| Multi-Agent System | One agent process can trigger another agent process | if llm_trigger(): execute_agent() |
OpenAI Agents SDK presentation:
An agent is an AI application consisting of
Lilian Weng (rejoin OpenAI in July 2026 after about 2 years in Thinking Machines Lab) , LLM Powered Autonomous Agents [1]:
Agent = LLM + memory + planning skills + tool use
Agent = runtime( AI model + prompts + tools + memory + guardrails + planning skills )
Large Language Model (LLM) is a neural network model trained on a large corpora of text data that predicts the next token based on the context
Large Language Model (LLM) is a neural network model trained on a large corpora of text data that predicts the next token based on the context
How to encode "I was reading an interesting book in New York"
Source: Hugging Face Agents Course
<eos>, <|endoftext|> — stop generation<|user|>, <|assistant|>, <|system|><translate>, <summarize>, <formal>
LLMs operate on tokens (including special ones), not words or meaning
LLMs predict the next token based on probabilities
Strengths:
Limitations:
Many limitations can be mitigated in agents using tools and memory.
CoT
Reasoning Traces $\rightleftarrows$ LM
RAG/Retrieval/Code/Tool use
Actions $\rightleftarrows$ LM $\rightleftarrows$ Env
Source: LLM Agents MOOC, Fall 2024
Source: LLM Agents MOOC, Fall 2024
| (NLP tasks) | (RL tasks) | ||
|---|---|---|---|
| PaLM-540B | HotpotQA (QA) |
FEVER (fact check) |
ALFWorld (Text game) |
| Reason | 29.4 | 56.3 | N/A |
| Act | 25.7 | 58.9 | 45 |
| ReAct | 35.1 | 64.6 | 71 |
So we need a long-term memory
Code-based controller
Instruction: ...
Thought: ...
Action: ...
Observation: ...
Thought: ...
...
def match_parens(lst): ...
Source: LLM Agents MOOC, Fall 2024
sp {blocks-world*opsub*proposal*clear}
(state <s> ^name blocks-world ^desired <d*1> ^clear <dobject>
<ontop2>)
^bottom-block <object>
[-0.3432, 2.444, 0.34342, ...
0.4545, 0.443, 3.34234]
Let’s think step by step…
The room is dark, so I need a
lamp, the lamp is in
bedroom, so I should ...
FireAct: Toward Language Agent Fine-tuning
SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
$\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Notebooks for this intensive: