nup_logo

Building AI Agents

Multi-Agent Systems & Multimodality


Alex Avdiushenko
October 29, 2026

What Is an AI Agent?

Agent = runtime( AI model + prompts + tools + memory + guardrails + planning skills )

Agent runtime
Short-term
Context
Long-term
Memory
Guardrails
Calendar()
Search()
... more
Tools
AI model
Prompts
System
User
Assistant
Planning
Reflection
Reasoning
Subgoal decomposition

Agenda

01 How Agents Think (TAO Loop)
02 Multi-Agent Systems (MAS)
03 Multimodal Agents
04 Summary & Practice

Thought → Action → Observation (TAO) Loop

The TAO cycle (Thought-Action-Observation) is the workflow of an artificial
intelligence agent. It allows agents to dynamically adjust their responses,
improve understanding, and improve the accuracy of the results

Thought → Action → Observation (TAO) Loop

User Request
THOUGHT LLM reasons and plans
ACTION Tool call / response
OBSERVATION Result from environment
Final Answer
Loop until task
is solved

THOUGHT: Agent's Internal Reasoning

THOUGHT Type Example
Planning "I'll break the task into 3 steps:
find booking → check policies → update"
Analysis "Based on the API error, the issue is in the date format"
Self-reflection "My previous answer was too generic"

Four Reasoning Strategies

CoT (Chain-of-Thought)

Step-by-step reasoning before answering.

Single generation.

«Let's think step by step.»

ReAct

Alternating Think ↔ Act ↔ Observe.

Suitable for tool-use tasks

Reasoning (o1-style)

Built into the model during training
(o1, DeepSeek R1).

<think>...</think> — self-reflection.

Four Reasoning Strategies

CoT
Question
↓
Reasoning
↓
Reasoning
↓
Reasoning
↓
Answer
Chain-of-Thought
One single generation
Linear logic
ReAct
Question
↓
Thought
↓
Action → Env
↓
Observation
↓
Done? → Answer
ReAct (Reason + Act)
Multi-step execution
Interacts with environment
Reasoning Models
Question
↓
<think>
...hidden reasoning
...self-reflection
</think>
↓
Answer
DeepSeek R1 / OpenAI o1
Trained with RL to reason
Internal hidden thoughts

Strategy Comparison Table

Characteristic CoT ReAct Reasoning (o1)
Step-by-step logic Yes Yes Yes
External tools No Yes Optional
Plan correction No Yes (Observation) Internal
Where defined Prompt Prompt Model training
Best for Logic, math Tool-based tasks Complex planning
Cost Low Medium (N calls) High (reasoning tokens)
Default for agents

ACTION  + Tool Call Formats

Format Description Example
JSON Agent Action as a JSON object {"action": "search_flights", ...}
Function-calling Native model format Messages with tool_calls
Code Agent Executable code generation results = search_flights(...)

Stop and Parse

LLM
generates
action JSON
[STOP]
parses
action + args
Runtime
calls tool
Tool (API)
result

OBSERVATION: Feedback from the Environment

Observation Type Example
API Data {"flights": [...], "count": 3}
System Error {"error": "rate_limit_exceeded", "retry_after": 30}
Confirmation {"status": "booking_confirmed", "id": "BK-12345"}
User Input "Yes, book this flight"
Three processing steps:
  • 1. Parse — identify function and arguments
  • 2. Execute — perform the action
  • 3. Append — result → into agent context

Structured Output for Planning

Pydantic Schemas
Structured LLM Call
from pydantic import BaseModel

class SubTask(BaseModel):
    agent_name: str
    description: str
    priority: int

class Plan(BaseModel):
    subtasks: list[SubTask]

structured_llm = llm.with_structured_output(Plan)
plan = structured_llm.invoke(
    "Break the rebooking task into subtasks"
)
# plan.subtasks -> [SubTask(agent_name="flight", ...), ...]

TAO Loop in Action: Flight Rebooking

1
Thought: Find the booking
Action: get_booking(user_id="U-1234")
Obs: {booking_id: "BK-789", flight: "SU2454"}
2
Thought: Searching flights for tomorrow
Action: search_flights("SVO", "CDG", "04-16")
Obs: [{SU2456: 37k₽}, {AF1147: 42k₽}]
3
Thought: I'll check the policies
Action: get_policy("rebooking")
Obs: {fee: 3000₽, same class only}
4
Thought: I'll offer options to the user
Action: final_answer("SU2456 or AF1147, $40 surcharge")
4 TAO loop iterations, 3 tool calls, 1 final answer

A task to solve (15-20 min)

  1. Custom Range

Agenda

01 How Agents Think (TAO Loop) ✓
02 Multi-Agent Systems (MAS)
03 Multimodal Agents
04 Summary & Practice

Why Do We Need MAS?

Problem — a single agent as task complexity grows:
Gets lost with too many tools
Mixes up subtask contexts
Difficult to debug
Becomes unreliable
Solution: specialized agents, each an expert in their domain

Single-Agent vs Multi-Agent

Single-Agent Multi-Agent
Strengths Simplicity, no coordination,
fewer resources
Complex tasks, specialization,
parallelism, fault tolerance
Weaknesses Limited by context,
confused with many tools
Coordination overhead,
debugging complexity
When to choose Simple tasks
(2–5 tools)
Many tools,
diverse expertise
Key Advantages of MAS:
  • More complete coverage of problem aspects
  • Mutual verification and result validation
  • Small specialized models instead of one large model
  • Resilience to individual agent failures

6 Principles of MAS

DECOMPOSITION
Breaking the task
into subtasks
SPECIALIZATION
Each agent is
an expert in their domain
COMMUNICATION
Efficient exchange
of information
COORDINATION
Aligning actions
and aggregating results
VERIFICATION
Critics validate
intermediate results
FEEDBACK
Correction based on
info from other participants

Three Types of Inter-Agent Interaction

DEBATE
Agent A: "$450"
↔ arguments ↔
Agent B: "$382"
▼
Consensus: "$382.50"
When:
Accuracy needed,
fact-checking, ethics/risks
COLLABORATION
Agent 1: Flight →
Agent 2: Policy →
Agent 3: Booking →
▼
Coordinator aggregates
When:
Clear pipeline,
diverse skills, efficiency
COMPETITION
Agent A: "cheapest"
Agent B: "fastest"
Agent C: "comfort"
▼
Evaluator picks the best
When:
Multiple strategies,
creative tasks

Handoff — Key Collaboration Pattern

User: "I want to rebook and ask about baggage"
Coordinator
"2 subtasks"
Rebook Agent
handoff: {flight: SU2456}
Baggage Agent
handoff: {rules: "23kg"}
← in parallel!
Coordinator
"Flight SU2456 + $40 fee, baggage 23kg"

MAS Architectures

1. Hierarchical (Vertical)
Coordinator
Flight
Policy
Booking

MAS Architectures

2. Decentralized (Horizontal)
Agent A
Agent B
Agent C
Agent D

MAS Architectures

3. Centralized (Star)
Central Router
Agent 1
Agent 2
Agent 3

4 MAS Architectures

4. Shared Message Pool
Agent A
Agent B
Agent C
Shared Message Pool

Architecture Comparison

Criterion Hierarchical Decentralized Centralized Shared Pool
Task complexity High Medium-High Low-Medium Medium
Fault tolerance Medium High ★ Low Medium
Scalability Medium High ★ Low High ★
Ease of debugging High ★ Low High ★ Medium
Predictability High ★ Low High ★ Medium

MAS Protocols  + A2A

MCP (Model Context Protocol)

Agent ↔ Tool

Analogy: USB-C — a universal
adapter for connecting AI
to tools

Anthropic, Nov 2024

A2A (Agent-to-Agent)

Agent ↔ Agent

Analogy: Wi-Fi Direct — direct
communication between devices

Google, Apr 2025

MAS Limitations

Category Problem Example
Resource Intensity Each agent = LLM call.
3 agents × 5 iterations = 15+ calls
Cost, latency
Error Accumulation One agent's error amplifies
down the chain
A found the wrong flight →
B booked it
Debugging Hard to tell which agent failed
in a long chain
Need tracing
of every step

Decision Guide

No
2–5
6+
Yes
Simple API
Need an agent?
Single-Agent
How many tools?
Need verification?
MAS + Critic

What is the task structure?

Clear hierarchy
→ Hierarchical
Peer experts
→ Decentralized
Many subsystems
→ Centralized
Async processing
→ Shared Pool

Anti-patterns — What NOT to Do

Anti-pattern Problem Solution
MAS for a simple task Overhead > benefit Start with single-agent
10+ agents at once Impossible to debug Incrementally: 2–3 → validate
No critic in
critical domains
Errors reach
the user
Always use Critic for
critical data
All agents on expensive models Expensive and slow Small models
for simple agents
No tracing Impossible to debug Log every
TAO step
Rule: architecture is chosen to fit the task, not the other way around!

Agenda

01 How Agents Think (TAO Loop) ✓
02 Multi-Agent Systems (MAS) ✓
03 Multimodal Agents
04 Summary & Practice

Multimodal Agents

Type Modality Example Tasks
Vision Images, Video stream UI analysis, document OCR, visual QA
Voice Audio Voice assistants, call centers

Multimodal Agents

Multimodal MAS:
VisionAgent
"I see a booking form
with fields: name, date, flight"
TextAgent
"Filling in the fields
based on the request"
VoiceAgent
"Confirming
with the user via voice"

Conclusion

1. How a single model thinks
TAO loop, strategies: CoT, ReAct, Reasoning, Plan & Execute
2. How multiple models work as a team
MAS: architectures, interaction, protocols
3. How agents go beyond text
Multimodality: vision, voice, video
Key Takeaways:
TAO loop — the fundamental pattern underlying every agent
ReAct > CoT for tool-based tasks
MAS solves specialization, verification, and scalability challenges
4 architectures — no "best" one, only the right fit for the task
Multimodality — the next frontier
Conclusion illustration

Practice

Notebook for this intensive: