=== Counting characters task ===
87
Ground truth: 97
=== Math ===
8293
Ground truth: 825.444444444445
=== Up-to-date info ===
I don't have access to real-time data,
including the current Bitcoin price.
USER: 1) Cancel reservation EHGLP3.
AGENT: I'm unable to access specific reservations
or cancel them directly.
A tool is a function that an LLM can call to interact with external systems
A tool should include:
LLMs can only generate text, don't execute code.
To the user, it looks like the LLM used the tool — but the agent did the actual work.
Providing tools to an LLM means:
|
Handling Partial Failures
Tools must handle partial failures while keeping the system consistent charge payment → reserve inventory → create delivery |
Idempotency & Safe Retries
Tools must safely handle repeated calls (after crashes, timeouts, agent restarts) Sending an email → gets timeout → retry must not send it twice |
Clear Error Messages
Errors must be understandable for LLMs so the agent can decide what to do next Good errors enable the agent to:
|
|
Guidance for Next Steps
Tools can return guidance that helps the agent continue the conversation properly. Examples
|
Structured Outputs
Return structured, machine-readable outputs instead of free-form text. Examples
|
Key Takeaway
A good tool for an AI agent is not just a function It is a reliable contract between the LLM and external systems, designed to withstand failures, retries, and changing environments. |
|
Yes, Tools have significantly expanded the capabilities of LLMs and enabled them to act as agents |
But
|
→
|
→
On November 25, 2024, Anthropic open-sourced the Model Context Protocol (MCP)
Model Context Protocol (MCP) is an open standard that enables secure, bidirectional connections between AI models and external systems (APIs, tools, data sources)
MCP is the “USB-C port” or “HTTP layer” of the AI ecosystem.
https://www.anthropic.com/news/model-context-protocol
https://modelcontextprotocol.io/docs/getting-started/intro
Client–server architecture with three roles
| Orchestrates LLMs and AI workflows | Manages connections: one MCP client per MCP server | Controls UI: User-facing application (e.g., Cursor, custom agents) |
| Security & permissions: auth, limits, policies | User consent: approvals for data sharing & tool use |
AI Chat
|
Server = provider of context and capabilities
Tasks:
Client = connector between Host and Server
| Problem | Without MCP | With MCP |
|---|---|---|
| Tool reusability | Each team reimplemented the same tools independently. | Tools live in MCP servers — standalone infrastructure components (like microservices). Write once → reuse across many agents and products. |
| API updates & maintenance | All agent teams had to update their integrations manually. | Responsibility shifts to MCP server maintainers. Agents automatically discover updated capabilities during initialization. |
| Scaling and modifications | Components were tightly coupled. Changing one component required rewriting many parts in each agent | Separation of responsibilities. Changes are localized to individual components. |
| Vibe |
|
|
Do not over-engineer. Start simple. Add MCP when complexity forces you.
In the middle ground you can use lightweight tool ecosystems
(e.g., AgentSkills) that provide:
A few tasks to solve (15-20 min)
Notebook for this intensive: