The AI Agent Wiki
A plain-language introduction to AI agents: what they are, how they work, the pieces they are made of, and how to build, evaluate, and safely operate them.
01 What is an AI Agent?
An AI agent is a software system that uses a large language model (LLM) as its decision-making core to perceive a situation, plan a course of action, act on it — often by calling tools — and then observe the result and keep going until a goal is reached.
The defining property is the loop: an agent is not a single question-and-answer exchange. It runs in cycles, each cycle turning the outcome of the previous action into the input of the next decision. This is what separates an agent from a plain chatbot.
Agent vs. plain LLM call
| Dimension | Plain LLM call | AI agent |
|---|---|---|
| Interaction | One prompt in, one response out | Multi-step loop with intermediate results |
| State | Stateless (context resets each call) | Stateful: memory of steps, results, and plans |
| Actions | Produces text only | Can call tools, run code, query APIs, control systems |
| Feedback | None during generation | Observes tool output, corrects course, retries |
| Goal | Answer the prompt | Accomplish a task, possibly with sub-goals |
| Failure handling | Responds anyway | Can detect failure, retry, or ask for help |
02 Core Components
Almost every agent can be described with six building blocks. You can build an agent from scratch with nothing but these; frameworks just package them more conveniently.
1. Model (reasoning core)
The LLM that decides what to do next. May be general-purpose or fine-tuned for tool use.
2. Tools
External capabilities the agent can invoke: web search, code execution, file access, APIs, databases.
3. Memory
Everything the agent remembers: conversation history, tool results, and long-term facts stored externally.
4. Planner
The strategy layer that decomposes a goal into steps and decides which step to tackle now.
5. Environment
The world the agent acts in: a sandbox, a browser, a codebase, a company's internal systems, or the open internet.
6. Guardrails
Rules and safety layers: tool allowlists, permission gates, output filters, budgets, and human approval points.
03 The Agent Loop
The canonical pattern for most agents is called ReAct (Reasoning + Acting, from a 2022 paper by Yao et al.). The agent alternates between thinking and doing until the task is complete:
In pseudocode, a minimal agent looks like this:
def run_agent(goal, tools, memory):
# ReAct loop: Reason -> Act -> Observe, until done
while not memory.is_done():
state = memory.current_state() # Observe
thought = model.reason(state, goal) # Reason: "I should search X"
action = model.choose_action(thought) # Act: pick a tool + arguments
result = tools.execute(action) # Tool runs in the environment
memory.record(action, result) # Remember, then loop
if guardrails.need_human(action): # Safety checkpoint
memory.pause_for_approval()
return memory.final_answer()
The loop terminates when the agent decides the goal is met, hits a maximum step budget, requests human help, or is stopped by a guardrail. Step budgets are important: without them, a loop with a bug can spend real money on API calls.
04 Types of Agents
Not all agents are the same shape. Common designs, roughly in order of increasing complexity:
ReAct agent
Interleaves reasoning traces and tool calls step-by-step. Simple, transparent, and the default for many SDKs.
Tool-calling agent
Relies on the model's native function-calling ability: the model emits structured calls (e.g. JSON) that the runtime executes.
Planner–executor
Writes a full plan up front, then executes it step-by-step, re-planning only when something goes wrong. Good for long tasks.
Reflection agent
Produces a draft, then critiques it in a second pass (sometimes a separate "critic" model) and revises. Improves quality at extra cost.
RAG agent
Retrieves relevant documents from a knowledge base before answering. The dominant pattern for domain question-answering.
Multi-agent system
Several specialized agents cooperate: an orchestrator delegates to workers, or agents debate and review each other.
Autonomous / long-running
Runs without a human in each step: scheduled jobs, background research, or agents that act on inboxes and queues.
Computer-use / browser agent
Operates a GUI like a human: clicking, typing, and reading screens in a real or virtual browser/desktop.
05 Reference Architecture
A realistic production agent is a pipeline. Here is the shape most systems converge on:
prompt, message, event, cron job
owns the loop, session state, and policy
decompose goal, pick strategy
history + vector store (RAG)
match intent to capability
permission gates, budgets, sandbox, audit log
Layers worth knowing
- Interface layer — how the agent is triggered and how results are delivered (chat, dashboard, API, messaging platform).
- Orchestration layer — the loop: state, step limits, retries, and routing. This is where frameworks like LangGraph or the OpenAI Agents SDK earn their keep.
- Cognitive layer — planning, prompting, and the model calls themselves.
- Capability layer — the tools: search, code execution, files, APIs. In 2025–2026 the Model Context Protocol (MCP) became the standard way to plug tools in (see section 10).
- Governance layer — identity, permissions, budgets, sandboxing, logging, and evals. Often the layer that separates a demo from a product.
06 Memory Systems
An LLM has no memory between calls. Everything an agent "remembers" must be stored and fed back in explicitly. Memory comes in two broad flavors:
| Short-term (working) memory | Long-term memory | |
|---|---|---|
| What it holds | Current task: conversation history, recent tool results, in-progress plan | Durable facts and past experiences across sessions |
| Where it lives | Inside the model's context window | External store: vector database, key-value store, files |
| How it's used | Sent with every model call as messages | Retrieved on demand (e.g. embedding similarity search) and injected as context |
| Main constraint | Context-window size and cost (tokens in = money) | Retrieval quality — garbage in, garbage out |
| Techniques | Trimming, summarization, context compaction | Embeddings, vector stores, RAG, knowledge graphs, memory files |
Key techniques
- RAG (Retrieval-Augmented Generation) — embed documents, store vectors, retrieve the most relevant chunks for each question, and stuff them into the prompt. The workhorse of knowledge-heavy agents.
- Context compaction — when a conversation gets long, summarize old turns into a compact digest to keep costs and latency down.
- Episodic vs. semantic — remembering what happened (episodes) vs. remembering what things mean (semantic facts). Mature agents keep both.
- Memory files / notes — many agent products (e.g. coding agents) persist a plain-text notes file the agent updates, which acts as durable long-term memory without a vector DB.
07 Tools & Tool Calling
Tools are how an agent stops being a text generator and starts affecting the world. The model doesn't run the tool — it proposes a call, and the runtime executes it with the model's chosen arguments.
How function calling works
- You declare available tools as JSON schemas in the model request.
- The model decides a tool is needed and returns a structured call (tool name + arguments), not free text.
- The runtime validates and executes the call in a controlled environment.
- The result is appended to the conversation as a new message; the model continues from there.
An example tool declaration:
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": { "type": "string" },
"unit": { "type": "string", "enum": ["celsius", "fahrenheit"] }
},
"required": ["city"]
}
}
}
Common tool classes
| Tool class | Examples | Risk |
|---|---|---|
| Read-only information | Web search, page fetch, DB queries | Low |
| Code execution | Python sandbox, notebook runner | Medium — must be sandboxed and rate-limited |
| File & repo operations | Read/write files, git commits, code edits | Medium — needs scoped permissions and review |
| External actions | Send email, post content, buy things, deploy code | High — irreversible, needs human approval |
08 Planning & Reasoning
The "smart" part of an agent is largely prompting structure plus model capability. The standard techniques:
- Chain-of-Thought (CoT) — asking the model to reason step-by-step before answering. Dramatically improves multi-step accuracy.
- Task decomposition — split a big goal ("prepare Q3 report") into verifiable sub-tasks ("gather data", "compute metrics", "draft sections"), each of which can be checked independently.
- Plan-and-Execute — produce the full plan first, then execute with re-planning on failure. Cheaper than re-deciding every step, but brittle if the world changes.
- Reflection / self-critique — a second pass that reviews the first output against a rubric and revises it.
- Verification steps — the agent tests its own output (run the code, check the numbers, ask a "tester" agent) before declaring success.
09 Multi-Agent Systems
Sometimes one agent is not enough — either because the task has conflicting requirements, or because specialization makes each role better. Common patterns:
Orchestrator–worker
A lead agent splits the work and delegates to specialist workers (research, coding, review), then assembles the results.
Debate
Two agents argue opposing positions; a judge agent decides. Reduces one-sided bias on controversial decisions.
Reviewer / critic loop
A writer agent produces; a critic agent reviews against criteria; back and forth until quality passes.
Swarm / role crews
Agents with roles ("researcher", "writer", "editor") hand off work in sequence, like a team in a chat room.
10 Frameworks & Standards
The agent ecosystem consolidated rapidly. As of 2026, the landscape looks roughly like this:
| Framework | Language | Key idea | Best for |
|---|---|---|---|
| LangChain | Python / JS | Chains, LCEL, huge tool/community ecosystem | General agent apps, quick prototypes |
| LangGraph | Python / JS | Graph-based state machines; explicit control flow | Complex, stateful, long-running agents |
| LlamaIndex | Python | Data frameworks, RAG pipelines | Retrieval-heavy agents and knowledge apps |
| AutoGen / AG2 | Python | Multi-agent conversation programming | Research and multi-agent experiments |
| CrewAI | Python | Role-based "crews" that hand off tasks | Team-style workflows, non-experts |
| OpenAI Agents SDK | Python / JS | Handoffs, guardrails, built-in tracing | Production apps on OpenAI models (successor to Swarm) |
| Claude Agent SDK | Python / TS | Agent loop + Claude Code integration | Production apps on Anthropic models |
| Google ADK | Python | Multi-agent, tool chaining, evaluation toolkit | Gemini-based agents |
| Semantic Kernel | C# / Python / Java | Enterprise SDK with planners and connectors | .NET / Microsoft enterprise stacks |
| Pydantic AI | Python | Type-safe agents built on Pydantic schemas | Typed, testable, maintainable agents |
Interop standards (the important part)
11 Evaluation & Observability
Agents are stochastic programs: the same input can produce different outputs. Evals are the unit tests of the agent world — automated checks that catch regressions before users do.
What to measure
| Metric | What it captures | Why it matters |
|---|---|---|
| Task success rate | % of end-to-end runs that meet the goal | The headline number |
| Tool-call correctness | % of tool invocations with valid args / right tool | Catches routing and schema bugs early |
| Trajectory efficiency | Steps & tokens per successful task | Cost and latency control |
| Cost per task | Total API spend per completed run | Agents that "work" but burn money still fail |
| Latency | Wall-clock time per step / per task | UX and timeout budget |
| Safety violations | # of policy breaches (banned tool use, etc.) | Release gate for anything with real actions |
Methods
- Golden sets — a curated batch of tasks with expected outcomes, run on every change. The minimum viable eval.
- LLM-as-judge — a second model scores outputs against a rubric. Cheap at scale, but the judge has its own biases; spot-check with humans.
- Trajectory tracing — record every step (prompt, tool call, result) in a trace store like Langfuse or LangSmith. You cannot debug what you cannot replay.
- Regression gates — block deploys when success rate drops below a threshold. Simple and brutally effective.
12 Safety & Risks
An agent with tools is a program that can take real actions. The risk surface is wider than a chatbot's, and the mitigations are mostly boring engineering:
| Risk | What it looks like | Mitigations |
|---|---|---|
| Prompt injection | Untrusted content (web page, email, doc) contains instructions that hijack the agent | Treat tool outputs as data, not instructions; separate system/user/tool roles; filter suspicious content; never grant tools to unverified input |
| Unsafe tool use | Agent deletes files, sends emails, or spends money unintentionally | Least-privilege tool set, allowlists, human approval for irreversible actions, sandboxes |
| Hallucination & over-reliance | Confident wrong answers presented as fact | RAG grounding, required citations, verification steps, confidence thresholds |
| Data exfiltration | Private data leaks into a prompt, a log, or an external tool | Redaction, scoped permissions, no-internet modes for sensitive tasks, audit logs |
| Goal drift / misalignment | Agent optimizes the letter of the goal, not the intent | Clear success criteria, step budgets, human checkpoints, periodic goal re-reading |
| Cost blowout | Runaway loop spends thousands of dollars | Hard step/token/cost budgets enforced in code, not in the prompt |
| Supply chain | Third-party MCP servers / plugins with malicious behavior | Vendor review, pinned versions, network isolation, same scrutiny as any dependency |
13 Glossary
| Term | Definition |
|---|---|
| Agent | A system that uses an LLM to plan and act in a loop, with tools and memory, toward a goal. |
| Agent loop | The observe→reason→act→observe cycle an agent repeats until its task is done. |
| Autonomy | How much an agent acts without human approval; production agents are usually only partially autonomous. |
| Context window | The maximum number of tokens a model can consider in one call. Everything an agent "sees" must fit here. |
| Context compaction | Summarizing old conversation turns into a shorter digest to fit the context window and cut cost. |
| Chain-of-Thought (CoT) | Prompting the model to reason step-by-step before answering; improves multi-step accuracy. |
| Embedding | A vector (list of numbers) representing text's meaning; similar texts have similar vectors. |
| Eval | An automated test that scores an agent's output or behavior; the unit test of agent development. |
| Function calling | A model capability to emit structured tool calls (name + arguments) instead of free text. |
| Golden set | A curated batch of tasks with expected outcomes used to regression-test agent changes. |
| Guardrails | Enforced rules and gates around an agent: permissions, budgets, sandboxes, approval steps. |
| Hallucination | A model producing confident output that is factually wrong or fabricated. |
| Handoff | Passing a task from one agent (or agent turn) to another, common in multi-agent systems. |
| Human-in-the-loop (HITL) | A design where a human reviews or approves agent actions, typically the risky ones. |
| LLM-as-judge | Using a model to score another model's output against a rubric in evaluations. |
| MCP (Model Context Protocol) | Open standard (Anthropic, 2024) for connecting agents to tools and data — "USB-C for agents". |
| Memory | Stored information an agent can draw on: working context, history, or long-term external stores. |
| Multi-agent system | Several cooperating agents with different roles, coordinated by patterns like orchestrator–worker. |
| Orchestrator | The component that owns the agent loop, session state, and routing between model, memory, and tools. |
| Prompt injection | An attack where untrusted content embeds instructions that override the agent's intended behavior. |
| RAG (Retrieval-Augmented Generation) | Retrieving relevant documents before answering and grounding the model's reply in them. |
| ReAct | Reasoning + Acting pattern: interleave reasoning traces with tool actions (Yao et al., 2022). |
| Sandbox | An isolated environment (container, VM, browser profile) where agent actions can't harm the host. |
| Step budget | A hard cap on loop iterations, enforced in code to stop runaway behavior and cost. |
| System prompt | The instructions (role, rules, tool descriptions) prepended to every model call. |
| Token | The atomic unit of text models process; ~0.75 words for English. Tokens in = cost. |
| Tool | An external capability an agent can invoke: search, code execution, files, APIs. |
| Tool router | Logic that matches the agent's intent to the right tool among many. |
| Trace | A recorded log of one agent run: every prompt, tool call, and result, used for debugging. |
| Vector database | A store that searches by embedding similarity, the backbone of RAG memory. |
14 FAQ
Do AI agents replace software developers?
Are agents "AGI"?
Can an agent run 24/7 by itself?
How much does an agent cost to run?
What is the difference between MCP and A2A?
Where should a beginner start?
Can agents be trusted with money or external actions?
15 Timeline: How we got here
- Oct 2022ReAct paper (Yao et al.)Formalizes the reasoning + acting loop that most agents still use today.
- Nov 2022ChatGPT launchesChat interface makes LLMs mainstream — the precursor to agent interfaces.
- Mar 2023AutoGPT & BabyAGIOpen-source "autonomous agent" demos go viral; the term "agent" enters the lexicon (with unrealistic expectations).
- Jun 2023OpenAI function callingLLMs gain native, structured tool use — the technical foundation of modern agents.
- Sep 2023AutoGen released (Microsoft)Multi-agent conversation programming becomes mainstream research territory.
- Late 2023CrewAI & the framework boomRole-based crews and a wave of agent frameworks (LangChain tooling, LlamaIndex, later LangGraph) make agents accessible to non-researchers.
- Mar 2024Devin demo (Cognition)"First AI software engineer" sparks the coding-agent wave.
- Oct 2024Anthropic computer useClaude 3.5 Sonnet operates a desktop GUI — agents learn to click, not just call APIs.
- Nov 2024MCP announced (Anthropic)Model Context Protocol standardizes tool connections; industry-wide adoption follows within months.
- Jan 2025OpenAI OperatorBrowser agent research preview — agents browse, fill forms, and transact on the web.
- Feb 2025Claude CodeTerminal-native coding agent becomes the reference for developer agents; Deep Research products popularize long-horizon agents.
- Mar 2025OpenAI Agents SDK; OpenAI adopts MCPVendor SDKs arrive; MCP becomes a cross-vendor standard.
- Apr 2025Google ADK & A2A protocolAgent Development Kit ships; the Agent2Agent protocol proposes inter-agent communication.
- Jun 2025A2A joins the Linux FoundationInter-agent standards go neutral-vendor.
- Sep 2025Claude Agent SDKVendor agent SDKs mature into production tooling with built-in loops, guardrails, and tracing.
- 2025–2026The production eraAgents move from demos to deployed systems: evals, observability (OpenTelemetry GenAI), compliance, and cost governance become the dominant topics.