AI Agents Wiki v2026.08
Open reference • Community-maintained style

The AI Agent Wiki

A plain-language introduction to AI agents: what they are, how they work, the pieces they are made of, and how to build, evaluate, and safely operate them.

1 core loop 6 building blocks 10+ frameworks 15 sections 40+ glossary terms

01 What is an AI Agent?

An AI agent is a software system that uses a large language model (LLM) as its decision-making core to perceive a situation, plan a course of action, act on it — often by calling tools — and then observe the result and keep going until a goal is reached.

The defining property is the loop: an agent is not a single question-and-answer exchange. It runs in cycles, each cycle turning the outcome of the previous action into the input of the next decision. This is what separates an agent from a plain chatbot.

Agent vs. plain LLM call

DimensionPlain LLM callAI agent
InteractionOne prompt in, one response outMulti-step loop with intermediate results
StateStateless (context resets each call)Stateful: memory of steps, results, and plans
ActionsProduces text onlyCan call tools, run code, query APIs, control systems
FeedbackNone during generationObserves tool output, corrects course, retries
GoalAnswer the promptAccomplish a task, possibly with sub-goals
Failure handlingResponds anywayCan detect failure, retry, or ask for help
💡 One-sentence summary
An agent is "an LLM wrapped in a loop with tools and memory, pointed at a goal." Everything else — frameworks, protocols, guardrails — is scaffolding around that idea.

02 Core Components

Almost every agent can be described with six building blocks. You can build an agent from scratch with nothing but these; frameworks just package them more conveniently.

🤖

1. Model (reasoning core)

The LLM that decides what to do next. May be general-purpose or fine-tuned for tool use.

🔧

2. Tools

External capabilities the agent can invoke: web search, code execution, file access, APIs, databases.

📚

3. Memory

Everything the agent remembers: conversation history, tool results, and long-term facts stored externally.

📊

4. Planner

The strategy layer that decomposes a goal into steps and decides which step to tackle now.

🌏

5. Environment

The world the agent acts in: a sandbox, a browser, a codebase, a company's internal systems, or the open internet.

🛡

6. Guardrails

Rules and safety layers: tool allowlists, permission gates, output filters, budgets, and human approval points.

✅ Mental model
Think of it like a well-staffed office: the model is the employee who decides, tools are the phone, laptop and filing cabinet, memory is the employee's notes and the archive room, the planner is the to-do list, the environment is the building, and guardrails are the manager's rules and the security guard at the door.

03 The Agent Loop

The canonical pattern for most agents is called ReAct (Reasoning + Acting, from a 2022 paper by Yao et al.). The agent alternates between thinking and doing until the task is complete:

Observe 👁
↓what is the current state?
Reason 🧠
↓what should happen next, and why?
Act 🚀
↓call a tool or produce output
Receive result 🔄
↻ loop back to Observe until goal is met
decision points actions feedback path

In pseudocode, a minimal agent looks like this:

agent_loop.py
def run_agent(goal, tools, memory):
    # ReAct loop: Reason -> Act -> Observe, until done
    while not memory.is_done():
        state   = memory.current_state()          # Observe
        thought = model.reason(state, goal)        # Reason: "I should search X"
        action  = model.choose_action(thought)     # Act: pick a tool + arguments
        result  = tools.execute(action)            # Tool runs in the environment
        memory.record(action, result)              # Remember, then loop
        if guardrails.need_human(action):   # Safety checkpoint
            memory.pause_for_approval()
    return memory.final_answer()

The loop terminates when the agent decides the goal is met, hits a maximum step budget, requests human help, or is stopped by a guardrail. Step budgets are important: without them, a loop with a bug can spend real money on API calls.

04 Types of Agents

Not all agents are the same shape. Common designs, roughly in order of increasing complexity:

🔄

ReAct agent

Interleaves reasoning traces and tool calls step-by-step. Simple, transparent, and the default for many SDKs.

📱

Tool-calling agent

Relies on the model's native function-calling ability: the model emits structured calls (e.g. JSON) that the runtime executes.

🗓

Planner–executor

Writes a full plan up front, then executes it step-by-step, re-planning only when something goes wrong. Good for long tasks.

🔈

Reflection agent

Produces a draft, then critiques it in a second pass (sometimes a separate "critic" model) and revises. Improves quality at extra cost.

🔍

RAG agent

Retrieves relevant documents from a knowledge base before answering. The dominant pattern for domain question-answering.

👥

Multi-agent system

Several specialized agents cooperate: an orchestrator delegates to workers, or agents debate and review each other.

⏳

Autonomous / long-running

Runs without a human in each step: scheduled jobs, background research, or agents that act on inboxes and queues.

🎭

Computer-use / browser agent

Operates a GUI like a human: clicking, typing, and reading screens in a real or virtual browser/desktop.

⚠️ Reality check
"Autonomous" rarely means unattended in production. The most reliable deployments are human-in-the-loop at the right moments: the agent proposes, a human approves the risky steps, and the agent executes the rest.

05 Reference Architecture

A realistic production agent is a pipeline. Here is the shape most systems converge on:

User / Trigger 👤
prompt, message, event, cron job
↓
Orchestrator 🔮
owns the loop, session state, and policy
Planner 📊
decompose goal, pick strategy
Memory 📚
history + vector store (RAG)
Tool Router 🔗
match intent to capability
↓
Web Search
Code Sandbox
Files & APIs
Databases
↑observations flow back to the orchestrator
Guardrails 🛡
permission gates, budgets, sandbox, audit log

Layers worth knowing

  • Interface layer — how the agent is triggered and how results are delivered (chat, dashboard, API, messaging platform).
  • Orchestration layer — the loop: state, step limits, retries, and routing. This is where frameworks like LangGraph or the OpenAI Agents SDK earn their keep.
  • Cognitive layer — planning, prompting, and the model calls themselves.
  • Capability layer — the tools: search, code execution, files, APIs. In 2025–2026 the Model Context Protocol (MCP) became the standard way to plug tools in (see section 10).
  • Governance layer — identity, permissions, budgets, sandboxing, logging, and evals. Often the layer that separates a demo from a product.

06 Memory Systems

An LLM has no memory between calls. Everything an agent "remembers" must be stored and fed back in explicitly. Memory comes in two broad flavors:

Short-term (working) memoryLong-term memory
What it holdsCurrent task: conversation history, recent tool results, in-progress planDurable facts and past experiences across sessions
Where it livesInside the model's context windowExternal store: vector database, key-value store, files
How it's usedSent with every model call as messagesRetrieved on demand (e.g. embedding similarity search) and injected as context
Main constraintContext-window size and cost (tokens in = money)Retrieval quality — garbage in, garbage out
TechniquesTrimming, summarization, context compactionEmbeddings, vector stores, RAG, knowledge graphs, memory files

Key techniques

  • RAG (Retrieval-Augmented Generation) — embed documents, store vectors, retrieve the most relevant chunks for each question, and stuff them into the prompt. The workhorse of knowledge-heavy agents.
  • Context compaction — when a conversation gets long, summarize old turns into a compact digest to keep costs and latency down.
  • Episodic vs. semantic — remembering what happened (episodes) vs. remembering what things mean (semantic facts). Mature agents keep both.
  • Memory files / notes — many agent products (e.g. coding agents) persist a plain-text notes file the agent updates, which acts as durable long-term memory without a vector DB.
💡 Cost angle
Context is the dominant cost driver. A 200,000-token context window is not a reason to use 200,000 tokens every call — every token in is billed every step of the loop.

07 Tools & Tool Calling

Tools are how an agent stops being a text generator and starts affecting the world. The model doesn't run the tool — it proposes a call, and the runtime executes it with the model's chosen arguments.

How function calling works

  1. You declare available tools as JSON schemas in the model request.
  2. The model decides a tool is needed and returns a structured call (tool name + arguments), not free text.
  3. The runtime validates and executes the call in a controlled environment.
  4. The result is appended to the conversation as a new message; the model continues from there.

An example tool declaration:

tool_schema.json
{
  "type": "function",
  "function": {
    "name": "get_weather",
    "description": "Get the current weather for a city",
    "parameters": {
      "type": "object",
      "properties": {
        "city": { "type": "string" },
        "unit": { "type": "string", "enum": ["celsius", "fahrenheit"] }
      },
      "required": ["city"]
    }
  }
}

Common tool classes

Tool classExamplesRisk
Read-only informationWeb search, page fetch, DB queriesLow
Code executionPython sandbox, notebook runnerMedium — must be sandboxed and rate-limited
File & repo operationsRead/write files, git commits, code editsMedium — needs scoped permissions and review
External actionsSend email, post content, buy things, deploy codeHigh — irreversible, needs human approval
✅ Design rule
Give the agent the smallest set of tools that can do the job, and make every risky tool require explicit human confirmation. Tool sprawl is the #1 cause of surprising agent behavior.

08 Planning & Reasoning

The "smart" part of an agent is largely prompting structure plus model capability. The standard techniques:

  • Chain-of-Thought (CoT) — asking the model to reason step-by-step before answering. Dramatically improves multi-step accuracy.
  • Task decomposition — split a big goal ("prepare Q3 report") into verifiable sub-tasks ("gather data", "compute metrics", "draft sections"), each of which can be checked independently.
  • Plan-and-Execute — produce the full plan first, then execute with re-planning on failure. Cheaper than re-deciding every step, but brittle if the world changes.
  • Reflection / self-critique — a second pass that reviews the first output against a rubric and revises it.
  • Verification steps — the agent tests its own output (run the code, check the numbers, ask a "tester" agent) before declaring success.
⚠️ Known failure mode
Agents are confident planners and poor estimators. They routinely under-decompose (too few steps) or over-plan (pages of steps for a trivial task). Strong system prompts that demand verifiable checkpoints fix most of this.

09 Multi-Agent Systems

Sometimes one agent is not enough — either because the task has conflicting requirements, or because specialization makes each role better. Common patterns:

🏠

Orchestrator–worker

A lead agent splits the work and delegates to specialist workers (research, coding, review), then assembles the results.

⚖️

Debate

Two agents argue opposing positions; a judge agent decides. Reduces one-sided bias on controversial decisions.

📝

Reviewer / critic loop

A writer agent produces; a critic agent reviews against criteria; back and forth until quality passes.

🕵️

Swarm / role crews

Agents with roles ("researcher", "writer", "editor") hand off work in sequence, like a team in a chat room.

💡 When to go multi-agent
Rule of thumb: start single-agent. Add a second agent only when you can name a concrete failure the extra agent fixes (e.g. "the draft agent never catches its own errors"). Multi-agent systems multiply cost, latency, and failure modes — a fan-out of 5 agents can turn one bug into five.

10 Frameworks & Standards

The agent ecosystem consolidated rapidly. As of 2026, the landscape looks roughly like this:

FrameworkLanguageKey ideaBest for
LangChainPython / JSChains, LCEL, huge tool/community ecosystemGeneral agent apps, quick prototypes
LangGraphPython / JSGraph-based state machines; explicit control flowComplex, stateful, long-running agents
LlamaIndexPythonData frameworks, RAG pipelinesRetrieval-heavy agents and knowledge apps
AutoGen / AG2PythonMulti-agent conversation programmingResearch and multi-agent experiments
CrewAIPythonRole-based "crews" that hand off tasksTeam-style workflows, non-experts
OpenAI Agents SDKPython / JSHandoffs, guardrails, built-in tracingProduction apps on OpenAI models (successor to Swarm)
Claude Agent SDKPython / TSAgent loop + Claude Code integrationProduction apps on Anthropic models
Google ADKPythonMulti-agent, tool chaining, evaluation toolkitGemini-based agents
Semantic KernelC# / Python / JavaEnterprise SDK with planners and connectors.NET / Microsoft enterprise stacks
Pydantic AIPythonType-safe agents built on Pydantic schemasTyped, testable, maintainable agents

Interop standards (the important part)

🔗 MCP — Model Context Protocol
Open standard released by Anthropic in November 2024, adopted by OpenAI, Google, Microsoft and most of the industry by 2025. Think "USB-C for agent tools": one protocol for connecting agents to tools, data sources, and services — instead of a custom integration per tool. An MCP server exposes capabilities; any MCP-compatible client can use them.
🔗 A2A — Agent2Agent
Protocol announced by Google in April 2025 and contributed to the Linux Foundation the same year. Where MCP connects agents to tools, A2A lets agents talk to other agents — discovery, task handoff, and results exchange between agents from different vendors.
📈 OpenTelemetry GenAI conventions
Standardized tracing spans for LLM calls, tool calls, and agent steps — the foundation for evaluating and debugging agents in production.

11 Evaluation & Observability

Agents are stochastic programs: the same input can produce different outputs. Evals are the unit tests of the agent world — automated checks that catch regressions before users do.

What to measure

MetricWhat it capturesWhy it matters
Task success rate% of end-to-end runs that meet the goalThe headline number
Tool-call correctness% of tool invocations with valid args / right toolCatches routing and schema bugs early
Trajectory efficiencySteps & tokens per successful taskCost and latency control
Cost per taskTotal API spend per completed runAgents that "work" but burn money still fail
LatencyWall-clock time per step / per taskUX and timeout budget
Safety violations# of policy breaches (banned tool use, etc.)Release gate for anything with real actions

Methods

  • Golden sets — a curated batch of tasks with expected outcomes, run on every change. The minimum viable eval.
  • LLM-as-judge — a second model scores outputs against a rubric. Cheap at scale, but the judge has its own biases; spot-check with humans.
  • Trajectory tracing — record every step (prompt, tool call, result) in a trace store like Langfuse or LangSmith. You cannot debug what you cannot replay.
  • Regression gates — block deploys when success rate drops below a threshold. Simple and brutally effective.
⚠️ The eval gap
The biggest operational surprise for most teams: an agent that passed demo day fails in production, because real inputs are messier than demo inputs. Build evals from real user traces, not from what you hope users will type.

12 Safety & Risks

An agent with tools is a program that can take real actions. The risk surface is wider than a chatbot's, and the mitigations are mostly boring engineering:

RiskWhat it looks likeMitigations
Prompt injectionUntrusted content (web page, email, doc) contains instructions that hijack the agentTreat tool outputs as data, not instructions; separate system/user/tool roles; filter suspicious content; never grant tools to unverified input
Unsafe tool useAgent deletes files, sends emails, or spends money unintentionallyLeast-privilege tool set, allowlists, human approval for irreversible actions, sandboxes
Hallucination & over-relianceConfident wrong answers presented as factRAG grounding, required citations, verification steps, confidence thresholds
Data exfiltrationPrivate data leaks into a prompt, a log, or an external toolRedaction, scoped permissions, no-internet modes for sensitive tasks, audit logs
Goal drift / misalignmentAgent optimizes the letter of the goal, not the intentClear success criteria, step budgets, human checkpoints, periodic goal re-reading
Cost blowoutRunaway loop spends thousands of dollarsHard step/token/cost budgets enforced in code, not in the prompt
Supply chainThird-party MCP servers / plugins with malicious behaviorVendor review, pinned versions, network isolation, same scrutiny as any dependency
✅ The three guardrail questions
Before shipping any agent, answer these in writing: (1) What is the worst thing this agent could do with its tools? (2) Which steps must a human approve? (3) How do we stop it — kill switch, budget, or timeout — and how fast?

13 Glossary

TermDefinition
AgentA system that uses an LLM to plan and act in a loop, with tools and memory, toward a goal.
Agent loopThe observe→reason→act→observe cycle an agent repeats until its task is done.
AutonomyHow much an agent acts without human approval; production agents are usually only partially autonomous.
Context windowThe maximum number of tokens a model can consider in one call. Everything an agent "sees" must fit here.
Context compactionSummarizing old conversation turns into a shorter digest to fit the context window and cut cost.
Chain-of-Thought (CoT)Prompting the model to reason step-by-step before answering; improves multi-step accuracy.
EmbeddingA vector (list of numbers) representing text's meaning; similar texts have similar vectors.
EvalAn automated test that scores an agent's output or behavior; the unit test of agent development.
Function callingA model capability to emit structured tool calls (name + arguments) instead of free text.
Golden setA curated batch of tasks with expected outcomes used to regression-test agent changes.
GuardrailsEnforced rules and gates around an agent: permissions, budgets, sandboxes, approval steps.
HallucinationA model producing confident output that is factually wrong or fabricated.
HandoffPassing a task from one agent (or agent turn) to another, common in multi-agent systems.
Human-in-the-loop (HITL)A design where a human reviews or approves agent actions, typically the risky ones.
LLM-as-judgeUsing a model to score another model's output against a rubric in evaluations.
MCP (Model Context Protocol)Open standard (Anthropic, 2024) for connecting agents to tools and data — "USB-C for agents".
MemoryStored information an agent can draw on: working context, history, or long-term external stores.
Multi-agent systemSeveral cooperating agents with different roles, coordinated by patterns like orchestrator–worker.
OrchestratorThe component that owns the agent loop, session state, and routing between model, memory, and tools.
Prompt injectionAn attack where untrusted content embeds instructions that override the agent's intended behavior.
RAG (Retrieval-Augmented Generation)Retrieving relevant documents before answering and grounding the model's reply in them.
ReActReasoning + Acting pattern: interleave reasoning traces with tool actions (Yao et al., 2022).
SandboxAn isolated environment (container, VM, browser profile) where agent actions can't harm the host.
Step budgetA hard cap on loop iterations, enforced in code to stop runaway behavior and cost.
System promptThe instructions (role, rules, tool descriptions) prepended to every model call.
TokenThe atomic unit of text models process; ~0.75 words for English. Tokens in = cost.
ToolAn external capability an agent can invoke: search, code execution, files, APIs.
Tool routerLogic that matches the agent's intent to the right tool among many.
TraceA recorded log of one agent run: every prompt, tool call, and result, used for debugging.
Vector databaseA store that searches by embedding similarity, the backbone of RAG memory.

14 FAQ

Do AI agents replace software developers?
Not as a category — but they change how development is done. Coding agents (like Claude Code, Cursor, and others) are already strong at scaffolding, refactoring, test writing, and well-specified tasks. Humans still set direction, review designs, and own the hard judgment calls. The pattern in 2026 is "agent pairs with developer", not "agent replaces developer".
Are agents "AGI"?
No. Agents are a product pattern — a loop with tools and memory — while AGI refers to a hypothetical machine matching general human intelligence. An agent can be useful and narrow (booking a flight, fixing a bug) without being general. Most production agents are very good at one thing and useless outside it.
Can an agent run 24/7 by itself?
Technically yes — scheduled, event-driven agents run around the clock. Operationally, you should still gate risky actions behind human approval and give every agent hard budgets. "Autonomous" in production usually means "autonomous within a fenced area".
How much does an agent cost to run?
The honest answer: it depends on steps and context. A simple agent task might cost a few cents (a handful of model calls); a research agent that reads 50 pages and runs 30 steps can cost dollars. The lever is context discipline — smaller prompts, compaction, and step budgets — not just a cheaper model.
What is the difference between MCP and A2A?
MCP standardizes how an agent connects to tools and data (agent→tool). A2A standardizes how agents talk to each other (agent→agent). They complement each other: an A2A mesh of agents, each using MCP to reach its own tools, is the emerging stack.
Where should a beginner start?
Start without a framework: pick an LLM API, implement a 20-line ReAct loop (see section 03), and give it one tool like web search. Once you feel the loop, pick a framework (Pydantic AI for typed Python, OpenAI or Claude Agent SDK for production, LangGraph for complex state) and build a real task you care about. Then write evals before adding features.
Can agents be trusted with money or external actions?
Only with the boring safeguards in place: least-privilege credentials, hard spend limits, human approval for irreversible steps, and full audit logs. The failures that make headlines are almost always deployments that skipped one of these — not the model "going rogue".

15 Timeline: How we got here

  • Oct 2022
    ReAct paper (Yao et al.)
    Formalizes the reasoning + acting loop that most agents still use today.
  • Nov 2022
    ChatGPT launches
    Chat interface makes LLMs mainstream — the precursor to agent interfaces.
  • Mar 2023
    AutoGPT & BabyAGI
    Open-source "autonomous agent" demos go viral; the term "agent" enters the lexicon (with unrealistic expectations).
  • Jun 2023
    OpenAI function calling
    LLMs gain native, structured tool use — the technical foundation of modern agents.
  • Sep 2023
    AutoGen released (Microsoft)
    Multi-agent conversation programming becomes mainstream research territory.
  • Late 2023
    CrewAI & the framework boom
    Role-based crews and a wave of agent frameworks (LangChain tooling, LlamaIndex, later LangGraph) make agents accessible to non-researchers.
  • Mar 2024
    Devin demo (Cognition)
    "First AI software engineer" sparks the coding-agent wave.
  • Oct 2024
    Anthropic computer use
    Claude 3.5 Sonnet operates a desktop GUI — agents learn to click, not just call APIs.
  • Nov 2024
    MCP announced (Anthropic)
    Model Context Protocol standardizes tool connections; industry-wide adoption follows within months.
  • Jan 2025
    OpenAI Operator
    Browser agent research preview — agents browse, fill forms, and transact on the web.
  • Feb 2025
    Claude Code
    Terminal-native coding agent becomes the reference for developer agents; Deep Research products popularize long-horizon agents.
  • Mar 2025
    OpenAI Agents SDK; OpenAI adopts MCP
    Vendor SDKs arrive; MCP becomes a cross-vendor standard.
  • Apr 2025
    Google ADK & A2A protocol
    Agent Development Kit ships; the Agent2Agent protocol proposes inter-agent communication.
  • Jun 2025
    A2A joins the Linux Foundation
    Inter-agent standards go neutral-vendor.
  • Sep 2025
    Claude Agent SDK
    Vendor agent SDKs mature into production tooling with built-in loops, guardrails, and tracing.
  • 2025–2026
    The production era
    Agents move from demos to deployed systems: evals, observability (OpenTelemetry GenAI), compliance, and cost governance become the dominant topics.