What Is an AI Agent? The One-Question Test

Ask ten people to define an AI agent and you'll get three different answers. Vendors slap the label on chatbots, researchers reserve it for autonomous systems, and engineering teams use it as a catch-all for whatever they've shipped. LangChain, four years into running agents in production, says it directly: the definition depends on who you ask.

Here's the part worth keeping: underneath the label fight, every definition shares one hard core. An agent is a system built on a large language model where the model decides the next step — not just once, but over and over, until the goal is done. LangChain's engineering phrasing: a system that uses an LLM to decide the control flow of an application. Anthropic's version: systems where LLMs dynamically direct their own processes and tool usage. Strip the jargon and it's one question: who calls the next move — your code, or the model?

This article settles the definitions, shows what actually happens inside an agent loop, and — the part most explainers skip — covers when you don't need one, and what agents really cost to run.

Chatty vs. capable: where chatbots end

IBM defines an agent as a system that autonomously performs tasks by designing workflows with available tools. A chatbot, by contrast, responds to prompts. It answers, summarizes, drafts. Then it stops.

There's a question that cuts through every vendor demo: can the AI actually do the job, or can it only talk about the job?

Cloudflare has the cleanest everyday example. Ask an LLM to "write an invitation email for my company's top 10 clients" and you get — an email. Good one, even. Ask an agent to "invite my top 10 clients to the dinner" and it looks up the top 10 in your CRM, drafts a personalized email for each, and sends them, permissions permitting. The first gives you homework. The second hands you a finished task.

DimensionChatbotAI agent
Primary behaviorResponds to messagesPursues a goal through steps
Typical outputAnswers, summaries, draftsCompleted actions
AutonomyLow — waits for each promptHigh — plans and continues within bounds
ContextChat history, pasted textTask state, tool results, product data
ToolsOptional or narrowCore to execution

The Chinese AI community has a phrase for this: giving the model "hands and eyes." The model was always the engine; an agent is the body that lets it act on the world.

Inside one turn of the loop

What makes an agent tick isn't a fancy model — it's a loop. The LLM runs in a cycle: read the current state, pick an action, call a tool, observe the result, update its memory, and decide whether to continue or stop.

Picture a research request: "summarize this quarter's top customer complaints." One turn of the loop might go — the agent reads the request, decides it needs the ticket database, queries it, gets back 400 tickets, realizes that's too many to summarize in one pass, chunks them into batches, summarizes each, merges, checks the merge against the original counts, and writes the report. Nobody scripted that path. The model chose each step based on what the last step returned.

Four components sit inside every one of them:

  • Model — the LLM that decides the next action. Strong reasoning and tool-calling matter more than raw size.
  • Tools — APIs, databases, code execution, retrieval, even other agents. This is what the agent acts with.
  • Memory — context that persists across turns: conversation history, long-term state. The physical limit here is the model's context window, which we've broken down in depth.
  • The loop — the reasoning cycle that ties it together. Think-Act-Observe, replayed until done.

Notice the tools list includes retrieval. A well-built retrieval pipeline is often the difference between an agent that stalls and one that finishes — we broke down how retrieval works in our RAG explainer.

Autonomy is a spectrum, not a switch

Nobody asks whether Level 2 driver-assist is "really" self-driving — the auto industry agreed on numbered levels, so engineers argue about which level they're shipping instead. AI hasn't gotten there yet, which is why "is it really an agent?" arguments never end.

The useful move is to stop arguing and place systems on a spectrum: handwritten logic at one end, then a single LLM call, then chains, then routers, then state machines, then — at the far end — a fully autonomous agent that picks its own tools and remembers what it built. The more control you hand the model, the more infrastructure you need around it: observability, evals, permissions, safe execution.

The practical dividing line is workflow vs. agent. A workflow orchestrates LLMs and tools through paths written in code. An agent lets the model decide the path at runtime. Anthropic files both under "agentic systems" and treats the split as one question: how much of the control flow is fixed in code versus decided by the model.

Both failure modes are real. Reach for autonomy when a workflow would do, and you've added non-determinism you didn't need — same input, different runs, different results. Reach for a workflow when the task needs an agent, and your hardcoded paths break the first time the input shifts. Most production systems blend: deterministic code for the well-understood steps, model decisions where the input is messy.

When you don't need an agent

Every explainer covers what agents can do. Almost none cover the opposite, so here it is — LangChain's field-tested advice, compressed:

  1. Start with one agent and good prompts. A well-prompted model with a handful of tools solves a surprising share of real problems.
  2. Add tools before adding agents. If yours lacks a capability, give it a tool. A tool is cheaper and far easier to debug than splitting reasoning across a team.
  3. Go multi-agent only at clear limits. Context overflow, capability sprawl, team boundaries — those justify sub-agents and routers. Until one of them actually bites, one agent is the right answer.
  4. Prefer retrieval over reasoning. A good retrieval pipeline outranks a mediocre agent on most knowledge tasks. (If your problem is "the model doesn't know X," you want RAG, not an agent.)
  5. Don't outsource judgment you can't evaluate. If you wouldn't recognize a correct answer, neither will your agent — you'll just be automating blind trust.

My own addition after watching teams adopt agents: if you can't name the step where the model must choose, you don't have an agent use case. You have a workflow, and a workflow is fine.

The real cost: 80% unglamorous work

Here's the number that doesn't make it into product launches. When MIT researchers deployed an AI agent to detect adverse events in cancer patients from clinical notes, the biggest challenge wasn't prompt engineering, and it wasn't model fine-tuning. The 2025 paper found that 80% of the work was consumed by data engineering, stakeholder alignment, governance, and workflow integration — the unglamorous parts.

The same MIT piece punctures the ROI math. "Just because an agentic AI model reclaims 20% of someone's time, that doesn't mean it's a 20% labor-cost savings." Reclaimed time only turns into savings if something else absorbs it.

The failure modes are specific, not sci-fi:

  • Infinite loops. An agent that can't reflect can call the same tool repeatedly, forever. IBM's mitigation is blunt: build in interruption.
  • Shared weaknesses. Multiple agents on the same foundation model inherit the same blind spots — one flaw can take down the whole chain.
  • Permissions and accountability. An agent that acts needs access; access needs audit logs, and high-stakes actions (mass emails, financial transactions) need human sign-off. A rogue agent rejecting a mortgage on faulty data does more damage than a hallucinated paragraph.

None of this means agents aren't worth it. JPMorgan uses them for fraud detection and loan-approval pipelines; Walmart for personal shopping and merchandise planning; one legal research assistant built on IBM's stack cut contract review from 90 minutes to 45 by routing simple queries to a cheap classifier first. It means the budget should include the boring 80%.

Getting started, without the hype

If you're an individual or a small team, the entry path is smaller than the marketing suggests:

  • Use before you build. The major chat products already ship agentic features — deep research, task execution, code agents. Run your real work through them for a month and note where they actually help.
  • One agent, few tools. When you do build, resist the multi-agent org chart. One agent with three well-chosen tools beats a committee of five agents with overlapping tools.
  • Watch the loop, not the demo. Demos show the happy path. Ask to see the trace — the full list of steps the agent took — before you trust it with anything that matters.

Quick answers

What is an agent, in one sentence? A system where an LLM decides the next step itself — choosing tools and actions in a loop until the goal is done, instead of only replying to prompts.

What's the difference between an AI agent and a chatbot? A chatbot responds to messages; an agent pursues goals. The chatbot writes the email, the agent finds the clients, personalizes it, and sends it. The test: who decides the next step.

What's the difference between agentic AI and an AI agent? An AI agent is a single software program; agentic AI is the broader field of systems — often multiple agents orchestrated together. Most people use the terms interchangeably, and that's fine until you're architecting one.

How do I actually use AI agents as a normal person? Start inside the chat tools you already use — research modes, scheduled tasks, code assistants are all agents now. Graduate to building only when you hit a repetitive multi-step task you can describe precisely.

Are AI agents reliable yet? On narrow, well-instrumented tasks, yes — that's where the 90-to-45-minute wins come from. On open-ended work, expect the 80% rule: most of your effort goes into data plumbing and guardrails, not intelligence.

The next time someone demos an agent at you, ask the one question that cuts through everything: who decides the next step? If the answer is "the model, repeatedly, with tools" — that's an agent. If it's "the code, and the model fills in the blanks" — that's a workflow wearing the label. Both are useful. Knowing which one you're looking at is what stops you from paying agent prices for workflow results.