Skip to content
Prompt Consulting
de
agentsimplementation

AI agent vs chatbot: which do you need?

A chatbot answers questions. An AI agent changes records in your systems. How the two differ in cost, risk and upkeep, plus a checklist to pick the right one.

Thilo Krause

A chatbot answers. An AI agent does something with the answer. That one difference decides what the project costs and who has to sign off before it goes live, and vendors blur it because "agent" is the word that sells this year.

Chatbot vendors now call their products agents, and some agent platforms put a chat window in front, so the label on the pricing page tells you little. The useful question is narrower. When the software finishes, has anything changed in one of your systems? If not, you are buying a chatbot, whatever the brochure says. If it has, you are buying an agent, and the cost of an agent build follows different rules from a chatbot subscription.

The difference in two sentences

A chatbot takes a question in plain language and returns an answer, usually drawn from documents or a knowledge base you give it. An AI agent takes a goal, works out the steps and carries them out through the tools and systems it can reach, for example by opening a ticket or updating a CRM record.

Both usually run on the same kind of language model. What separates them is what the model is allowed to touch.

What a chatbot does

A customer asks when their order ships. The chatbot looks the question up in your help pages, maybe in a read-only view of the order status, and writes a reply. The conversation is the product. When the window closes, nothing in your business has moved unless a person reads the transcript and acts on it.

That makes a chatbot cheap to get wrong. A bad answer is embarrassing and occasionally expensive, and the fix is usually a better source document or a tighter instruction. You can launch one in weeks, read the transcripts and correct it as you go.

What an AI agent does

Give an agent the same customer email and it can check the order in your ERP, see that the shipment is stuck, open a claim with the carrier, note it on the customer record and send a reply with the new date. Each of those is a call into a system with write access. The reply is the last of five steps, and the four before it are where the value and the risk sit.

An agent also works without anyone typing at it. Plenty of useful agents never have a chat window. They run when an invoice lands in a mailbox or when the month closes, and the first a person sees of them is an item in a review queue.

Where agents and chatbots differ in practice

Actions in your systems. A chatbot needs read access to content. An agent needs credentials for every system it changes, scoped per system, and a person in your company who is accountable for what it writes. Agreeing that access often takes longer than writing the code.

Multi-step goals. A chatbot handles one turn at a time. An agent breaks a goal into steps, and the result of step two decides step three. It can take a wrong turn early and keep going, so each step needs its own check.

Memory and state. A chatbot remembers the conversation it is in. An agent has to track work in progress: which invoices it has matched, which wait on a human, which failed. That state lives in a database you own, and when the agent restarts halfway through a batch it has to resume without doing anything twice.

Autonomy and approvals. With a chatbot, the person reading the answer decides what happens next. With an agent, you decide in advance which actions run on their own and which wait for a release. I put that line in code and never in the prompt, because a model can ignore a prompt. The piece on why hard limits belong in code shows where they go.

Failure modes. A chatbot fails by saying something wrong. An agent fails by doing something wrong, and some actions cannot be undone. An email that went out stays out. A posted journal entry needs a reversal and an audit trail. Agents therefore need a queue where uncertain cases go to a person, and a way to measure the error rate before they touch real data. I run every new agent in shadow mode next to the team first.

Cost and maintenance. A chatbot's running cost is mostly model usage and keeping the source content current. An agent's cost comes from how many systems it touches, how reversible its actions are and how many cases drop out as exceptions. It also breaks when a connected system changes its API or its data format, so somebody has to watch it after go-live.

Side by side

ChatbotAI agent
What it producesAn answer in a conversationA changed record, a sent message, a finished task
Access it needsRead access to documents or a knowledge baseScoped write access to each system it works in
How work startsSomeone types a questionAn event, a schedule or a request, often with no chat at all
Steps per taskOne question, one answerSeveral, each depending on the last
State it keepsThe current conversationWork in progress across runs and restarts
Who decidesThe person reading the answerRules you set, with approval steps for risky actions
Typical failureA wrong or invented answerA wrong action in a live system
Main cost driversModel usage and content upkeepIntegrations, exception handling and monitoring
Testing before launchReading transcripts, fixing sourcesShadow runs on real cases with a measured error rate

When a chatbot is enough

If the job ends when someone has the information, a chatbot is the right tool, and an agent would buy you integration work you don't need. Answering policy and product questions, helping staff find the right internal document and qualifying website visitors before a person calls them back all fall in this group.

A chatbot is also the better choice when the follow-up actions are rare, varied or need judgement each time. If a person would check every step anyway, give them a good answer fast and let them act on it.

When an agent pays off

An agent earns its build cost when the same multi-step task runs many times a week across systems you already have, and the person doing it spends the time copying data from one screen to another. Matching invoices to purchase orders, routing a shared inbox, chasing missing documents and reconciling accounts at month end fit that description.

Two conditions have to hold. The rules must be clear enough to write down, and the share of cases that need human judgement must be small. If a quarter of the cases need someone to think, an agent removes the typing and leaves the hard part where it was. In that case, fix how the work arrives before you automate it.

A checklist before you decide

Answer these for the task you have in mind:

  • Does the task end with an answer, or with something changed in a system?
  • How many systems does it read from, and how many does it write to?
  • What does a wrong action cost, and can you undo it?
  • How often does the task run each week, and how long does it take a person each time?
  • Could someone write the rules down on two pages today?
  • What share of cases needs a human decision?
  • Who in your company would own the agent's actions and review its exceptions?

If the first answer is "an answer", get a chatbot and stop there. If it is "something changed" and the answers on volume, rules and ownership hold up, you have an agent candidate. Our agent build service starts from one workflow like that and puts it into production on the software you already run.

All notes

Next step

Tell us what your team still does by hand.

Thirty minutes on a call. You describe the work that eats the week. We tell you whether an agent can take it and what building it would cost, including when the answer is that it cannot.

  • Built on your current stack
  • Nothing to migrate
  • Three clients at a time

Analytics and spam protection

We would like to count visits with Google Analytics, and to load Google's spam check on the contact form. Both load only if you accept. Either way we store one entry in your browser so this does not ask again, and the contact form works the same whichever you press.

What we collect, in full