Jun 16, 2026

When an AI agent is the wrong answer

Agents are useful for a specific kind of work. Here is how to tell whether your problem is one of them, and what to build instead when it isn't.

A lot of the agent projects we are asked to scope should not be agents. That isn’t a criticism of the people asking. “Agent” has become the default word for any software that uses a language model, and the market is full of demos that make everything look like an agent problem. But an agent is a specific design with specific costs, and picking it for the wrong job gives you a system that is slower, more expensive and less reliable than the boring alternative.

This is the checklist we use before recommending one.

What an agent actually is

For this article, an agent is a program where a language model decides which steps to take next. It is given a goal, a set of tools (search this database, send this email, update this record) and some instructions. It then loops: look at the situation, choose a tool, observe the result, repeat until done or stuck.

That flexibility is the whole point. It is also the source of every problem. A model choosing its own steps can choose wrong steps, take a different path on identical inputs, and run up costs in loops you didn’t anticipate. You accept those risks because the task genuinely needs judgement at run time. If it doesn’t, you are paying for flexibility you won’t use.

Sign 1: you can write the steps down

Ask the people who do the work to describe it. If the answer is a fixed sequence (“take the order number from the form, look it up, check the status, send template B if shipped and template C if not”), you have a workflow, not an agent problem.

Build it as a workflow. A script, a scheduled job or a no-code automation will run the same way every time, cost almost nothing per run and fail in ways you can predict. If one step needs a model (say, classifying a free-text message into one of five categories), call the model for that step only, with a constrained output. That is a workflow with an LLM step, and it is the right shape for most business automation.

Sign 2: the inputs are already structured

Models are good at turning messy input into structure: emails, PDFs, call transcripts, scanned forms. If your input is already a form submission, a database row or an API payload, most of the reason to use a model has gone. Validation rules and a lookup table will beat a model on accuracy, speed and cost.

We often see teams adding a model to interpret data that they themselves control upstream. The better fix is usually to change the form.

Sign 3: a wrong answer is expensive and hard to spot

Agents make mistakes. The question is whether those mistakes are cheap, visible and reversible.

A drafted reply that a person reviews before sending is cheap and visible. An extracted invoice total that is checked against the purchase order is visible. A refund issued automatically, a record silently overwritten, or a message sent to a customer without review is none of those things.

If the task’s failure mode is costly and quiet, either keep a person approving every action (in which case be honest that you have built an assistant, which is fine) or don’t use a model for that decision at all.

Sign 4: volume is low

Building a reliable agent is real work. You need an evaluation set of real examples with correct answers, guardrails, logging, a review interface and someone to watch it in production. That effort pays off when the task happens hundreds or thousands of times a month.

If it happens twenty times a month and takes ten minutes each, the honest answer may be a better checklist, a template, or a shared inbox rule. Or a person with a good chat assistant open in another tab.

Sign 5: you can’t say what “correct” means

Before building, we ask for fifty real examples and the right answer for each. If the team can’t agree on the right answers, the agent won’t find them either. It will produce plausible output that nobody can evaluate, and the project will end in an argument about whether it “works”.

This is not a reason to give up. It is a sign that the process needs defining first. Often that definition work alone removes most of the pain.

When an agent is the right answer

Agents earn their complexity when most of these are true:

  • The inputs are unstructured and varied: free text, documents, conversations.
  • The right next step depends on what you find, and the possible paths can’t be listed in advance.
  • The task is frequent enough to justify building evaluation and monitoring.
  • Mistakes can be caught by a review step, a downstream check or a cheap undo.
  • You can collect examples of correct outcomes and measure against them.

Inbound request triage, document-heavy intake, research and summarisation across several systems, and first-draft responses that a person approves all fit this well.

What to build instead

When an agent isn’t the right fit, the alternatives are usually simpler:

  • A deterministic workflow with one or two model calls for the steps that need language understanding.
  • A better form or interface so the data arrives structured in the first place.
  • An internal tool that puts the right information in front of a person, so their decision takes seconds instead of minutes.
  • A report or alert that surfaces the cases needing attention, instead of software trying to handle them.

None of these sound as exciting as an agent. They tend to ship faster, cost less to run and keep working when the model provider changes something.

A practical way to decide

Take the process you want to automate and mark each step as one of three things: fixed rule, judgement, or exception. Automate the fixed rules with ordinary code. Look at the judgement steps and ask whether a model, constrained to a narrow output, could make that call well enough given a review step. Leave the exceptions to people, and make sure the software routes them quickly.

If, after that exercise, you still have a chain of judgement steps where the path depends on what each one finds, then you have an agent problem. Build it narrow, measure it against real examples, and give it more autonomy only as it earns it.

Have something you want built?

Tell us what you are trying to do and where it is stuck. We will reply with questions, an honest view on the right approach, and next steps.

Start a conversation