AgentsUse

How to Choose an AI Agent for a Specific Task

A job-first framework for choosing an AI agent: score five criteria, run a same-task test, and cost it with a worksheet before you commit.

Published 2026-10-04 · 15 min read

Most people choose an AI agent backwards. They watch a demo, read a launch post, and then go looking for a problem the agent can solve. A month later they have a subscription they barely use, or worse, an agent wired into real work that fails in ways nobody tested for. This guide flips the order around. You start with one specific task, write down exactly what done looks like, and only then look at candidates. The framework below has five parts: define the job, score the candidates on five criteria, run a small same-task test, cost the winner with a worksheet, and set the review point before anything touches real work. It takes an afternoon. It saves you from the two most common outcomes in agent adoption: paying for capability you never needed, and trusting an agent with work it was never checked on.

Step one: write the job down before you look at any product

An agent choice is only as good as the job description behind it. Write down the task in plain language, in this order: the input it starts from, the output it must produce, how often it runs, and what happens when the output is wrong. Most people skip this because it feels obvious. It is not. Two people who both say I need an agent to handle research usually mean different things: one wants a brief with sources they can check, the other wants a summary they can paste into a report. Those are different agents, different costs, and different risks.

Keep the job description to one task. Agents that do everything exist mostly in marketing. A candidate that is genuinely strong at one named task will beat a generalist on that task almost every time, and when the generalist fails you will not know which part of its job it failed at. If you have five tasks, run this framework five times. It sounds slow. It is faster than untangling a tangled setup later.

  • Name the task in one sentence that includes the input and the output: turn these three supplier invoices into rows in this spreadsheet, every Friday.
  • Say how often it runs and roughly how much volume it carries. Once a week is a different product decision than four hundred times a day.
  • Write the cost of a mistake in plain words: a wrong number in a client report, a mis-filed invoice, a reply sent to the wrong person.
  • List what the agent is not allowed to touch, out loud, before you test anything: no sending, no paying, no deleting, no publishing.

You now have the only document that matters for the rest of this process. Every criterion below, every test, and every cost line gets judged against this job description and nothing else. When a vendor demo shows the agent doing something impressive that is not on your job description, that is a distraction, not evidence.

Step two: score candidates on five criteria

Once the job is written down, most candidates eliminate themselves quickly. Score each one on these five criteria, in this order. A candidate that fails an early criterion does not get rescued by a later one.

Criterion one: can it actually do the task shape?

Not the demo shape, the task shape. Your task has a specific input (a PDF, an inbox, a web page, a spreadsheet) and a specific output (a file, a draft, a row, a shortlist). Ask whether the candidate accepts that input and produces that output through a real mechanism: a tool it can call, a file it can write, a system it can reach. If the answer involves you copying and pasting between the agent and the real system, you are not evaluating an agent, you are evaluating a chatbot with extra steps. That can still be the right choice for low volume, but name it honestly, because your cost and error expectations will be completely different.

Criterion two: can you check the output faster than doing the task?

This is the criterion that decides whether the agent saves time at all. Time yourself doing the task once. Then ask how long checking a finished output would take. If checking takes nearly as long as doing, the agent adds coordination cost without removing work. Good agent tasks have checkable outputs: a table you can scan, a draft you can skim, a brief with sources you can spot-check. Bad ones produce smooth, plausible work you have to redo mentally to verify. Prefer candidates whose output format makes checking cheap, and treat output format as a feature worth paying for, because it is the difference between supervision that takes two minutes and supervision that takes twenty.

Criterion three: what does one run cost, in full?

Not the headline price, the full cost shape. That includes the model calls the agent makes along the way, any tools or APIs it calls that bill separately, the human minutes spent reviewing, and the cost of the runs that fail and have to be repeated. Agents that take many steps cost more per task than they appear to, because every step is work billed or time spent. You will build the actual numbers in the worksheet below. At the scoring stage, just note the shape: is pricing per seat, per task, per token, per tool call, or some blend? A pricing shape that punishes your exact usage pattern is a real disqualifier, even when the product itself is strong.

Criterion four: what can it touch, and can you fence it?

List the systems the candidate would reach: files, inboxes, calendars, databases, the public web, payment tools. Then ask what fences exist. Can you restrict it to one folder, one inbox, one account? Can external actions be set to draft-only? Can you cap spending or the number of steps per run? A candidate with strong output and no fences is a candidate you can only use on work that does not matter. Match the fence to the stakes: the closer the task sits to money, customers, or your name, the harder the fences need to be, and if a candidate cannot be fenced hard enough for its task, that is your answer.

Criterion five: what does it do when it is stuck?

Every candidate will eventually meet an input it cannot handle: a missing file, an odd format, a step that fails. The question is what it does next. The good behavior is narrow and boring: try one reasonable alternative, then stop and report where it got stuck and what it needs. The bad behaviors are quiet guessing, silent skipping, and confident continuation down a wrong path. You cannot learn this from a feature list. You learn it in the test, which is the next step, and it is the step people skip most often.

Step three: run a same-task test on your own inputs

Do not trust any comparison that did not run your task on your inputs. Vendor benchmarks, launch demos, and other people's write-ups all test somebody else's job. A small test of your own beats all of it. Here is the routine, sized for one afternoon.

  • Pick five to ten real past examples of the task: old invoices, old requests, old documents. Use work you already finished, so you know what a correct output looks like.
  • Run each candidate on exactly the same examples, with exactly the same instructions, in the same order. Changing the instructions between candidates makes the comparison meaningless.
  • Score every output against a short rubric you wrote before the test: correct, checkable in under two minutes, flagged its own uncertainty, stayed inside the fences.
  • Note the stuck behavior separately. Deliberately include one example with something missing or malformed and watch what each candidate does with it.
  • Time the checking, not just the running. The candidate whose output you can verify fastest is often the right choice even when raw speed favors another.

Keep the test small on purpose. You are not certifying the agent forever; you are finding out whether it handles your task shape, on your inputs, well enough to earn a closer look. Five honest examples will show you the failure pattern that matters. Fifty would mostly show you the same pattern fifty times. If a candidate cannot get through five of your real examples without a serious failure, it does not matter what it scores anywhere else.

One more rule for the test: write the rubric before you run anything. A rubric written after you have seen the outputs quietly bends toward whichever candidate you liked. Four lines is plenty. Did it produce the right output? Could you check it quickly? Did it say so when it was unsure? Did it stay out of everything you fenced off? Score each output line by line and keep the scores. They become the baseline you compare against later, when the agent has been running for a month and you want to know whether it is getting better or worse.

Step four: cost it with a worksheet before you commit

Agent costs surprise people because the sticker price is only one line. Use this worksheet shape for each finalist. The figures in the worked example below are hypothetical and invented for illustration, not quotes for any real product. Replace them with the real numbers from your test runs and the pricing pages you checked yourself.

  • Subscription or seat cost per month, if any, divided by the number of tasks you will actually run that month.
  • Per-run usage cost: model calls and any separately billed tool or API calls, measured from your test runs rather than estimated from a pricing page.
  • Review cost: your checking time per run, at an honest rate for your own time, multiplied by runs per month.
  • Failure cost: the share of runs that fail or need redoing, multiplied by the cost of fixing one, including your time.
  • Setup cost, paid once: your hours to configure, write instructions, and build the fences, counted honestly rather than ignored.

Here is the worksheet filled in as a worked hypothetical. Imagine a weekly task: turning supplier invoices into spreadsheet rows, twenty invoices a week. Candidate A is a hosted agent at a flat monthly seat price. Candidate B is a build-it-yourself setup where you pay per model call. The numbers are invented round figures chosen to make the arithmetic visible, and they are labeled hypothetical because they are not prices for any real product.

For Candidate A, suppose the seat costs a hypothetical 40 dollars a month and covers the runs. Checking takes two minutes per invoice batch, runs weekly, so roughly forty minutes of review a month. If you value your time at a hypothetical 50 dollars an hour, review costs about 33 dollars a month. Failures: suppose one batch in four needs a partial redo costing fifteen minutes, roughly another 12 dollars a month in expectation. Setup takes three hours once, about 150 dollars of your time, paid a single time. Monthly running total once set up: about 85 dollars of cash plus your time counted, against doing the task by hand at perhaps two hours a month.

For Candidate B, suppose there is no seat price but each invoice costs a hypothetical 30 cents in model and tool calls: twenty invoices a week is about 24 dollars a month. Review is slower because the output format is plainer, say four minutes a batch, so about 27 dollars a month of your time at the same rate. Failure handling is on you either way, but configuration is heavier: suppose six hours to set up, about 300 dollars of your time once. Monthly running total: about 51 dollars, cheaper per month, slower to set up, and the bill grows directly with volume while Candidate A stays flat until you cross a plan limit.

The worksheet's job is not to crown a winner. It is to make the trade visible: flat cost with faster checking versus usage cost with slower checking and heavier setup. Notice that review time is a real line in both columns and it is bigger than the usage line in one of them. That is normal, and it is why criterion two matters more than most pricing pages want you to think. If neither candidate beats doing the task by hand once review and failure costs are counted, the honest answer is to not use an agent for this task yet, and that is a perfectly good outcome for an afternoon of testing.

Step five: decide, then set the review point before real work

Pick the candidate that passed the five criteria, survived your test examples, and costed out honestly. Then, before it touches real work, set the review point in writing: who looks at the output, at what stage, and what that person is checking for. Start every new agent task in draft-only mode, where the agent proposes and a person disposes, and widen from there only after a run of clean reviews. The graduation rule is simple: when reviews have been clean long enough that checking feels boring, spot-check instead of reviewing everything. When a failure shows up, tighten back up and find out why before widening again.

Write down what would make you stop using the agent for this task. A cost ceiling, an error pattern, a change in the task itself. Deciding the exit rule in advance sounds pessimistic. In practice it is what lets you widen the agent's scope calmly later, because you already know the conditions under which you will take the task back, and so does everyone working with you.

Keep the whole selection in one short document: the job description, the five criterion scores, the test results, the worksheet, and the review point. It fits on two pages. That document does three jobs after the decision is made. It tells a new teammate why this agent has this job and where the fence sits. It gives you the baseline when you rerun the test set after a version change. And it is the honest record you want when someone asks, six months in, why the team pays for this tool. Selections made by demo and memory get relitigated. Selections made on paper hold.

Mistakes that wreck an otherwise good selection process

  • Choosing on demo tasks instead of your own past examples. Demos are chosen to flatter the product. Your old work is chosen by reality.
  • Testing with different instructions for each candidate, then comparing scores. You measured your instructions, not the agents.
  • Counting only the subscription in the cost. Review and redo time usually decide the real winner.
  • Skipping the malformed-input test. The stuck behavior you never tested is the failure you will meet in production.
  • Granting broad access on day one because setting up fences felt slow. Fences are the product. An unfenceable candidate is a different, riskier product than its feature list suggests.
  • Picking a generalist for a specific job because the roadmap mentions your task. Judge the task in front of you, not a future release.

Where to look once you know the job

With the job description written, the directory on this site becomes more useful, because you can read a profile as an answer to your criteria instead of as a feature list. The tool categories group products by the job they do, such as browser automation and web data extraction, so you can shortlist by task shape first. Inside a profile, look for the setup path, the limits, and the alternatives, and read them against your five criteria: input and output fit, checkable output, cost shape, fences, and stuck behavior. A profile that names its limits plainly is worth more to this process than one that promises everything, because your test still has to confirm the claims that matter.

Questions people actually ask

How many candidates should I test? Two or three. One gives you no comparison. Five turns selection into a project of its own, and the differences between finalists are usually visible by the third. Shortlist by criteria one and four on paper, test the survivors, and cost the winner.

The task is small. Is all of this overkill? Scale the process to the stakes, not the clock. For a cheap, reversible task, the job description and a three-example test are enough, and they take twenty minutes. Save the full worksheet and rubric for tasks that touch money, customers, or anything with your name on it. The framework compresses well; skipping the job description entirely is the only step that reliably causes trouble later.

What if no candidate passes? Then the task stays manual for now, and you have learned that in an afternoon instead of a quarter. Keep the job description and the test examples. The next candidate you meet gets judged against the same documents in an hour, and the task itself often turns out to be a good script candidate instead, since writing the evaluation usually reveals that the steps were specifiable all along.

Should I rebuild the test when the task changes? Yes, and cheaply. Keep your five to ten scored examples as a small test set. When the task, the inputs, or the agent version changes, rerun the set before trusting the new setup on real work. Ten minutes of regression testing is what separates a workflow that quietly drifts from one you actually control.

What about agents I build on a framework instead of buying a product? The same five criteria apply, with one addition: you become the vendor, so setup and failure costs land on you entirely. Test your build against the same examples and rubric you would use for a purchased candidate, and be stricter, not kinder, because nobody else will be. A framework comparison is a separate decision from a task decision, and the task decision comes first.

Keep reading