Which work to hand to an AI agent first

In short: The first process you automate should be the one that repeats most often, can be described in writing, has a survivable cost of error, and whose current cost you can measure. Applied honestly, that test disqualifies the customer-facing chatbot most companies start with, and points at unglamorous internal work that pays for itself within a quarter.

Most companies choosing their first AI project pick the one that will look best in a demonstration. It is usually customer-facing, usually impressive, and usually the worst possible starting point — because it is the one where a wrong answer is most expensive and where success is hardest to measure.

There is a better order, and it comes down to four questions asked in sequence.

Question What a good answer looks like What it decides
How often does it happen? Dozens of times a week, not once a quarter Whether you learn fast enough to improve it
Could a new hire do it from written instructions? Yes, from one page Whether the process exists in a transferable form
What does a wrong answer cost? Recoverable within a day How much human review stays in the loop
Can you measure it today? Hours, error rate or delay, already known Whether you can prove it worked

1. How often does it happen?

Frequency beats drama. A task performed forty times a week by three people is worth more automated than a quarterly exercise that takes a week, even though the quarterly one feels larger.

Two reasons. The obvious one is arithmetic: small savings multiply. The less obvious one matters more — a high-frequency process gives you a fast feedback loop. You will find out within days whether the automation is right, and you will have forty examples a week to improve it against. On a quarterly task you get four data points a year, which is not enough to learn from before everyone loses interest.

Start where the repetitions are.

2. Could a competent new hire do it from written instructions?

This is the sharpest test we know, and it is really a question about whether the process exists in a form that can be handed over at all.

If you could write a page of instructions and a capable newcomer would produce acceptable work, an agent can probably be made to do the same. If the newcomer would need to ask three colleagues what the company usually does in this situation, then the process lives in people's heads, and the agent will not ask — it will invent something plausible and present it with confidence.

That failure mode is worth being precise about. The model is not lying; it was given the what and left to infer the why, and inference fills the gap. The fix is not a better prompt. The fix is writing down the judgment that was never recorded.

Most companies discover here that the useful first deliverable is not software. It is the specification.

3. What does a wrong answer cost?

This question does not decide whether to automate. It decides how much human review stays in the loop, and for how long.

Sorting an inbox wrongly costs a minute. Sending a client the wrong contract clause costs considerably more. Both can be automated; they need different shapes:

  • Low cost of error — the agent acts, a person spot-checks a sample.
  • Medium — the agent prepares, a person approves before anything leaves the building.
  • High — the agent assembles the material and the reasoning, a person makes the decision and is accountable for it.

The pattern that works in every one of these is the same: the agent does the preparation, the human keeps the judgment. That is also the arrangement people accept. Teams do not resist automation because it threatens them; they resist it when they are made accountable for output they did not see.

4. Can you measure what it costs today?

If nobody can say how long the work currently takes, how often it is wrong, or how much it delays something downstream, then you will not be able to demonstrate that the automation helped. The project will be judged on impressions, and impressions favour whoever is most enthusiastic in the meeting.

Spend the week before the build measuring the before. It is the least interesting part of the project and the reason anyone will believe the after.

A caution from the research literature here: measured productivity and perceived productivity come apart more often than people expect. There are credible studies in which participants using AI tooling believed they had been substantially faster while the measurements showed the opposite. Whatever you think of any individual study, the implication is practical — if you rely on how the team feels about the tool, you will sometimes be measuring enthusiasm.

What this rules out, on purpose

Applied honestly, this framework disqualifies most first projects that get proposed. It rules out the customer-facing chatbot, because the cost of error is high and the process was never specified. It rules out the quarterly analysis, because you will not learn fast enough. It rules out anything where nobody internally will own the output.

What tends to survive is unglamorous: moving data between systems that were never designed to talk to each other, drafting the reply that is written forty times a week against precedent that already exists, assembling the same report from the same five sources every Monday.

These are the projects that pay for themselves within a quarter, and they are the ones that build the internal confidence to attempt something harder next.

The order, briefly

  1. Find the work that repeats most.
  2. Check whether it could be written down. If not, write it down — that is the project.
  3. Decide where the human stays accountable, and build that in from the start.
  4. Measure the before, so the after is not a matter of opinion.

None of this is about which model you use. By the time that question matters, the hard part is already done.

Frequently asked questions

What should a company automate with AI first?

The work that repeats most often and can be described in writing. High frequency gives you a fast feedback loop and multiplies small savings; a written specification is what stops an agent from inventing the reasoning nobody recorded. Impressive, customer-facing projects usually fail both tests.

Why do most AI pilots fail?

Rarely because of the model. Three causes dominate: the process was never specified, so the agent infers the reasoning and gets it plausibly wrong; nobody is accountable for the output, so quality drifts until trust collapses; or the tool sits outside the systems people already work in, so it is quietly abandoned.

Should a human still review what an AI agent produces?

Yes, and the amount of review should scale with the cost of a wrong answer. Low cost: the agent acts and a person spot-checks a sample. Medium: the agent prepares and a person approves before anything leaves the building. High: the agent assembles the material and a person makes the decision and owns it.

How do you measure whether an AI automation actually worked?

By measuring the before. Record how long the work takes, how often it is wrong and what it delays, in the week before the build. Without that baseline the project gets judged on impressions — and measured productivity and perceived productivity are known to come apart.

Do we need to choose a model before starting?

No. By the time the model choice matters, the hard work — specifying the process, deciding where a human stays accountable, and establishing the baseline — is already done. Companies that start with the model selection usually have not done any of it.

About the author

Filip Salamon

Filip Salamon

CEO / CTO, Salamon Capital

Filip has spent his career between media and technology — filming for ŠKODA, Pilsner Urquell and Range Rover, then co-founding a startup in San Francisco, working with a YC-backed company and serving as CIO at Renato. He built ZEUS Legal AI and now runs the systems Salamon Capital operates on.

BACKGROUND

  • Founder, ZEUS Legal AI
  • Former CIO, Renato
  • 15+ years across media and technology

IN PRACTICE

We ran this order on our own company before we ran it for anyone else.

SEE OUR WORK →

MORE READING

Research and analysis on measurement, automation and what the evidence actually shows.

ALL INSIGHTS →

SERVICES

Growth and acquisition, custom AI and automation, and legal through our own law office.

WHAT WE DO →

WORK

What we built for clients and for ourselves, and what measurably changed.

CASE STUDIES →