G.32Guides · Decision brief
Which work AI can take on, and which it cannot
Whether AI can take on a piece of work depends on the task, not on the job title. A task is a good candidate when its output can be checked, its mistakes are cheap to reverse, and someone owns the exceptions; fail any one of the three and the task needs a person in the loop.

The frame
What is being decided?
A task is a good candidate for AI when its output can be checked, its mistakes are cheap to reverse, and someone owns the exceptions; fail any one of the three and the task needs a person in the loop. That is the argument of this page. It is offered as reasoning for a reader to test against their own work, not as a scoring system or a published method, and the test is meant to be run by the person who knows the task.
The reason to sort by task is that the question people actually ask, can AI do my accounts, or my customer service, or my hiring, has no single answer. Each of those is a bundle. Some of the tasks inside pass the test as they are, some pass with one change, and some should stay with a person however good the tooling becomes. The free consultation Praxis offers covers the same ground out loud: which jobs AI can do now, which it cannot do well yet, and which should never be handed to a machine.
- 01
Can the output be checked?
Checking means someone who did not produce the output can tell whether it is right, quickly, without redoing the work. Sorting an incoming enquiry into a category can be checked at a glance: the category is visibly right or wrong. Deciding whether a customer deserves a refund cannot, because the only way to check the judgment is to make it again.
Where the output cannot be checked, the machine is not the problem. The problem is that nobody would know if it were wrong, and a mistake nobody can see is a mistake that repeats.
- 02
Is a mistake cheap to reverse?
This is a question about what happens next, not about how often the mistake occurs. A wrong label on an internal record is undone in a second. A wrong figure in a message already sent to a customer cannot be unsent, and it may have been acted on. The cost of the mistake is set by whether the output leaves the building before anyone has looked at it.
That is also where the cheapest improvement sits. Many tasks fail this question only because the output goes straight out. Put a checkpoint between the draft and the send, and the same task becomes one where a mistake costs a glance.
- 03
Does someone own the exceptions?
Every task that is handed over still produces cases it cannot settle: the odd order, the unusual request, the document that did not arrive. Someone has to receive them, and that someone needs a name, not a shared inbox. A task with no owner for its exceptions does not fail loudly. It fails by the exceptions piling up where nobody looks, which is worse than the task being done by hand.
Owning exceptions also means having the time for them. An exception pile that lands on a person whose week was already full is a hidden cost of the handover, and it belongs in the sum as hours, as the comparison with hiring does.
- 04
Two tasks, one change apart
Take two tasks that are invented for this illustration, in an invented small business. The first is sorting incoming enquiries by type and sending a standard acknowledgement. The output is checked at a glance, a wrong sort is corrected in seconds, and the owner of the leftover pile is the person who reads the enquiries labelled other. It passes all three as it stands.
The second is sending each customer a quote. A person can read a quote, so the output can be checked. But a wrong figure sent to a customer is not cheap to undo, so the task fails the second question. The single change is a checkpoint: the system prepares the draft, a named person approves it, and only then does it go out. The task now passes, and a person remains in the loop by design, not by accident.
- 05
Where the answer is no
Some work fails more than one question and no checkpoint fixes it. Decisions about individual people are the plainest example: the output cannot be checked without making the decision again, and a mistake lands on a person. In hiring, that is the reason software may sort, flag and prepare, while the decision that someone is not moving forward stays with a named human being, as the hiring page sets out.
Refusing a task is a result, not a failure of the method. A sort that returns nothing to hand over is useful, because it stops money being spent on automating something that would have cost more to supervise than to do.
- 06
What this test does not say
It says nothing about any named model or product, because what those can do changes faster than a guide can be updated, and a capability claim written today would be stale soon. It is a way of asking about the task, which stays the same while the tools change.
Nor does it say how much a handed-over task is worth. That is a separate question about hours and baselines, and the savings estimator is where it starts. Passing the test makes a task a candidate. It does not make it worth doing.
Side by side
The 3 questions, and what failing each one asks for.
| Dimension | A task that passes | A task that fails | What failing asks for |
|---|---|---|---|
| Can the output be checked? | A sorted enquiry, which is visibly right or wrong | A judgment that can only be checked by making it again | A person makes or confirms the decision |
| Is a mistake cheap to reverse? | A label on an internal record | A figure already sent to a customer | A checkpoint between the draft and the send |
| Does someone own the exceptions? | A named person who reads the leftover pile | A shared inbox nobody is responsible for | Name an owner, and count the hours they will need |
| Does it pass all three? | A candidate for handing over | A task that fails two or more | Leave with a person; use tools to prepare, not to decide |
The call
What to do with a task list
- 01
Take the work apart into tasks first.
A job title passes or fails nothing. Its tasks do, one at a time, so the sorting starts from the list and not from the department.
- 02
Ask the 3 questions of each task, in writing.
Can the output be checked, is a mistake cheap to reverse, and who owns the exceptions. Write the answer down, including the name of the owner.
- 03
For each failure, look for the single change that fixes it.
Often it is a checkpoint before the output leaves. If no single change fixes it, the task stays with a person.
- 04
Only then put hours and a baseline against the candidates.
Passing the test makes a task a candidate. Whether it is worth doing is a question about how many hours it takes now, measured from a system that records the work.
A note on interest. Praxis sells consulting, so treat this page as an informed party’s brief, not a referee’s ruling. The discipline we hold ourselves to is written down: category-level comparisons only, no named competitors, and a public page on when we are not the right fit.
Questions
Asked before scoping.
- Is this the method Praxis uses to assess a business?
- No. It is an argument, set out so that a reader can test it against their own work and disagree with it. It is not a published scoring system, and no step of it is presented as how an engagement is run. What Praxis does commit to publicly is the conversation on the free consultation about what AI can do now, what it cannot do well yet, and what should stay with people.
- Can AI do my accounts, my customer service or my hiring?
- Parts of each, and not the whole of any. Each is a bundle of tasks that pass the test to different degrees, which is why the function pages, such as customer service, finance and sales and CRM, are written task by task. The sorting on this page applies to all of them and sits above any one of them.
- What if a task fails only one of the three questions?
- Look for the single change that makes it pass, which is often a checkpoint before the output leaves or a named owner for the exceptions. If one change fixes it, the task is a candidate with a person in the loop. If no single change does, the task stays with a person, and tools can still be used to prepare the work for them.
- Does a task that passes all three have to be automated?
- No. Passing makes it a candidate and nothing more. A task that takes ten minutes a month passes every question and is still not worth building anything for. The test says whether a task can be handed over, and the hours and the baseline say whether it should be. The readiness scorecard is a reasonable next step for the wider picture.
Where this leads on the site
- How AI implementation works
- The free AI consultation
- Automating customer service
- Automating the back office
- AI automation vs hiring
- AI readiness scorecard
Other decision guides
- G.01Strategy vs management
- G.02Consultant vs contractor vs fractional
- G.03Boutique vs Big Four
- G.04Consultant vs in-house
- G.05Change vs transformation
- G.06Fractional vs retainer
- G.07AI consultant vs implementation partner
- G.08Data strategy vs engineering
- G.09Consulting vs coaching
- G.10Interim vs consultant
- G.11SEO consultant vs agency
- G.12Transformation vs modernization
- G.13What drives cost
- G.14Fee structures
- G.15Fixed vs T&M
- G.16First-engagement budget
- G.17Questions to ask
- G.18Red flags
- G.19Writing an RFP
- G.20Evaluating proposals
- G.21Do you need one?
- G.22Getting the value
- G.23When to hire strategy
- G.24SEO for construction
- G.25Board vs advisory board
- G.26Marketing consultant vs agency
- G.27Piloting an advisor
- G.28AI implementation cost
- G.29Strategy vs ESG reporting
- G.30Checking an AI savings claim
- G.31AI automation vs hiring
Clearer on what you are deciding?
Then the next conversation is about fit and scope. Tell us what you are deciding, and we will tell you honestly whether we are the right resource.
No obligation · a scoping conversation first