Skip to content
PRAXIS

G.32Guides · Decision brief

Which work AI can take on, and which it cannot

Whether AI can take on a piece of work depends on the task, not on the job title. A task is a good candidate when its output can be checked, its mistakes are cheap to reverse, and someone owns the exceptions; fail any one of the three and the task needs a person in the loop.

A modern facade of layered panels and glass seen from below, an illustrative image for a whole assembled from separate parts.

The frame

What is being decided?

A task is a good candidate for AI when its output can be checked, its mistakes are cheap to reverse, and someone owns the exceptions; fail any one of the three and the task needs a person in the loop. That is the argument of this page. It is offered as reasoning for a reader to test against their own work, not as a scoring system or a published method, and the test is meant to be run by the person who knows the task.

The reason to sort by task is that the question people actually ask, can AI do my accounts, or my customer service, or my hiring, has no single answer. Each of those is a bundle. Some of the tasks inside pass the test as they are, some pass with one change, and some should stay with a person however good the tooling becomes. The free consultation Praxis offers covers the same ground out loud: which jobs AI can do now, which it cannot do well yet, and which should never be handed to a machine.

  1. 01

    Can the output be checked?

    Checking means someone who did not produce the output can tell whether it is right, quickly, without redoing the work. Sorting an incoming enquiry into a category can be checked at a glance: the category is visibly right or wrong. Deciding whether a customer deserves a refund cannot, because the only way to check the judgment is to make it again.

    Where the output cannot be checked, the machine is not the problem. The problem is that nobody would know if it were wrong, and a mistake nobody can see is a mistake that repeats.

  2. 02

    Is a mistake cheap to reverse?

    This is a question about what happens next, not about how often the mistake occurs. A wrong label on an internal record is undone in a second. A wrong figure in a message already sent to a customer cannot be unsent, and it may have been acted on. The cost of the mistake is set by whether the output leaves the building before anyone has looked at it.

    That is also where the cheapest improvement sits. Many tasks fail this question only because the output goes straight out. Put a checkpoint between the draft and the send, and the same task becomes one where a mistake costs a glance.

  3. 03

    Does someone own the exceptions?

    Every task that is handed over still produces cases it cannot settle: the odd order, the unusual request, the document that did not arrive. Someone has to receive them, and that someone needs a name, not a shared inbox. A task with no owner for its exceptions does not fail loudly. It fails by the exceptions piling up where nobody looks, which is worse than the task being done by hand.

    Owning exceptions also means having the time for them. An exception pile that lands on a person whose week was already full is a hidden cost of the handover, and it belongs in the sum as hours, as the comparison with hiring does.

  4. 04

    Two tasks, one change apart

    Take two tasks that are invented for this illustration, in an invented small business. The first is sorting incoming enquiries by type and sending a standard acknowledgement. The output is checked at a glance, a wrong sort is corrected in seconds, and the owner of the leftover pile is the person who reads the enquiries labelled other. It passes all three as it stands.

    The second is sending each customer a quote. A person can read a quote, so the output can be checked. But a wrong figure sent to a customer is not cheap to undo, so the task fails the second question. The single change is a checkpoint: the system prepares the draft, a named person approves it, and only then does it go out. The task now passes, and a person remains in the loop by design, not by accident.

  5. 05

    Where the answer is no

    Some work fails more than one question and no checkpoint fixes it. Decisions about individual people are the plainest example: the output cannot be checked without making the decision again, and a mistake lands on a person. In hiring, that is the reason software may sort, flag and prepare, while the decision that someone is not moving forward stays with a named human being, as the hiring page sets out.

    Refusing a task is a result, not a failure of the method. A sort that returns nothing to hand over is useful, because it stops money being spent on automating something that would have cost more to supervise than to do.

  6. 06

    What this test does not say

    It says nothing about any named model or product, because what those can do changes faster than a guide can be updated, and a capability claim written today would be stale soon. It is a way of asking about the task, which stays the same while the tools change.

    Nor does it say how much a handed-over task is worth. That is a separate question about hours and baselines, and the savings estimator is where it starts. Passing the test makes a task a candidate. It does not make it worth doing.

Side by side

The 3 questions, and what failing each one asks for.

DimensionA task that passesA task that failsWhat failing asks for
Can the output be checked?A sorted enquiry, which is visibly right or wrongA judgment that can only be checked by making it againA person makes or confirms the decision
Is a mistake cheap to reverse?A label on an internal recordA figure already sent to a customerA checkpoint between the draft and the send
Does someone own the exceptions?A named person who reads the leftover pileA shared inbox nobody is responsible forName an owner, and count the hours they will need
Does it pass all three?A candidate for handing overA task that fails two or moreLeave with a person; use tools to prepare, not to decide

The call

What to do with a task list

  1. 01

    Take the work apart into tasks first.

    A job title passes or fails nothing. Its tasks do, one at a time, so the sorting starts from the list and not from the department.

  2. 02

    Ask the 3 questions of each task, in writing.

    Can the output be checked, is a mistake cheap to reverse, and who owns the exceptions. Write the answer down, including the name of the owner.

  3. 03

    For each failure, look for the single change that fixes it.

    Often it is a checkpoint before the output leaves. If no single change fixes it, the task stays with a person.

  4. 04

    Only then put hours and a baseline against the candidates.

    Passing the test makes a task a candidate. Whether it is worth doing is a question about how many hours it takes now, measured from a system that records the work.

A note on interest. Praxis sells consulting, so treat this page as an informed party’s brief, not a referee’s ruling. The discipline we hold ourselves to is written down: category-level comparisons only, no named competitors, and a public page on when we are not the right fit.

Questions

Asked before scoping.

Is this the method Praxis uses to assess a business?
No. It is an argument, set out so that a reader can test it against their own work and disagree with it. It is not a published scoring system, and no step of it is presented as how an engagement is run. What Praxis does commit to publicly is the conversation on the free consultation about what AI can do now, what it cannot do well yet, and what should stay with people.
Can AI do my accounts, my customer service or my hiring?
Parts of each, and not the whole of any. Each is a bundle of tasks that pass the test to different degrees, which is why the function pages, such as customer service, finance and sales and CRM, are written task by task. The sorting on this page applies to all of them and sits above any one of them.
What if a task fails only one of the three questions?
Look for the single change that makes it pass, which is often a checkpoint before the output leaves or a named owner for the exceptions. If one change fixes it, the task is a candidate with a person in the loop. If no single change does, the task stays with a person, and tools can still be used to prepare the work for them.
Does a task that passes all three have to be automated?
No. Passing makes it a candidate and nothing more. A task that takes ten minutes a month passes every question and is still not worth building anything for. The test says whether a task can be handed over, and the hours and the baseline say whether it should be. The readiness scorecard is a reasonable next step for the wider picture.

Clearer on what you are deciding?

Then the next conversation is about fit and scope. Tell us what you are deciding, and we will tell you honestly whether we are the right resource.

No obligation · a scoping conversation first