A couple of weeks ago, Paul Guyer from The Nexus and I sat down and talked about taxtech, AI, and career choices. But I wanted to share much more than I could fit into one interview. Last week I wrote up the first, extended piece from that conversation: TaxTech Manager vs. Tax Product Manager, which was about who’s allowed to say no to what. This is part two. It’s about the work itself, and what AI is actually doing to it, based on my observations at Uber and other companies.
But let’s start with the basics.
In 2020, EY claimed that a tax function spends 40 to 60% of its time gathering and reshaping data. Not analyzing anything, just moving it from one format to another so somebody or some system downstream can use it. In 2025, a newer EY survey found the same shape of problem under a different label: 53% of time still goes to routine work, and the people doing it said they'd rather spend much less time on those tasks. Why has that number barely moved in 5 years of “AI transformation”?
Because most AI transformation questions start from the job title and work down. As I mentioned in the Nexus interview, one day AI makes the taxtech manager obsolete; the next, she is back. Same with the Tax Product Manager. We can stay in this useless limbo for a long time, or we can change how we approach answering the question. I doubt that when automobiles were replacing horse carriages, people spent their days arguing what the person driving the former would be called and what authority over critical decisions that person would have. So, let’s start by discussing the specific tasks that AI affects, rather than arguing about roles and responsibilities.
The three-property test
I’ve observed at Uber that AI replaces tasks with three properties:
high volume
a ground truth you can check, and
cheap verification
Here are some examples:
Take document extraction and first-pass classification, the stuff a TaxTech Manager clears out most weeks. Plenty of volume, you know what right looks like, checking one takes seconds. Now look at the product side of the same chart. Generating test cases for a tax determination API. Triaging the support tickets that pile up when a jurisdiction changes a rate overnight. All of these have the three properties and are fully automated at Uber.
Looking through the second half of the table gives a different picture. Tasks like permanent establishment characterization, negotiating with an authority, a Tax PM prioritizing a roadmap amid conflicting regulatory deadlines, or signing off on a launch that quietly creates a tax fact are as far from AI automation as it gets. A human is still making a call in all those areas and living with it.
None of that splits along job titles, which most of the conversation is focused on. A TaxTech Manager and a Tax PM each carry a stack of tasks that clear the bar and a stack that doesn’t. Which pile a task lands on has nothing to do with seniority or authority. It comes down to whether the task includes repetition, checkable facts, and cheap verification, or a judgment call.
What the Tax Authorities Say
Turns out three tax authorities published AI guidance this year. Not together, not even close in timing. And they all said basically the same thing, which is a clear pattern.
HMRC went first, back in January, and interestingly they wrote it for the software vendors, not the practitioners. AI “must support, never replace, human judgement,” and the product itself has to remind the user they’re still the one on the hook for the return.
Then in June the IRS’s Office of Professional Responsibility put out its own guidance under Circular 230. Practitioners “remain fully responsible for the accuracy of their work.” Whatever the AI drafted is “a starting point only” until you’ve gone through it yourself, line by line.
Australia followed in July. Their Tax Practitioners Board said practitioners are “still ultimately responsible for the tax agent services they provide.” Same conclusion as the other two, in different words.
So three regulators, three separate documents, no coordination that I know of, and the same sentence every time. This is by far the most solid argument of what AI will automate and what will definitely stay in the near future.
My Framework For AI Automation: Score It Before You Automate It
Here’s how to actually run this on a real backlog. Take every task and score the three properties one at a time: how much volume is genuinely there, how checkable the ground truth really is, not how checkable you’d like it to be, and what it costs to verify the output every time, not just the first time. Force a number or a plain yes/no on each one before you move to the next task:
Then use the pattern of passes and fails to decide what happens next, because the four patterns need four different responses, not one automate-or-don’t switch:
1️⃣ All three clear: automate it properly. Build the real pipeline. This is where the volume actually is, so it’s worth the engineering time.
2️⃣ Volume and ground truth clear, verification doesn’t: fix verification first. Build the check as its own piece of infrastructure before you touch automation. Once checking is cheap, the automation part gets easy on its own.
3️⃣ Ground truth doesn’t clear, whatever else does: leave it with a person. No amount of volume or cheap tooling turns a judgment call into a checkable fact. This bucket doesn’t get smaller with better AI. It gets smaller with better regulation.
4️⃣ Volume doesn’t clear: don’t build a pipeline for it. Use AI on it anyway if you want, supervised, one task at a time, the way I do. Just don’t put engineering time behind something that happens twice a quarter.
Of course, this framework comes from one PM’s backlog at one company, checked against three regulators who so far agree with each other. Tell me where it breaks.
If you’ve actually automated something that lives in bucket three, I want to hear how. Right now I don’t think it’s possible unless you change the ground truth itself. Better tooling alone won’t do it. And if you think the “human stays the backstop” conclusion from three regulators won’t hold once the guidance catches up with the technology, argue it in the comments. I read and respond to all of them.
Which bucket is your team currently arguing about right now? 👇



