Part 3 of the series I’m writing from my Nexus interview with Paul Guyer. Part 2 covered what AI actually replaces in taxtech: tasks, not roles.
When cars replaced carriages, I doubt anyone spent much time arguing about what to call the driver. Instead, somebody still had to think of what the price should be. I’ve been following the same argument right now, about roles impacted by AI in TaxTech. Hardly anyone asks whether the fees for AI use make sense or how to get the best outcome for your buck.
How AI Gets Bought and Capped
Let’s start with how AI gets bought. Microsoft lists its Copilot Business add-on at $18 per user per month. Salesforce lists its Agentforce add-ons at $125 per user per month. Both are priced per person. (Salesforce also sells agents at $2 per conversation, which is a bit closer to paying per task.)
The labs that build the models sell a seat plus tokens. Anthropic lists Claude Enterprise at $20 per seat per month plus usage at API rates, and The Register reported in April that bundled tokens are going away from its enterprise seats as contracts renew. OpenAI lowered ChatGPT Business to $20 a seat the same month, and usage of its coding agent, Codex, is billed in credits priced per token.
Then there’s how companies keep the bill under control. Uber capped its agentic coding tools at $1,500 a month per employee, per tool. SemiAnalysis talked to enterprises in June and heard per-employee limits and budgets anywhere from $250 to $4,000 a month. Meta is moving to formal token budgets in 2027, The Information reported. So the limit is set per person, or per token.
Well, if AI replaces tasks (as I would argue), what’s the correlation between the price plans above and the end user flow? Where's the task in that table and how much ROI are we getting by using AI? Difficult to say. Imagine the following scenario: two invoices go out on the same Monday. One clears in seconds. The other is rejected, and someone spends the afternoon fixing it. A seat or a cap can't tell which was which. The tokens were just used.
What a Task Really Costs
There’s a second problem with counting people (seats). The AI call is only part of a task's cost. Somebody has to check the result, and when it’s wrong, somebody has to redo it (as in the example above).
A June MIT Sloan piece on research by Christian Catalini and his co-authors puts it well: “AI makes it cheap to produce work, but not to judge whether that work is any good.” BetterUp Labs and Stanford looked at what they call workslop, polished-looking AI output that doesn’t hold up. 41% of workers had received some, and each instance cost nearly two hours of rework. That’s general office work, not tax troubleshooting. I wouldn’t bet the latter is kinder.
A cap doesn’t fix any of that. Picture two teams under the same cap. One spent its budget clearing failed invoices that the authority then accepted. The other spent it producing drafts someone now has to redo. Same cap, same spend. How do you determine which one has a higher ROI? Are you paying per seat but still using the plan for workslop or troubleshooting more than for meaningful automations?
The Coming of the Outcome Age
Here’s the gap. We use AI to do tasks, but we pay for it by the seat or the token. Tokens are billed whatever they were spent on: a clean filing, a retry loop, a fix for a mistake in the vendor’s own system. The customer pays for all of it.
Outcome pricing closes the gap, because only the result is billed. Deloitte’s accounting guidance puts it plainly: “unsuccessful attempts, incomplete actions, or results that do not meet the contractual success criteria typically do not trigger payment.” The retries and the vendor’s own flaws come out of the vendor’s margin, not your budget.
Is it actually happening? Yes, but it’s early. The Big Four are already writing the accounting rules for it. Deloitte did in June, and EY published its take on revenue recognition last year. PwC’s Strategy& says it is “seeing a move away from flat, seat-based pricing.” OpenAI’s CFO, Sarah Friar, wrote in January: “Licensing, IP-based agreements, and outcome-based pricing will share in the value created.” McKinsey’s Michael Birshan told Business Insider that about a quarter of the firm’s global fees now come from outcome-based pricing.
It’s also still small. Among AI-native and seat-based software companies adding a hybrid AI meter, Bain found that only about 10% rely on outcome-based meters. A Gartner analyst told CIO Dive the rise is “more buzz than reality”. Where it has “found a real home,” Bain says, is customer support, “where a resolved conversation is observable, attributable, and contractible.”
In the Nexus interview, I said I suspect tax will follow, with customers asking to pay per successful VAT return or per pound recovered. For now, here’s who sells that way, from Intercom and Zendesk to HubSpot and Sierra.
Look at the last column. Each of these prices needs a definition of “done”, and two of the vendors publish theirs. At Intercom, a customer who goes quiet for 24 hours counts as resolved. At Zendesk it’s 72 hours. That’s a stand-in for “the customer was actually helped”, and the vendor picked it. (Intercom and Zendesk also still charge for seats underneath.)
So “outcomes, not seats” comes with a catch. You stop paying for the vendor’s retries and flaws, but you start paying for whatever the vendor decides counts as a result.
How to Prepare for What’s Coming
If you’re the one buying, here’s how to get ready, in the order it happens.
Before the renewal
1️⃣ Map your tasks. List what the AI actually does for you and what each task really costs: the call, the check, and the redo.
In the contract
2️⃣ Make the task the unit, and define “done”. Ask for a single price per verified task that covers the model, the vendor’s support, and integration into your systems. If the answer is “seats” or “tokens,” push back. Put “done” on the order form so you can audit it. If only the vendor defines it, you’re paying for their stand-in.
3️⃣ Cap the total, not the seat. A cap per person won’t help when an agent works all night. Ask for a ceiling on the total task bill instead.
4️⃣ Say who owns which failure. The vendor will push back fairly when the failure is in your master data, so the contract has to specify who owns what.
After you sign
5️⃣ Check a sample. Take some of the tasks billed as done and see whether they really were. If the definition and the results drift apart, that’s your reason to renegotiate it.
Do you agree? Let me know in the comments. 👇



