Stop Measuring AI Adoption
The Wrong Way to Push AI in Tax Teams
No one used to measure how many spreadsheets you built.
No one tracked how many Word documents you opened, how many Outlook emails you sent, or how many Excel formulas you wrote. Tools were invisible. The work was visible. You were judged on the memo, the return, the advice. Not on the software you used to produce them.
That has changed. Tax firms are now rolling out AI adoption dashboards. Partners are asking why usage rates are low. Some teams have internal leaderboards. The tool has become the metric.
This is a mistake. AI usage is a bad KPI. It measures activity, not value. And it quietly punishes the people doing the best work.
From Adoption Push to Adoption Target
The pattern is visible across tax teams. Leadership sends a message about “embracing AI.” A platform is rolled out. Then come the dashboards. Usage is tracked by team, by office, by individual. Partners start asking why certain groups are below the average. The conversation shifts from “should this task use AI” to “why aren’t you using it more.”
Uber is the clearest public example of where this logic leads. In April 2026, CTO Praveen Neppalli Naga told The Information that Uber's full-year 2026 AI budget was already gone. Engineers had been actively encouraged to use Claude Code, Cursor, and other agentic coding tools. Internal leaderboards ranked teams by AI tool usage. By March 2026, around 84% of Uber's developers were classified as agentic coding users. AI-related costs at the company had risen roughly six-fold since 2024.
Uber is a tech company with fast feedback loops. Code compiles or it doesn’t. Tests pass or they don’t. Tax does not work that way. Feedback loops run in months and years, not seconds. Errors surface in audits, not unit tests. Applying the same adoption playbook to tax is not just expensive. It is risky.
What the Metric Actually Measures
The first problem is Goodhart’s Law. When a measure becomes a target, it stops being a good measure. Track AI usage and people will run queries they do not need. They will paste emails into a chatbot to “summarize” them. They will ask the model to rephrase something they already wrote. The dashboard goes up. The work does not get better. The metric is hit. The point is lost.
The second problem is that the metric rewards the wrong behavior. Consider two tax professionals working on the same platform tax liability question. One knows how the rules work across jurisdictions. She knows that EU marketplaces are deemed suppliers only for specific transaction types, that US marketplace facilitator laws apply broadly, and where the exceptions sit. She writes the memo in 30 minutes. The other does not know the rules. He asks three different LLMs. He compares outputs. He re-prompts when they disagree. Two hours and a lot of tokens later, he produces a weaker memo. Partly because he did not know what to ask in the first place. On the dashboard, he looks like the engaged one. She looks disengaged. The person producing the better work, faster, is penalized.
The third problem is that high usage is not good usage. Reaching for an LLM on every task is not a sign of skill. Often it is a sign of the opposite. A usage metric cannot tell the difference between thoughtful use and reflexive use, between the professional who uses AI where it genuinely helps and the one who uses it because the dashboard is watching. It rewards the latter because the latter generates more data points.
Cost sits behind all of this. AI is not free. Every prompt, every agent call, every long context window has a price. “Use it more” is a policy that scales directly with spend. Rewarding usage means rewarding expense. Uber hit the wall in four months. Tax teams have smaller budgets.
Why Firms Do It Anyway
Leaders are not being stupid. Outputs in tax are hard to measure. Quality surfaces months later, sometimes in audits, sometimes never. Activity is easy to count. A usage dashboard gives partners something concrete to show when the board asks how the firm is adapting to AI. The pressure is real, and the dashboard is the easiest answer to it.
The other common argument is that adoption targets are necessary to overcome inertia. Without a push, people will not use the tools. There is some truth to this. But “measurable” and “worth measuring” are not the same thing. And there is a difference between making tools available, training people well, and giving them time to integrate AI into their workflows, versus mandating usage. The first builds capability. The second builds compliance theater. People learn to generate usage without generating value.
Judge the Work, Not the Tools
The fix is not a better AI KPI. It is no AI KPI.
Focus on outputs over inputs. Quality over process. Outcomes over activity. If a tax professional delivers strong work without touching an LLM, that is fine. If another delivers strong work with heavy AI use, that is also fine. What matters is the deliverable, the risk it carries, and the cost of producing it.
The question isn’t how much AI your team is using. It’s whether the work is any good.



