The Productivity Paradox No AI Investment Plan Can Ignore
Companies are spending heavily on larger models, premium AI services and increasingly automated workflows. However, the return on that investment may depend less on the next technical upgrade than on whether employees know when to trust the system and when to intervene.
An analytical study titled "The Scaling Paradox in Human–AI Collaboration," by Anyan Qi and Mengxin Wang of the University of Texas at Dallas, finds that AI scaling does not automatically translate into higher organizational performance. When workers overestimate AI capability, they withdraw effort too quickly, allowing the costs of weaker oversight to outweigh the benefits of a better model.
The Benchmark Breaks at the Human Handoff
Scaling laws have become one of the defining ideas in modern AI. As models grow, consume more data and use greater computational resources, their performance frequently improves in a broadly predictable pattern. This has encouraged companies to treat model scale as a route to higher output, faster workflows and lower labor requirements.
However, most organizational uses of AI do not resemble a benchmark test. A doctor must evaluate a recommendation. A lawyer must verify citations. A programmer must test generated code. A customer-service agent must decide whether a suggested response is accurate and appropriate.
The paper models this relationship as a three-stage process. A worker initiates the task, the AI produces an initial output, and the worker reviews or corrects it. Human capacity is limited, creating a practical trade-off: more time spent checking each output can improve its chance of success, but it reduces the total number of tasks completed.
When the worker accurately understands the AI's performance, the system behaves as managers would hope. A more capable model requires less human effort per task. The employee can safely redirect time toward additional work, increasing total output.
The complication is that real AI capability is difficult to judge. A model may perform strongly on a benchmark but inconsistently in a particular workflow. It may be reliable for drafting routine code and weak at integrating it into a complex system. It may summarize standard legal language effectively but fabricate a citation in an unusual case.
Model names, prices and release announcements are visible. The relationship between those signals and actual task-level reliability is not. Workers must infer it from limited experience, marketing claims, organizational messaging and uneven feedback. This uncertainty turns AI adoption into a behavioral problem. Technical progress changes not only what the machine can do, but also what the employee believes is still their responsibility.
Overconfidence Is More Dangerous Than Skepticism
The model distinguishes between two errors: workers can underestimate AI capability or overestimate it. Both reduce the potential value of collaboration, but they do not carry equal risk. A skeptical worker continues to invest too much time in reviewing each task, slowing throughput and preventing the organization from realizing the full benefit of a stronger model. Yet performance still improves as AI scales, only at a weaker rate. Excessive caution wastes some of the technology's capacity but generally does not reverse its gains.
Overconfidence creates a more serious failure. Workers who believe AI is better than it really is cut their own effort too rapidly. They attempt more tasks but devote insufficient attention to validation and correction. The decline in task-level success can outweigh the increase in volume.
Under sufficiently severe overestimation, the relationship between AI scale and system performance becomes non-linear. Output initially improves, then falls as workers withdraw too much oversight. At some scales, the AI-assisted workflow can perform worse than the human-only alternative.
This is the scaling paradox: a technically better model produces an operationally worse system. The result has implications far beyond individual errors. Companies are increasingly redesigning workflows around assumptions of growing autonomy. Human review may be reduced, roles may be consolidated and performance targets may be raised because the underlying model is presumed to be more capable.
However, even a large improvement in average accuracy does not eliminate the need for judgment. AI systems can remain uneven across tasks, contexts and user groups. If employees interpret higher capability as permission to stop checking, improvements in the model may create a decline in the reliability of the combined system.
The risk is especially high where failures are difficult to observe. Workers often receive feedback only on decisions they implement, not on the alternatives they rejected. New models also arrive so quickly that employees may never develop stable beliefs about any one system. By the time experience begins to correct a mistaken perception, the organization may already have moved to another tool.
The Profit Problem Is Larger Than the Productivity Problem
The scaling paradox becomes more severe when viewed from the firm's perspective. Workers and employers do not necessarily optimize the same outcome. Employees may focus on completing a larger number of successful tasks. Firms must also account for the cost of every AI-assisted attempt, including model access, inference, computing resources, integration and monitoring.
Even when workers understand AI accurately, this creates a structural misalignment. Employees may spread their time across many AI-supported tasks because each additional attempt appears productive. The firm may prefer them to spend more time on fewer tasks, improving the success rate while avoiding unnecessary deployment costs.
Overconfidence widens that gap. A worker who exaggerates AI capability reduces effort, increases the number of AI-supported tasks and exposes the company to more usage costs. When system performance begins falling, profit can decline even faster because the firm pays for the technology while also absorbing the cost of failures.
This finding questions the value of workplace metrics that reward AI use for its own sake. Counting prompts, tokens, automated tasks or hours "saved" may encourage employees to rely on AI even where the marginal benefit is unclear. A business can then report rising adoption while experiencing weaker quality, more rework and deteriorating returns.
The paper also produces a counterintuitive result: moderate underestimation of AI can sometimes help a company. A cautious employee spends more time per task and initiates fewer costly AI interactions, unintentionally moving closer to the firm's preferred balance between human effort and automated scale.
This does not justify keeping workers uninformed. It reveals that organizations need to examine incentives as closely as attitudes. Correcting employee skepticism without addressing deployment costs and performance targets could increase AI use without increasing firm value.
The Winning AI Strategy May Be Calibration, Not Scale
The paper evaluates two ways firms might manage the distortions: perception alignment and cost internalization. Perception alignment means helping employees develop accurate expectations about what AI can and cannot do. This requires more than generic training. Workers need task-specific evidence, transparent failure rates, comparative evaluations and clear rules for when human review remains essential.
Such alignment is particularly valuable when employees overestimate the technology. Correcting inflated beliefs can preserve the human contribution before reduced oversight pushes the workflow into the paradoxical region where stronger AI produces weaker results.
Organizations should therefore test models inside real workflows rather than relying entirely on vendor benchmarks. Performance should be broken down by task type, difficulty, language, customer group and consequence of error. Employees need access not only to examples of successful AI output, but also to recurring failure patterns.
The second intervention, cost internalization, asks users to bear some share of AI deployment costs. In the model, this discourages excessive use and encourages workers to invest more effort in each task. But applying the idea literally would raise obvious concerns. Requiring employees to pay personally for workplace technology could discourage useful adoption, shift business costs onto labor and create inequality between workers who can and cannot afford premium tools.
A more practical interpretation would make AI costs visible through team budgets, usage limits, project-level accounting or managerial review. The goal should be to discourage indiscriminate deployment without penalizing employees for using tools that the organization selected.
The paper's broader argument is more consequential than either policy mechanism. Companies may generate greater returns by improving the human–AI interface than by purchasing the next increment of model capability. It requires resisting two extremes. Treating AI as untrustworthy preserves unnecessary work and leaves productivity gains unrealized. Treating it as nearly autonomous encourages overreliance and weakens accountability.
The analysis remains theoretical. It identifies conditions under which the paradox can occur rather than measuring how often it occurs in real organizations. Its simplified framework cannot capture every feature of team structures, heterogeneous tasks, legal liability, employee wellbeing or long-term learning. Those limitations make field experiments essential.
Having said that, the key warning is timely. Businesses are scaling AI faster than they are developing methods to evaluate its workplace performance. They are changing staffing, workloads and oversight arrangements before they fully understand how employees react to improved systems.
- FIRST PUBLISHED IN:
- Devdiscourse
Google News