Wensu Li, Atin Aboutorabi, Harry Lyu, Kaizhi Qian, Martin Fleming, Brian C. Goehring, Neil Thompson
8 min
Abstract
This paper develops a unified framework for evaluating the optimal degree of task automation. Moving beyond binary automate-or-not assessments, we model automation intensity as a continuous choice in which firms minimize costs by selecting an AI accuracy level, from no automation through partial human-AI collaboration to full automation. On the supply side, we estimate an AI production function via scaling-law experiments linking performance to data, compute, and model size. Because AI systems exhibit predictable but diminishing returns to these inputs, the cost of higher accuracy is convex: good performance may be inexpensive, but near-perfect accuracy is disproportionately costly. Full automation is therefore often not cost-minimizing; partial automation, where firms retain human workers for residual tasks, frequently emerges as the equilibrium. On the demand side, we introduce an entropy-based measure of task complexity that maps model accuracy into a labor substitution ratio, quantifying human labor displacement at each accuracy level. We calibrate the framework with O*NET task data, a survey of 3,778 domain experts, and GPT-4o-derived task decompositions, implementing it in computer vision. Task complexity shapes substitution: low-complexity tasks see high substitution, while high-complexity tasks favor limited partial automation. Scale of deployment is a key determinant: AI-as-a-Service and AI agents spread fixed costs across users, sharply expanding economically viable tasks. At the firm level, cost-effective automation captures approximately 11% of computer-vision-exposed labor compensation; under economy-wide deployment, this share rises sharply. Since other AI systems exhibit similar scaling-law economics, our mechanisms extend beyond computer vision, reinforcing that partial automation is often the economically rational long-run outcome, not merely a transitional phase.
Alex: Huh... so it's not just about AI being right more, but shrinking the doctor's decision space.
Sam: Exactly. Humans complement by resolving what's left, and evidence from experts like Agarwal shows joint work shines when allocating cases smartly—high-confidence to AI, tricky to people. The framework ties this to costs: when AI supply curves bend up fast, firms pick the sweet spot where savings match expense.
Alex: That explains why full replacement feels rare... firms balance the curves. But walk me through the supply side—why does pushing AI accuracy higher get so pricey?
Sam: To build an AI that gets things right more often, companies spend upfront on gathering lots of examples to train it—like sorting through thousands of photos to teach it patterns. Bigger, smarter models need more computer power to learn those patterns, and tweaking them takes extra time and energy. Those one-time setup costs climb quickly as accuracy rises, because tiny improvements demand way more effort near perfection. Then there's running the AI: each time it checks an image, it uses compute based on its size. Overall, this creates a supply curve for AI help that's steep—cheap for decent performance, but expensive for flawless.
Alex: Huh, so fixed setup plus per-use fees... and that steepness is key. On the demand side, better AI saves time in a straight-line way up to full replacement—firms pick where the rising AI bill matches those steady savings.
Sam: Yes. The paper calls this the supply-demand optimum for partial automation. Evidence from psych studies, like Hick-Hyman, backs the time savings being linear with reduced uncertainty. This even holds if AI services get shared across companies—fixed costs spread out, but the curve stays convex, favoring teams over full takeovers for vision tasks. It explains why radiologists still thrive with AI triage.
Alex: Okay, so no need for perfect AI because the last bits cost too much relative to human fixes. How did they actually measure that steepness for vision tasks?
Sam: They built a fuller picture: AI accuracy ties to three knobs—amount of training examples, how long you train, and model brainpower—plus task hardness from how many categories to sort, like two choices versus hundreds. They tested this on image classification by fine-tuning a vision model across 80 setups, tweaking those factors. The key measure is cross-entropy loss—it scores how mismatched the AI's probability guesses are from reality; lower loss means sharper predictions.
Alex: Okay, so loss drops as you pump in more data or compute... but what did the tests reveal about the trade-offs?
Sam: The fits were strong—a single equation captured over 96% of the patterns. Data and training time reliably cut loss, but bigger models helped less on tough tasks with many categories, showing diminishing gains. Those inputs act as complements—you need balanced boosts in all three for best results. This non-standard shape fits AI's quirk: more resources shrink errors, bending costs convexly upward.
Alex: Huh, so for real jobs with few categories mostly, the curve stays steep. They surveyed thousands on 461 vision tasks from job data, finding most have under ten categories—like checking if a truck part works or not?
Sam: Yes. Plugging survey accuracy needs and these laws into the cost math confirms partial setups rule: AI hits good-enough cheaply, but perfection explodes expenses versus human tweaks on rares.
Alex: That ties the lab fits to economy-wide tasks neatly. Those lab tests sound solid... but to see how this plays out, look at actual jobs—like security screeners checking bags with X-rays versus zoologists identifying animal species.
Sam: Exactly. Screeners deal with straightforward visual checks, often just a handful of categories like hazard or no hazard, so AI handles most of the work—over 90% substitution—while humans step in rarely for doubts. Zoologists face hundreds of species variations with added uncertainty, making full AI too costly; partial help covers little, humans dominate. This shows two forces: technical feasibility, whether a task is standardized enough for AI, and economic feasibility, if costs make it worth doing at scale.
Alex: Right, so standardized low-variety jobs like screening get hit harder... but high-variety ones stay mostly human. And economy-wide, that's not a huge shake-up?
Sam: The evidence points to under 4% of U.S. labor pay for vision tasks overall—transformative for niches like inspection, gradual elsewhere. Partial setups reshape those jobs via teams, boosting output without mass displacement. The framework's logic fits other AI too, like language models.
Alex: Makes sense... though it's tuned to vision, right? What else might limit it?
Sam: Yes, calibrated just for computer vision; it assumes tasks stay fixed, no new ones created by AI. As a partial equilibrium, it skips broader shifts like wage or compute price changes, or future AI jumps. Still, it offers a solid base for weighing human-AI fits.
Alex: So the takeaway is this unified view: not race-to-replace, but smart optimization favoring collaboration... a meaningful lens on where work heads.
Sam: Precisely. It bridges AI tech metrics to real labor shifts, explaining augmentation over wipeouts. That's the paper's core contribution.
Alex: Well put, Sam. Thanks for breaking it down—this gives a clear, grounded picture of human-AI teamwork in practice. Thanks for listening to ResearchPod.