Two routes to impact: automate tasks, expand what people can do
Cheaper models do not guarantee cheaper delivery. A new lens on AI and work: automation, human capabilities, and the evidence needed to tell them apart.
The two clocks measure different things: what models can do and cost, and what happens in work and the economy. Daron Acemoglu's essay in Humanist Review adds a useful question: are we building applications that help people do better work, or mainly trying to reproduce and replace what they already do? His argument is a perspective on the direction of development, not proof that the economic outcome is already determined.
Automating an existing task can reduce the human effort needed for an activity. Complementing a worker can improve quality, increase capacity, or make a new service practical. An application can do both. These mechanisms are relevant across countries, but their effects depend on local institutions, adoption, demand, and how gains are shared. Neither route by itself tells us whether a particular occupation will grow or shrink.
The important test is delivered value, not a benchmark alone. Count integration, review, exceptions, rework, and responsibility alongside model costs. Hold quality fixed when comparing speed or cost; if the benefit is higher quality or new capacity, measure that directly. A task becoming faster is not the same as a complete workflow becoming faster.
The studies linked from the essay must be read on their own terms. A model benchmark, a randomized task experiment, an enterprise survey, and an economy-wide estimate answer different questions. Findings tied to particular workers, software, and study periods are not universal verdicts on AI. Positive and negative evidence should both inform the tracker.
One linked experiment makes the distinction concrete: METR studied 16 experienced open-source developers completing 246 tasks in repositories they knew. Using early-2025 AI tools increased completion time by 19 percent. For balance, a separate 2023 experiment by Noy and Zhang, already in our research library but not linked in this essay, randomized ChatGPT access for 453 professionals doing specified writing tasks: average time fell 40 percent and rated quality rose 18 percent. These studies concern different work and tools. Their percentages cannot be averaged into an economy-wide effect or used to rank today's models.
The linked 2025 enterprise reports from Project NANDA and BCG, and the software-delivery survey, describe adoption and reported business outcomes. They do not establish that AI generally fails or that model improvements have no value. Some links are journalism, vendor advice, opinion essays, or book pages rather than research. We have not imported their headlines, or the essay's broader claims about human and machine cognition, as measured facts.
The article's approximate 5 percent of work automated and 1.5 percent GDP gain over ten years are Acemoglu's US estimates, not global facts or a prediction that 5 percent of jobs disappear. Our existing forecast entry instead quotes his total factor productivity estimate: GDP and TFP are different measures. The essay does not resolve that forecast, and this note makes no new scenario call.
On this site, observed Claude use remains evidence of how people use one product, not an employer adoption rate or a replacement probability. The new homepage section makes the two routes explicit and connects them to the task explorer and team audit. The practical question is what an application improves, under what conditions, and for whom.
Source: Linked sources distinguish commentary from research; results retain their study-specific scope. · Acemoglu's essay in Humanist Review (commentary) · The Simple Macroeconomics of AI (macroeconomic model) · METR: experienced developers with early-2025 AI (randomized experiment) · Combining Human Expertise with AI: radiology (experimental research) · Noy and Zhang: professional writing (2023 randomized experiment, abstract)