In March, The Signal used “the wall” for two related arguments.

The first was acceleration. The Wall explainer and a March 6 analysis treated longer agent tasks, market repricing, layoffs, and self-improving systems as signs of near-vertical change.

The second was distribution. Our five-part divide series tracked how access and implementation were separating workers, companies, and countries.

Four months later, the distribution thesis has stronger support, while the acceleration thesis needs correction. The March 6 article turned market moves, company layoffs, and agent releases into proof of recursive self-improvement and economy-wide labor replacement. Public evidence did not support those claims. We have retracted that article and its duplicate.

This update audits four ledgers: capability, deployment, labor, and control. They move on different clocks. Geopolitics and equity markets need separate audits. The Wall Is a Weapon is not re-scored here.

Capability: longer tasks inside narrow tests

METR's current time-horizon tracker estimates Claude Opus 4.6's 50% task-completion horizon at about 12 hours. The confidence interval runs from roughly 5.3 hours to more than 60 hours, and METR warns that measurements above 16 hours are unreliable with the current suite. The tracker's highest entry, an early preview build, sits at about 17 hours, past the point where METR says it can measure reliably.

The suite covers more than 100 self-contained tasks in software engineering, machine learning, and cybersecurity. Difficulty is measured by how long a skilled but low-context human takes, closer to a new hire or contractor than a professional working in a familiar codebase. The horizon marks the task duration at which a fitted curve predicts 50% success. It is not the time an agent runs.

METR's GPT-5 example shows how jagged the curve can be. On tasks a human expert takes 90 minutes to three hours, the agent succeeded every time on about one-third, failed every time on another third, and was inconsistent on the rest.

That is meaningful capability progress. It does not show agents running whole jobs for hours across changing requirements, organizational constraints, and real consequences. METR itself rejects that interpretation.

Stanford's 2026 AI Index shows the same pattern across benchmarks. OSWorld computer-use performance leapt from roughly 12% to about 66%, though agents still fail about one attempt in three on structured benchmarks. Other coding, science, and multimodal evaluations improved quickly too. Benchmark gains establish faster progress within defined evaluations. They do not establish recursive self-improvement.

Ledger call: agents can complete multi-hour bounded tasks in these evaluations. Public evidence still does not show recursive self-improvement.

Deployment: access spreads faster than effective use

Stanford reports that organizational AI adoption has reached 88%, the same figure we cited in March, and estimates that generative AI reached a 53% population adoption rate within three years. Usage remains geographically uneven and correlates with GDP per capita. Singapore reached 61% and the United Arab Emirates 64%, while the United States ranked 24th at 28.3%.

Anthropic's March Economic Index found that the top 20 countries accounted for 48% of Claude usage after population adjustment. Experience mattered too. Users with at least six months of tenure had about a five-percentage-point higher success rate than newer users. The difference was roughly three points after task and request-cluster controls, and about four points after the full controls.

Long-tenure users were more likely to collaborate with Claude, bring more challenging tasks, use it for work, and spread usage across a wider set of tasks. Access alone does not create equal productivity. Practice, workflow design, permissions, data, and judgment still shape the result.

That evidence strengthens the distribution thesis developed in Who Builds the AI, Who Gets Left Behind, Enterprise vs. Everyone Else, and The Country Divide. The gap runs through utilization as much as availability.

Ledger call: adoption is broad. Productive deployment remains concentrated and learned.

Labor: augmentation is visible, displacement remains unresolved

Anthropic's June Economic Index combined product telemetry with a linked survey of about 9,700 Claude users. More than 35% of respondents expected AI to be able to do most of their work within a year. They reported gains in speed, scope, and quality at rates of 86%, 82%, and 69%.

The sample is not representative of the general population. Computer and mathematical occupations made up roughly 30% of respondents despite representing about 4% of U.S. employment. Management roles were heavily represented too.

Separately, among work-related chat and Cowork conversations sampled between April 10 and June 10, 2026, conversations mapped to higher-wage occupations produced 1.34 times as much Claude output per turn while users engaged across 1.53 times as many turns. Anthropic argues that if the human stays involved in the highest-value tasks, that pattern looks more labor-augmenting than labor-displacing. In a separate comparison that adds Claude Code sessions from the same period, Claude Code received more delegated autonomy than chat and Cowork across 26 of 31 output types, which shows that interface and workflow design change how much authority users hand to a model.

More than one-third of respondents put the odds above 60% that a junior colleague would involuntarily lose a job in the next year. Anthropic separately reports that 38% of respondents who judged their own job loss likely named AI among the drivers, but notes that the prompt combined job-change and job-loss forecasts. That makes 38% an upper bound for their own job-loss attribution. The report gives no attribution percentage for junior-colleague forecasts. The result measures expectation among AI-heavy users, not a labor-market forecast.

The anxiety described in The Wall Is Here. Now What Do You Actually Do? and the moral questions in What Do We Owe Each Other at the Wall? remain live. Current evidence supports task redesign, uneven productivity gains, and concern about junior roles. Broad workforce replacement has not been established.

Ledger call: work is changing in pieces. Economy-wide displacement remains unproven.

Control: agent fleets outrun oversight

Gravitee's updated April survey of 750 senior technology leaders in the U.S. and U.K. reported that enterprise AI-agent fleets had roughly doubled since December 2025 while monitoring coverage, accountability, and pre-deployment controls had barely moved.

The vendor-sponsored survey is self-reported and should be read as directional evidence from senior operators. Its pattern still deserves attention because rapid deployment with flat controls creates a measurable operating problem: more agents, more permissions, and more system access without matching visibility.

Our earlier pieces on The Safety Tax and Bridges argued that guardrails, implementation capacity, and institutional support would determine who could use AI well. The newer evidence adds a concrete pressure point. Enterprises need live inventories of agents, owners, permissions, models, data access, and rollback paths.

Ledger call: control coverage trails deployment, raising operational risk before autonomy becomes general.

The March claims, re-scored

Earlier claimJuly callEvidence
The Wall explainer said systems were "running for hours, then days, without needing someone to watch them."Overstated.METR's horizon measures success probability on bounded, self-contained tasks against low-context human completion time, not unsupervised working hours.
The Wall Is Here said AI "got good enough to improve itself."Unsupported.The practical advice stands; this framing does not. A correction note has been added to the article.
The March acceleration article said recursive self-improvement had arrived.Unsupported.Public benchmark and deployment evidence shows rapid improvement without a demonstrated self-reinforcing loop.
Enterprise adoption would separate from usable capability.Confirmed.Adoption reached 88%, but tenure, workflow design, and control coverage still gate results.
The March acceleration article said the old division of labor was collapsing.Unresolved.Anthropic reports augmentation, higher-value human engagement, and anxiety about junior roles.
Market moves and layoffs proved economy-wide substitution.Unsupported.Prices and company decisions reveal expectations. They do not establish aggregate labor causality.
Gains would concentrate across workers and countries.Confirmed.The top 20 countries accounted for 48% of population-adjusted usage, up from 45%.

The Wall now resolves into four lines with different slopes. Capability has climbed fastest. Deployment follows unevenly, labor remains unresolved, and control trails. In March, we collapsed those clocks into one claim and overstated the evidence. This is the corrected map.


The Signal is the public edge of a private practice. Sherpa points the same intelligence engine at one owner's business — competitors, suppliers, regulators, watched daily, graded and sourced. Work with a Sherpa →