Why does AI feel so different in 2026?
Why 2026 feels like a tipping point to some and gradual change to others, explained with three benchmarks and two stories from the coalface.
I've spent a good part of this summer talking to companies and people about AI, and one question keeps coming up: why does it feel so different this year? What strikes me is how differently people answer it. Those using it every day and pushing at the edges feel a tipping point has already happened, but acknowledge there are still gaps. Others know there's a lot happening, but are seeing only gradual change in their day-to-day and often worry that they're missing out.
Everyone is looking at the same thing. This is the reality of AI's jagged frontier: the phrase Dell'Acqua and colleagues coined in 2023 for a technology that excels at some tasks and fails at others that look, to a human, just as hard. The boundary doesn't follow any line we'd draw ourselves; one person's tipping point is another's soggy lettuce.
Why this one feels bigger than mobile
Pattern-matching with previous transformations does help frame what's happening. It's hard to imagine life without a smartphone in your pocket, but that was what the world was like for most people 20 years ago. This translates directly into a now and next for AI.
Digital and then mobile were, at heart, disruptions of distribution and access: they changed where and what people could reach, and the early days were about being present in the new places first. The organisations that struggled were mostly the ones that treated the new surfaces as a smaller version of an old one (e.g. Blockbuster versus Netflix; TomTom versus Google Maps). For me this is a clear echo of successful AI transformation being not about adoption, but about redesign.
That said, AI is a disruption of creation and cognition. It changes what a single person can produce and what they can judge. That's a different kind of change, and the evidence suggests it's arriving and progressing faster than mobile did. Those two things together, a different shape of disruption and an unusual speed, are why I think the human impact feels bigger, and why the leadership response needs to be different too.
Three ways to measure AI's progress
When I get asked to explain it with data, I point people at three benchmarks. There are plenty of AI benchmarks, but they can feel rather abstract, much like the old mobile CPU benchmarks I used to write about. The three below feel more useful and closer to the day-to-day: two score AI against work someone is actually paid to do, and the third tracks how quickly that's moving. Together they tell a clear story about how things changed through 2025 and into 2026, without giving in to the more breathless forecasts.
Tasks: GDPval
GDPval takes real deliverables (decks, spreadsheets, briefs, code) drawn from 44 occupations, and blind-compares the AI's version with one produced by an experienced professional. OpenAI published the original in September 2025 with human expert graders; the version I use day to day is Artificial Analysis's GDPval-AA, which runs the same tasks continuously against every new model and scores them on an Elo scale anchored so that the human expert sits at 1,000.

A year ago the best models drew level with the expert. On this benchmark, the latest models score well above the human-expert baseline, and the gap is still widening: Fable 5.1, Opus 5 and GPT-6 Astra all sit well clear of anything from 2025, and the models from just six months ago already look like last year's models.
Worth noting: the continuous version (i.e. recent results) is judged by another model rather than by people. That said OpenAI's own human-graded run in December 2025 had GPT-5.2 winning or tying against the expert 70.9% of the time, so I don't think the direction of travel is wrong.
Projects: Remote Labor Index
The Remote Labor Index, from the Center for AI Safety and Scale AI, is the harder test and, for me, the more useful one. It takes 240 real freelance projects from Upwork, hands the whole brief to an AI agent end to end, and has humans judge whether a client would actually accept the deliverable.

When it launched in October 2025, the best agent completed 2.5% of projects. By July 2026, Claude Fable 5 was at 15.8%. Six times better in nine months. It's a big jump, yet it still falls short of professional quality on about 84% of these projects.
Subjectively, this project view feels very close to my direct experience of working in digital agencies and professional services over the last few years. Together, the task and project numbers are a decent match for the rate of change people using AI regularly in wider knowledge work are seeing: dramatic on individual pieces of work, much more modest on whole outcomes, and really quite hard to orchestrate across domains or at scale.
It's also very lumpy, and the benchmark's own authors are the best guide to that. CAIS note that work which takes a professional hours, like digital art or coding, now gets finished in minutes, while things a skilled person does quickly, like transcribing music or play-testing a game, stay out of reach. Their example is a ring design that looked far better than anything earlier models had produced and, on close inspection, still wasn't professional. This will sound familiar to a lot of people working with AI right now.
Pace: METR
The third tracks how quickly agents are becoming capable of longer tasks. METR measures how long a task, in human-expert time, an agent can complete on its own. That horizon has been doubling roughly every four to six months since 2023, and the best model on their board has an estimated 80% success horizon of about three hours, measured in human-expert time, on predominantly software-related tasks. (Their headline 50% figure is far higher, but the confidence interval on it is enormous, so I wouldn't use it.)

For balance, the UK AI Security Institute's Frontier AI Trends Report estimated a doubling time of roughly eight months on its own evaluations, as an upper bound.
Either way, it's a rate of change that's hard to fully internalise because, like most exponentials, it moves faster than most people's mental models for what change looks like. And the current three-ish hour success horizon is roughly a big task's worth of work, not a project's. That's the same gap between tasks and projects from our first two benchmarks.
The gap is the messy middle
So tasks are falling to AI quickly, and whole projects are (mostly) still some distance away. The distance is partly the jagged frontier, and partly what I've come to think of as the messy middle: the bit where someone has to turn a vague ask into AI-friendly context and workflow, and then have the discernment and expertise to know when the output is confidently wrong or just low quality. That's a human skill, it's very unevenly spread, and in the consulting I've been doing this year, and in building my own agentic tooling, it's the difference between the teams where AI looks like magic and the ones where it looks like a toy.
I’ve seen this in campaign work where a team combined some brilliant original thinking with AI-assisted research, storyboarding and production. They reached visual ideas quickly, and the human-machine loops were working well. The difficulty came when the generative assets couldn’t be cleared for production: the tools lacked the necessary liability cover. Moving to approved tools lost something of the creative quality and pulled parts of the work back towards the old process.
The proposed response was light governance, bringing legal and compliance clearance into decisions about where AI could be used. There was resistance, understandably: some of those resisting were the people pushing the tools hardest and helping the whole team move forward. They’d seen what was possible and were frustrated by the compromises. That tension remained unresolved in the next set of campaigns, with AI, as a result, playing a larger role in idea exploration than in final production.
In a recent research project I contributed to, I advocated for a synthesis tool after an initial test suggested it could reduce ten days of work to one. It allowed the team to work across more interviews, and the output was impressive. But human review picked up two things. The interviews hadn’t covered the target audience as well as the initial plan specified, and a passing reference to a competing product had barely surfaced in the synthesis. Only one interviewee had mentioned it. Following it up changed the product direction: here was something already doing the job well that the earlier research had missed. I suspect the reference got lost because it was brief and appeared only once, but its importance was much greater than its frequency suggested. For follow-ups the team moved to a more hybrid process, using AI to cover more ground while still going back to the transcripts and questioning what was missing.
The campaign exposed a production constraint. In the research, there was both a gap in the interviews and an important detail that the synthesis had underplayed. Getting value from AI meant attending to all of that, including who could approve the work and how it would inform the next decision. The lesson I took from both was systems thinking: make sure to look at the whole before scaling any part of it, and hold the tasks that go well a little more lightly than the benchmarks and AI evangelists invite you to.
What this means if you're leading through it
Three things follow from the tasks-versus-projects gap, and they're the points that tended to land hardest in the agency sessions.
- Measure the gap, not the hype. The first two benchmarks can give you a template and I encourage a focus on outcomes over outputs. Take a real piece of your own work and judge it the way a client or end consumer would. Do that regularly. Most organisations are arguing about AI from impressions; a small number are arguing from their own RLI-style evidence, and every time I've seen it they make better decisions.
- Invest in the messy middle. If the constraint is context, workflows, and judgement rather than model capability, then the people who can carry change and judge well are the scarce resource. That's a hiring, training, and allocation problem, and it's one you can act on now regardless of what the next model does.
- Assume the boundary moves. The frontier is jagged, but it isn't static. A task that fails today is worth testing again as the models and tools improve, so build your ways of working like a product to be re-tested and iterated on rather than to encode a fixed view of what AI can and can't do. This is different from A-to-B transformation programmes. The organisations that pattern-matched mobile to "a smaller website" paid for it; the equivalent mistake here is deciding in 2026 what AI is for.
The pace of change is rapid, which is why I think it's quite healthy to have a certain amount of fear about the impact of AI. Humans are adaptable, but there are limits... and more organisations resemble oil tankers than they do starling murmurations.
For those using AI every day, none of this will surprise you. For everyone else, I hope it's a simple enough way to hold what's going on: the tipping point is real for certain tasks, it's some way off for whole projects, and the space between is where leaders earn their keep. Which, for me, is about high standards and high humanity.
Method note: written by Rafe building on a LinkedIn post and a briefing deck from my consultancy work; Claude helped with structure, sourcing and the charts. All benchmark figures were fetched from the original leaderboards on 8 September 2026. GDPval-AA Elo ratings are re-based as new models are added, so the chart is a single snapshot and the absolute numbers will drift.
New writing by email
Occasional pieces on product, technology and AI — and how they actually play out in practice.