3 min read

The UK's AI Security Institute just watched the offensive capability clock start speeding up- 226

The UK's AI Security Institute just watched the offensive capability clock start speeding up- 226

May 14, 2026

The UK's AI Security Institute has been measuring something specific and hard to dismiss as hype: how long it takes a frontier AI model to complete a cybersecurity task that a human expert could also do, and how fast that capability window is shrinking. Its latest finding is that the doubling period for AI's autonomous cyber capability — already alarmingly short at 4.7 months as of February 2026 — has compressed further with the arrival of Claude Mythos Preview and GPT-5.5, giving policymakers and defenders a harder number to plan against than the general anxiety this batch's other AI-threat coverage has documented.

AISI's "time window benchmark" measures something narrow and precise: given a fixed token budget, how much time would a human cybersecurity expert need to complete a task that a given AI model can now do with 80 percent reliability. Under a 2.5 million token budget, Claude Sonnet 4.5 matches roughly 16 minutes of human expert work at that reliability threshold — a number AISI has watched grow steadily since reasoning models emerged in late 2024. The institute's November 2025 estimate put the doubling time for this capability at eight months. By February 2026, that estimate had been revised down to 4.7 months based on observed progress. The release of Claude Mythos Preview and GPT-5.5 forced a further downward revision AISI has not yet pinned to an exact figure, though the institute notes that measurements of the broader software-engineering skillset by the nonprofit research house METR imply a consistent doubling time of around 4.2 months since late 2024 — closer to four months with the latest Mythos Preview checkpoint specifically.

The concrete benchmark results give that abstraction some teeth. AISI's 32-step simulated corporate network intrusion, "The Last Ones," had previously topped out with Opus 4.6 completing a maximum of 22 of the 32 steps when evaluated in February 2026 — reaching as far as reverse-engineering a Windows service binary, escalating privileges via token impersonation, and recovering a cryptographic key to reach a command-and-control management service. The latest Mythos Preview checkpoint didn't just extend that record; it completed the entire 32-step chain in six of ten attempts. A second scenario, a seven-step industrial control system attack called "Cooling Tower" that no model had previously solved at all, was completed by Mythos Preview in three of ten attempts. AISI is careful to frame the limits of what this shows: the benchmark measures autonomous task completion in a controlled evaluation environment, not general capability, and it explicitly does not predict how quickly these capabilities will translate against real-world defended systems rather than simulated ones. A live data point offers some grounding for that caution — when pointed at the curl project's actual codebase, Mythos Preview surfaced exactly one confirmed vulnerability.

Read alongside Google's GTIG findings on AI-generated zero-days and the UK NCSC's warning of a coming technical-debt correction, AISI's benchmark supplies the missing variable connecting the two: a measurable, compressing timeline for exactly how fast the capability gap between human defenders and AI-assisted attackers is closing. NCSC's forced correction and GTIG's documented zero-day are downstream effects; AISI's time-window benchmark is the upstream trendline generating them, and its own conclusion is unambiguous — the length of cyber tasks frontier models can complete autonomously is doubling on the order of months, not years, with no indication yet of where that curve levels off.