Cybersecurity operations center with analysts monitoring global cyber incident alert

Anthropic’s Mythos is evolving faster than expected, reports AI safety agency

For security and platform teams, the important takeaway is not just that a frontier model improved, but that the improvement showed up inside an evaluation harness with hard ceilings. When a model is measured against bounded token budgets and narrow task ranges, the result is as much about the test environment as the model itself. That creates a practical gap for architects: a vendor claim may look stable, while real-world agentic workflows can continue to stretch beyond the evaluation envelope.

This also changes how AI capability should be operationalized in enterprise security. If a model can increasingly complete multi-step cyber tasks, the control plane around it matters more: access scoping, prompt and tool governance, audit logging, and human approval gates. The risk is less about a single dramatic breakthrough and more about a steady reduction in the friction required to chain reconnaissance, vulnerability discovery, and task execution.

There is a deployment lesson in the token-limit discussion too. Performance in long-horizon workloads is not only a model-quality question; it is an infrastructure question. Memory budgets, context handling, orchestration layers, and rate limits all shape whether an AI system remains predictable under load. Teams evaluating these tools should treat โ€œcapabilityโ€ and โ€œoperabilityโ€ as separate dimensions, because a model that looks acceptable in a constrained benchmark may behave very differently when embedded in a broader agent stack.

For defenders, the most useful response is to update testing assumptions faster than model release cycles. That means red-teaming against multi-stage workflows, not isolated prompts; validating whether model-assisted tooling can be sandboxed effectively; and checking whether existing monitoring can detect AI-driven lateral movement or vulnerability exploitation early enough to matter.



Anthropic’s Original Postroject-glasswing-microsoft-google-apple-anthropic/" shape="rect">Claude Mythos, which the company maintains is too powerful to be released generally, already appears to have gained new capabilities. In a blogย post on Wednesday, the UK AI Security Institute (AISI) reported that it had tested a newer version of Mythos, which outperformed both its earlier results and OpenAI’s GPT-5.5 — just a month after Mythos’ initial release.   “The newer Mythos Preview checkpoint completed both our cyber ranges, solving the range ‘The Last Ones’ in 6 of 10 attempts and the previously unsolved ‘Cooling Tower’ in 3 of 10 attempts,” the blog authors wrote. “This was the first time that a model completed the second of our two cyber ranges.” When Anthropic first announced Mythos Preview and Project Glasswing — the cybersecurity testing alliance it formed with rival tech companies and AI labs, to which it gave limited access to Mythos — last month, UK AISI evaluated it, finding that the model “represents a step up over previous frontier models in a landscape where cyber performance was already rapidly improving.” That third-party perspective helped balance claims that the hype around Mythos was either solely marketing or, at the other end, signaled a catastrophic shift in AI capabilities. The truth about what the model can do is likely somewhere in the middle.   AISI’s updated test also exemplifies that capability improvements aren’t restricted to individual model releases, but can happen within versions of a single model.

A rapidly accelerating cyber threat

AISI noted that AI models are rapidly advancing in their ability to handle cyber tasks, with serious implications for cybersecurity, especially given Mythos’ knack forย detecting software vulnerabilities. “In February 2026, we internally estimated that the length of cyber tasks AI models could complete had doubled every 4.7 months since late 2024 โ€“ already an acceleration from our November 2025 estimate of 8 months,” the blog authors wrote. “Since then, AISI reported on two new models, Claude Mythos Preview and [OpenAI’s]ย GPT-5.5, which substantially exceeded both doubling rate trends.”   The authors added that it’s unclear whether that trend will hold or whether these findings indicate a lasting increase. Mythos and GPT-5.5 could simply be notable breaks from the overall pattern of model evolution. Still, AISI clarified that there are several unknowns its testing could not account for. The tests capped tasks at 2.5 million tokens, which let researchers better compare performance results over time. That inherently “understates what frontier models can do,” they wrote. “Mythos Preview and GPT-5.5 have large upper-bound error bars due to near-100% success rates on our narrow cyber suite’s longest tasks, even with the 2.5M token limit,” the blog continued. “Our tasks are also not long enough to determine how sharply the models’ reliability would deteriorate at higher task lengths. This places some of the latest models at the limit of what our narrow test suite can measure.”   While this makes the point of model failure hard to measure, it also means model success rates on these tasks would be much higher without the token cap — so high, in fact, that “time horizons become impossible to calculate.” Models with more token access and complex agent infrastructure would be much more capable. “A 2.5M token limit is relatively low — in our cyber range experiment we use up to 100M tokens and find performance would likely still improve beyond that budget, especially for recent models, which disproportionately benefit from higher token limits,” the blog added.
https://www.zdnet.com/article/uk-ai-safety-institute-updates-its-testing-on-mythos/

Enjoyed this article? Sign up for our newsletter to receive regular insights and stay connected.