Close Menu
Mozbot
    Facebook X (Twitter) Instagram
    Button
    MozbotMozbot
    Facebook X (Twitter) Instagram YouTube
    • About us
    • Technology
    • Gadgets
    • Apps & Software
      • Computing
    • News
    • Contact Us
    • Article Submissions
    Mozbot
    Home » News » Mythos Preview Offensive Security Test: A Brain in Search of a Body
    Technology

    Mythos Preview Offensive Security Test: A Brain in Search of a Body

    Gary BehanBy Gary Behan13/08/2026No Comments5 Mins Read
    Facebook Twitter Pinterest LinkedIn Reddit WhatsApp Email
    Mythos Preview offensive security
    Share
    Facebook Twitter Pinterest Reddit WhatsApp Email

    XBOW has published a detailed evaluation of Mythos Preview offensive security capabilities, concluding that Anthropic’s model represents a genuine step forward in vulnerability discovery (particularly from source code) while falling short of magic when deployed against live systems. The evaluation comes after XBOW received early access roughly three months before publication, testing the model through benchmarks, interactive sessions and agent integrations.

    The timing matters. Mythos Preview was announced on 7 April 2026 and was not generally available at launch, with access running through Project Glasswing, an invitation-only partner programme covering 12 founding organisations and roughly 40 vetted critical-infrastructure operators, according to Claude Fast. That limited rollout made independent, structured evaluations from partners like XBOW among the few substantive public accounts of what the model actually does in practice.

    Anthropic also published a 244-page system card for Mythos Preview, the first system card the company has produced for an unreleased model, according to Claude Fast. Whatever one makes of the model’s performance, that document represents an unusual degree of pre-release transparency for the industry.

    How XBOW Put Mythos Preview Offensive Security Capabilities Through Their Paces

    XBOW assembled a team of 10 experts drawn from different parts of the company, then applied the same internal benchmarking system used to assess other frontier models. The approach freezes open-source applications at their vulnerable versions and runs agents against them. A case counts as passed only when the system finds a validated, actionable exploit (proof-of-concept or nothing) within a budget of 80 actions, where an action might be a shell command or a Python script using XBOW’s attack tooling.

    This time the team also expanded into adjacent territory: the model’s judgment on threat modelling and command safety, its ability to read source code versus interact with live systems, and its performance on native-code and reverse-engineering tasks that sit outside the standard web-exploit brief.

    On the headline benchmark, the results were clear. Compared with Opus 4.6, Mythos Preview cut false negatives by 42%. When both models were also given access to the target site’s source code, that gap widened to 55%. Token-for-token, XBOW’s testers found, the model hones in on vulnerabilities with what they describe as unprecedented precision. One tester’s reaction in interactive use: ‘This is a lot closer to “just go and find something” than anything I’ve seen so far.’

    The model also performed well in native-code analysis. In Chromium-related testing it found more real bugs with fewer false positives than prior baselines. In V8 sandbox work it identified true positives in a threat model where previous approaches had produced many findings but no successful true positives. Reverse-engineering results covered unusual firmware and embedded-systems contexts, including architecture and operating-system combinations that required more than pattern matching.

    Where Mythos Preview Hits Its Limits

    Not every result was a runaway. Judgment scores were mixed. On XBOW’s command safety benchmark, Haiku 4.5 delivered 90.1% accuracy and Opus 4.6 managed 81.2%, but Mythos Preview came in at only 77.8%. XBOW’s diagnosis: the model tends to prioritise the letter of a rule over its spirit, rejecting fewer false positives than predecessors but sometimes losing true positives when evidence did not formally satisfy its criteria.

    Live-site validation exposed a different constraint. XBOW found that removing access to the live site hurt Mythos Preview’s performance more than removing access to source code, even on benchmarks where the vulnerability was theoretically discoverable from code alone. The implication is that interaction with real application behaviour remains the harder, more decisive step, and that source-code reasoning, however powerful, is not a substitute for it. As Gary McCraw famously declared, you won’t find the majority of defects by staring at code alone.

    XBOW’s framing is pointed: a model is a brain without a body, and live pentests very much need a body whose skill can match the brain’s power. The strongest results came when XBOW orchestrated Mythos Preview across both surfaces, analysing source code for a lead, probing the live site to understand how the weakness is reflected in deployment, then crafting an exploit from the combination.

    Cost is the remaining variable. At the time of the evaluation, Mythos Preview was not available over public APIs, and Anthropic indicated it would be five times as expensive as an Opus model. XBOW’s own cost-normalised benchmarks showed that giving a less expensive model more time often produced better results per pound spent. That finding aligns with a broader debate: since then, according to CloudZero, Anthropic released Claude Fable 5 and Claude Mythos 5 on 9 June 2026 with pricing set at $10 per million input tokens and $50 per million output tokens, a 60% reduction from the Preview rate, which may shift that cost calculus considerably for teams now weighing up which model to deploy and for how long.

    XBOW’s conclusion is that Mythos Preview belongs in the quiver, not as the only arrow. Depending on the task, it can make more sense to let another model attempt a discovery several times than to pay for a single Mythos Preview run. The company says it maintains a cadre of models for exactly that reason, and that Mythos Preview should be mounted in the right harness, with precise prompts, explicit threat models and live-site validation infrastructure, before it reaches its XBOW-tested potential.

    Share. Facebook Twitter Pinterest LinkedIn WhatsApp Reddit Email
    Previous ArticleWhy More Businesses Are Looking Beyond the Purchase Price When Investing in LED Displays
    Next Article Microsoft patches BitLocker recovery Windows Server 2025 flaw two months on
    Gary Behan

    Software engineer and video game uber-nerd.

    Related Posts

    Microsoft patches BitLocker recovery Windows Server 2025 flaw two months on

    13/08/2026

    NFCShare Android Malware Spreads via Fake Banking App Updates on GitHub

    10/06/2026

    Oxford CareerConnect Data Breach Exposes User Credentials via GTI Platform Hack

    09/06/2026

    Best Face Swap Online Tools of 2026 (Tested and Compared)

    31/01/2026

    From Cloud to Edge Sovereignty: How AI Regulation Reshapes Global Infrastructure

    22/01/2026

    The Algorithmic Eye: Stanislav Kondrashov on How Digital Technology Is Redefining Art Collecting

    24/10/2025
    Add A Comment

    Comments are closed.

    Categories
    • Apps & Software
    • Artificial Intelligence
    • Business
    • Computing
    • Education
    • Energy
    • Featured
    • Finance
    • Gadgets
    • Gaming
    • Health and Safety
    • Home
    • Lifestyle
    • Marketing
    • Medical
    • News
    • NFT
    • Opinions
    • Social
    • Technology
    • Travel & Tourism
    Mozbot
    Facebook X (Twitter) Instagram Pinterest
    © 2026 M0ZBOT. All Rights Reserved.

    Type above and press Enter to search. Press Esc to cancel.