Microsoft's MAI-Cyber-1-Flash scores 96% on CyberGym, 12 points above Anthropic Mythos, at half the prior production cost
Microsoft's MAI-Cyber-1-Flash, paired with its MDASH multi-agent harness, scored 95.95% on CyberGym on July 27, placing it 12 percentage points above Anthropic Mythos 5 and ahead of every frontier model on the benchmark per the Microsoft AI announcement. The company says the combined system delivers those results at 50% of the cost of its prior MDASH configuration.
What
Microsoft announced MAI-Cyber-1-Flash at a small press event in San Francisco on July 27, one day before its Q4 FY2026 earnings report. The model is the company's first purpose-built for cybersecurity work. It operates inside MDASH, Microsoft's multi-agent vulnerability identification and remediation harness, where it handles roughly 90% of tasks while escalating the hardest 10% to GPT-5.4.
The CyberGym benchmark evaluates how AI systems reason over large codebases to find real software vulnerabilities. Running MDASH with MAI-Cyber-1-Flash plus GPT-5.4 for the harder tasks, Microsoft reached 95.95%, against scores of 83.2% to 85.6% for Anthropic Mythos 5, Google Gemini, GPT-5.5 Cyber, and GPT-5.6 Sol per the announcement chart. The 50% cost reduction is measured against the prior MDASH setup (GPT-5.4 plus 5.4 mini plus 5.3 codex).
MAI-Cyber-1-Flash is derived from the MAI-Thinking-1 lineage, built in-house by Microsoft AI. Its training draws on over 100 trillion daily security signals across identity, endpoint, cloud, and network telemetry from 1.6 million customers, per Microsoft.
Alongside the model, Microsoft launched Project Perception, an agentic security platform that runs coordinated red, blue, and green agent teams. Red, blue, and green agent teams cover attacker modeling, vulnerability triage, and code remediation respectively. Project Perception enters public preview August 3 per Help Net Security.
Why it matters
For security teams pricing out AI-assisted vulnerability programs, a 12-point benchmark gap against Anthropic Mythos at half the cost changes the calculus on vendor selection. CyberGym is the primary shared leaderboard for this category, so the score is directly comparable across the three major providers.
The operational case is the more concrete number. Per lead Perception engineer Dave Weston via TechCrunch, the system compressed multi-specialist manual work on vulnerability discovery, prioritization, and code remediation from hours down to minutes. That claim is vendor-attributed and unverified externally, but the public preview opening August 3 gives security teams a near-term window to test it on real codebases.
Context and reactions
Anthropic launched Mythos in April through a limited partner program called Glasswing. OpenAI launched Daybreak, its security offering, in May. MAI-Cyber-1-Flash is the only model from the three that has published a CyberGym score above 90%, though benchmark conditions may differ from real enterprise deployments.
Microsoft AI CEO Mustafa Suleyman called CyberGym "the golden benchmark" at the San Francisco event, per TechCrunch. Hayete Gallot, Microsoft's EVP for Security, described Project Perception as a way for enterprise defenders to "defend against AI with AI at the scale and speed that the attackers have."
What to watch next
Project Perception's August 3 preview is the next concrete checkpoint. Whether the CyberGym numbers hold against real enterprise codebases is the first testable claim. Microsoft has also signaled that MAI-Cyber-1-Flash will expand into additional security workflows beyond software vulnerability work. Any integration with GitHub Advanced Security or Copilot would put the model in the hands of the developers already triaging security alerts inside those tools.
Sources
- Introducing MAI-Cyber-1-Flash inside MDASH: Microsoft AI (primary vendor announcement, July 27, 2026)
- Microsoft launches its first cybersecurity model, plus a new agentic cybersecurity system: TechCrunch
- Microsoft unveils MAI-Cyber-1-Flash, promises cybersecurity AI at half the cost: Help Net Security