Nverg:~$ _

Microsoft launched MAI-Cyber-1-Flash. The 95.95% benchmark belongs to the multi-model system.

3 min read  1.2×

Microsoft announced MAI-Cyber-1-Flash yesterday. Satya Nadella posted a number: 95.95% on CyberGym. The press ran with it. What that number measures is narrower than the coverage suggests.

what the model is

MAI-Cyber-1-Flash is a sparse mixture-of-experts transformer: 137 billion total parameters, 5 billion active, 256,000-token context window. It is a cybersecurity fine-tune of MAI-Code-1-Flash, which itself comes from a mid-training checkpoint of MAI-Thinking-1. Microsoft's first cybersecurity model is, architecturally, a coding model that was trained further on security tasks.

It is not available as a standalone model or API. It ships only inside MDASH, Microsoft's multi-agent vulnerability identification and remediation harness, and only to approved MDASH customers through Azure AI Foundry private preview.

the score belongs to the system

The 95.95% on CyberGym is an MDASH system score. MAI-Cyber-1-Flash handles up to 90% of tasks; GPT-5.4 handles the hardest 10%. The model card doesn't report what MAI-Cyber-1-Flash scores on CyberGym by itself. That number is not in the announcement.

"Microsoft's cybersecurity AI scores 95.95%" and "Microsoft's multi-model system, with GPT-5.4 handling the hard cases, scores 95.95%" are different claims. The coverage is running the first version.

CyberGym Level 1 also measures something specific: given a vulnerability description and the unpatched source code, can the agent reproduce a working proof of concept? It is a known-vulnerability reproduction test. It does not measure blind vulnerability discovery. The public leaderboard did not list the 95.95% result when checked on July 28. On ExploitGym, which asks agents to turn supplied vulnerabilities and crashing inputs into working code-execution exploits, the model scored zero across kernel, userspace, and browser categories.

the cost claim

Microsoft says the new MDASH configuration delivers comparable performance at 50% of the cost of its previous best model mix. The baseline is GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex. Token use, call volume, task mix, and compute allocation behind that comparison are not disclosed. The figure is probably accurate under Microsoft's specific workload. Whether it holds under a different organization's task distribution is a separate question the announcement doesn't answer.

the sentence worth keeping

Buried in the launch post: "As the cost of finding a flaw collapses, the old model of security, where you scan occasionally and patch eventually, is now obsolete."

That's the thing to sit with. If AI agents can scan large codebases for known vulnerability patterns faster and cheaper than before, the assumption baked into most vulnerability management programs — scan on a schedule, triage periodically, patch on a risk-adjusted timeline — starts to break. Offense has the same access to these tools. The dynamic shifts whether Microsoft announces it or not.

The benchmark number is a product claim. The capability description is a security architecture problem. They are not the same thing to read, and the coverage is mostly focused on the first one.

references

  1. Microsoft: Introducing MAI-Cyber-1-Flash inside MDASH
  2. The Hacker News: Microsoft Says New Cybersecurity AI Model Helps MDASH Hit 95.95% at Half the Cost
  3. Help Net Security: Microsoft unveils MAI-Cyber-1-Flash, promises cybersecurity AI at half the cost
← back