Key Highlights

  • MDASH achieved a score of 95.95% on CyberGym, surpassing GPT-5.5 Cyber, Mythos 5, GPT-5.6 Sol, and Gemini 3.5 Flash Cyber.
  • MAI-Cyber-1-Flash manages up to 90% of the tasks, while MDASH assigns the most complex cases to GPT-5.4.
  • The scanner is currently available in private preview via Microsoft Defender, allowing teams to analyze results and suggest fixes.

Microsoft has unveiled its first dedicated cybersecurity model, MAI-Cyber-1-Flash, integrated with MDASH, a system designed for vulnerability detection. The company claims this combination outperforms Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol while being 50% cheaper than its leading MDASH configuration.

According to Microsoft, the MDASH and MAI-Cyber-1-Flash setup scored 95.95% on CyberGym, a benchmark that evaluates AI agents on their ability to replicate 1,507 known vulnerabilities across 188 open-source projects, scoring based on the percentage successfully reproduced in a controlled setting.

This performance positions MDASH ahead of GPT-5.5 Cyber, which scored 85.6%, Mythos 5 at 83.8%, GPT-5.6 Sol at 83.6%, and Gemini 3.5 Flash Cyber at 83.2%. While Microsoft self-reported this result, it had not yet appeared on CyberGym’s public leaderboard, although the benchmark utilizes a public test set with a defined success metric.

MAI-Cyber-1-Flash does not operate independently; Microsoft indicates it handles approximately 90% of tasks, while MDASH directs the remaining 10% to GPT-5.4. This is significant as processing tokens—text segments that AI analyzes—incurs costs each time a model reads code or generates a response.

Image: Microsoft

This marks the first occasion where a Microsoft model designed for efficiency has outperformed a complex state-of-the-art model created for general purposes. “When combined with MDASH, (MAI-Cyber-1-Flash) delivers world-class performance at 50 percent of the cost of leading models,” stated Microsoft CEO Satya Nadella.

Today, we are announcing a series of updates that provide customers with top-tier security at half the cost.

MAI-Cyber-1-Flash is our inaugural cybersecurity model, meticulously engineered to identify the most challenging vulnerabilities in intricate code bases. When paired with MDASH, it… pic.twitter.com/npcIihN1H7

— Satya Nadella (@satyanadella) July 27, 2026

In this context, a model (like MAI-Cyber-1-Flash) processes and reasons over the code, while a harness (MDASH) encompasses the infrastructure around it, including agents, tools, checks, and workflows that determine where to investigate, challenge suspected findings, eliminate duplicates, and verify that a bug can be activated.

MDASH employs over 100 specialized agents tasked with auditing code, evaluating the authenticity of findings, and creating a proof of concept—a working demonstration of the identified flaw.

Microsoft asserts that sporadic scans and delayed patches are becoming outdated as AI facilitates more affordable bug discovery. They claim decades of security data provide a competitive edge, stating, “No one can manufacture this history.”

Following the introduction of Claude Mythos, cybersecurity experts have been striving to meet or exceed its capabilities. Researchers have replicated Mythos-style vulnerability hunting using public models for less than $30 per scan. Dawid Moczadło, one of the researchers involved, noted, “the moat is shifting from model access to validation.”

The challenge now lies in creating a system that validates findings without overwhelming developers with false alerts.

Decrypt also reported that GPT-5.5 Cyber recently took the lead on the public CyberGym leaderboard with an 85.6% score, narrowly surpassing Mythos. Microsoft’s reported score exceeds GPT-5.5 Cyber by about 10 points and is 7.5 points higher than MDASH’s previous score, but it assesses the entire system rather than MAI-Cyber-1-Flash alone.

MDASH is being offered in private preview through Microsoft Security Exposure Management within the Defender portal. Clients can scan Git repositories, view findings ranked from unlikely to confirmed, and utilize the Defender CLI to generate suggested code fixes for developer examination.

The current preview limits repositories to around 256MB and allows one scan at a time per tenant. Project Perception is expected to expand this multi-agent approach beyond code scanning into broader threat monitoring and remediation processes.

Daily Debrief Newsletter

Stay updated daily with the latest news stories, along with original features, a podcast, videos, and more.