Security Strategies
Microsoft Says New Cybersecurity AI Model Helps MDASH Hit 95.95% at Half the Cost
Cyber RTJuly 28, 20263 min read

Microsoft has introduced a cybersecurity-specific model, MAI-Cyber-1-Flash, within its MDASH system, achieving a 95.95% score on CyberGym. This model, integrated with GPT-5.4, is cost-effective, reducing expenses by 50% compared to previous configurations. MAI-Cyber-1-Flash handles 90% of tasks, with GPT-5.4 tackling the most challenging 10%. Access is limited to approved MDASH customers, with broader applications planned under Project Perception.
Microsoft has introduced its first cybersecurity-specific model within MDASH, its multi-model system for identifying and addressing vulnerabilities. This new model, MAI-Cyber-1-Flash, combined with GPT-5.4, achieved a score of 95.95% on CyberGym, a known-vulnerability reproduction test. The configuration is touted to be 50% more cost-effective than the existing MDASH setup, which includes GPT-5.4, GPT-5.4 mini, and GPT-5.3 Codex. Access to this new model is currently restricted to approved MDASH customers via a private preview on Azure AI Foundry.
MAI-Cyber-1-Flash is engineered to manage up to 90% of MDASH tasks, with the more complex 10% being handled by GPT-5.4. However, it is not available as a standalone model or for general-purpose API use. The impressive CyberGym score is attributed to the combined operation of MAI-Cyber-1-Flash and GPT-5.4, rather than the new model on its own. CyberGym's Level 1 test involves providing an agent with a vulnerability description and unpatched source code to see if it can produce a proof of concept, but it does not assess blind vulnerability discovery or patch accuracy.
Despite the high score, Microsoft's result was not listed on CyberGym's public leaderboard as of July 28, 2026. The leaderboard still showed an earlier MDASH submission from May 12 at 88.4%. Microsoft's public communications did not clarify whether the new result was submitted for official listing. A previous MDASH score of 96.55% in June included any crash, not just target vulnerabilities, making it difficult to compare directly with the July score.
The MAI-Cyber-1-Flash model is described as a sparse mixture-of-experts transformer with 137 billion total parameters, of which five billion are active, and it has a 256,000-token context window. It is a cybersecurity-focused version of MAI-Code-1-Flash, derived from a mid-training checkpoint of MAI-Thinking-1. The new configuration reportedly replaced 80% of MDASH's existing models, boosting the CyberGym score from 88.4% to 95.95%.
The central technical feature of MAI-Cyber-1-Flash is its task routing capability, handling most tasks while leaving the most challenging ones to GPT-5.4. Microsoft claims this setup offers comparable performance at half the cost of leading models, though specific details about token use, call volume, latency, task mix, or compute allocation were not disclosed, making independent verification difficult.
Microsoft's vice president of agentic security, Taesoo Kim, emphasized the importance of the system surrounding the model, rather than the model alone. Under a lightweight terminal harness, the model achieved various scores on different benchmarks, such as 0.314 on CVEBench and 0.651 on CRSBench. However, it scored zero on ExploitGym categories, which test the ability to turn vulnerabilities into working exploits.
All benchmark tests were conducted in a network-isolated environment without access to production systems or external services. The model card cautions that generated text and code might be inaccurate or incomplete, necessitating review before use in critical applications. This cautious approach underscores the complexity and potential risks involved in deploying AI models for cybersecurity.
The introduction of MAI-Cyber-1-Flash within MDASH marks the first scenario announced for Project Perception, Microsoft's broader initiative for coordinating defensive security agents. Project Perception is set to enter public preview on August 3, with plans to expand the model's application beyond software vulnerability management to other security workflows, signaling Microsoft's commitment to enhancing cybersecurity through advanced AI solutions.


