Microsoft's cyber model finds 1,507 bugs and skips the $30 club 🐛
Back to feed

Microsoft's cyber model finds 1,507 bugs and skips the $30 club 🐛

Microsoft has introduced a dedicated cybersecurity model, MAI-Cyber-1-Flash, and paired it with its MDASH vulnerability-hunting system in a configuration the company says surpasses Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol while costing 50% less than its previous top MDASH setup. The combined system scored 95.95% on CyberGym, a benchmark that asks AI agents to reproduce 1,507 known vulnerabilities across 188 open-source projects and measures the percentage successfully reproduced in a controlled setting. Microsoft CEO Satya Nadella wrote, "When combined with MDASH, (MAI-Cyber-1-Flash) delivers world-class performance at 50 percent of the cost of leading models." The figure is self-reported by Microsoft and had not appeared on CyberGym's public leaderboard at publication time, though the benchmark uses a public test set and a defined success metric.

According to Microsoft, the MAI-Cyber-1-Flash and MDASH combination routed most tasks to the new model, sending only the hardest 10% of cases to GPT-5.4, with tokens representing the per-call cost of an AI reading code or producing an answer. The configuration placed MDASH ahead of GPT-5.5 Cyber at 85.6%, Mythos 5 at 83.8%, GPT-5.6 Sol at 83.6%, and Gemini 3.5 Flash Cyber at 83.2%. Microsoft described MAI-Cyber-1-Flash as "our first cybersecurity model, built ground up to find the most challenging vulnerabilities in complex code bases" in a post on X dated July 27, 2026.

MDASH itself is the surrounding machinery, using more than 100 specialized agents assigned to audit code, debate the validity of findings, eliminate duplicates, and produce a proof of concept demonstrating that a flaw is exploitable. Microsoft said periodic scans and delayed patches are becoming outdated as AI-driven bug discovery lowers costs, arguing that "No one can manufacture this history" of decades of internal security telemetry.

The announcement arrived alongside independent research showing Claude Mythos–style vulnerability hunting had been reproduced with public models for under $30 per scan, with researcher Dawid Moczadło noting that "the moat is moving from model access to validation." Decrypt also reported that GPT-5.5 Cyber had entered the competitive field as companies race to automate the discovery and verification of software flaws at scale.

Share:
Publishercryptonewsroom.xyz
Published
CategorySecurity

Disclaimer: This content is for information and entertainment purposes only. It does not constitute financial, investment, legal, or tax advice. Always do your own research and consult with qualified professionals before making any financial decisions.

See our Terms of Service, Privacy Policy, and Editorial Policy.