Technology & InnovationNeutral
37

Microsoft MDASH Tops CyberGym Benchmark, Outperforming GPT-5.6 Sol

Microsoft’s MDASH, with MAI-Cyber-1-Flash, scored 95.95% on CyberGym, beating Anthropic’s Mythos and OpenAI’s GPT-5.6 Sol. The system costs 50% less and handles 90% of scans itself, routing hardest cases to GPT-5.4, signaling a shift toward AI-driven, continuous vulnerability detection.

DecryptJose Antonio Lanz

Quick Take

1

MDASH scored 95.95% on CyberGym, topping GPT-5.6 Sol and Claude Mythos by over 12 points.

2

MAI-Cyber-1-Flash handles 90% of scans, sending hardest cases to GPT-5.4 to reduce costs 50%.

3

Available in private preview via Microsoft Defender, it generates fixes and offers review workflows.

4

Microsoft says AI-driven discovery will make periodic scans obsolete, pushing continuous security.

Market Impact Analysis

Neutral

The news is about AI and cybersecurity benchmarks, with no direct impact on cryptocurrency markets.

Timeframeshort

Speculation Analysis

Factuality80/100
RumorsVerified
Speculation Trigger10/100
MinimalExtreme FOMO

Key Takeaways

  • MDASH scored 95.95% on CyberGym, topping GPT-5.6 Sol and Claude Mythos by over 12 points.
  • MAI-Cyber-1-Flash handles 90% of scans, routing hardest cases to GPT-5.4 and cutting costs by 50%.
  • Private preview available through Microsoft Defender with fix generation and review workflows.
  • Microsoft expects AI-driven vulnerability discovery to make scheduled scans obsolete.
Benchmark Score95.95%vs GPT-5.6 Sol 83.6% on CyberGym
Cost Reduction50% lessthan previous best MDASH config
Workload Routing90% handledby MAI-Cyber-1-Flash, 10% to GPT-5.4
Benchmark Scale1,507 vulnsacross 188 open-source projects

What Happened

Microsoft dropped MDASH, an AI-powered vulnerability scanner that scored 95.95% on the CyberGym benchmark. Paired with its new MAI-Cyber-1-Flash model, the system crushed rival models from OpenAI and Anthropic. It’s now in private preview through Microsoft Defender, letting teams review findings and auto-generate fixes. The result signals a leap in automated security — flagship models like GPT-5.6 Sol scored just 83.6%.

The Numbers

MDASH’s 95.95% score left competitors in the dust: GPT-5.5 Cyber hit 85.6%, Mythos 5 83.8%, GPT-5.6 Sol 83.6%, and Gemini 3.5 Flash Cyber 83.2%. The CyberGym benchmark tests AI agents on reproducing 1,507 known vulnerabilities across 188 open-source projects. MAI-Cyber-1-Flash handles up to 90% of tasks, while the hardest 10% get routed to GPT-5.4. This routing slashes costs by 50% compared to Microsoft’s previous top configuration.

Why It Happened

General-purpose models are expensive and inefficient for specialized security tasks. Microsoft built MAI-Cyber-1-Flash from scratch to hunt vulnerabilities in complex codebases. By integrating it with MDASH’s multi-agent harness—over 100 agents that audit code, debate findings, and build proofs of concept—the system achieves both high accuracy and cost efficiency. This targeted approach addresses the token-cost problem of running massive models on every scan.

Broader Impact

AI-driven vulnerability discovery is moving from periodic audits to continuous, autonomous scanning. MDASH’s performance suggests a future where scheduled scans and delayed patches become obsolete. For enterprises, this could shrink the window between vulnerability detection and remediation. Competitors will likely respond with specialized models of their own, accelerating the shift to always-on security.

What to Watch Next

  • Public rollout timeline beyond private preview and integration with broader Microsoft security products.
  • How vulnerability disclosure norms change when AI finds and proves exploits automatically.
  • Counter-moves from Anthropic and OpenAI to close the benchmark gap.

Source: Decrypt

This article is for informational purposes only and does not constitute financial advice.

SourceRead the full article on Decrypt
Read full article

Always late to trends?

Join for the latest news, insights & more.

Disclaimer: Bytewit is an independent media outlet that delivers news, research, and data.

© 2026 Bytewit. All Rights Reserved. This article is for informational purposes only.

Read Next

Most Read

Top StoriesBearish
71

SparkKitty Malware Targets Crypto Wallets via App Stores

A new malware campaign called SparkKitty has infiltrated Apple's App Store and Google Play, scanning photo libraries for crypto wallet seed phrases. Check Point reports that trojanized apps like 币coin and SOEX exfiltrated sensitive data, highlighting the danger of storing recovery phrases as screenshots.

80% confidence
Jul 27, 2026, 8:09 PM UTC · Decrypt
Microsoft MDASH Outperforms GPT-5.6 Sol on CyberGym | Bytewit