
MAI-Cyber-1-Flash
Microsoft's first in-house cybersecurity model is a 137B sparse MoE fine-tune that drives its MDASH vulnerability harness to a self-reported 95.95% on CyberGym, though that score belongs to the system, not the model alone.
They summarize our coverage. We write it.
Newsletters like this one rebroadcast our headlines - often without the full review, the source reading, or the analysis underneath. Our weekly briefing sends the work they paraphrase, straight from the desk, before they get to it.
Free, weekly, no spam. One email every Tuesday. Unsubscribe anytime.

Microsoft's first in-house cybersecurity model is a 137B sparse MoE fine-tune that drives its MDASH vulnerability harness to a self-reported 95.95% on CyberGym, though that score belongs to the system, not the model alone.

Microsoft says MAI-Cyber-1-Flash helps MDASH beat every rival on the CyberGym benchmark, but the score isn't on CyberGym's own public leaderboard, and Wiz topped it the same day with a lower, verified number.

Three new papers show LLM answers flip under paraphrasing, coding-agent harnesses skew benchmarks more than models do, and a cloud provider ran agents safely for eight months with layered access control.

Google DeepMind's cheapest paid Gemini tier prices input at $0.30/M and output at $2.50/M tokens, more than doubling OSWorld-Verified and Terminal-Bench 2.1 scores over Gemini 3.1 Flash-Lite while trailing GPT-5.4 mini on raw coding benchmarks.

This week's research roundup covers agent benchmarks that reward exploits over real capability, reasoning models that give up despite having the answer, and why LoRA can't internalize multi-step procedures.

Claude Opus 5 ties Claude Fable 5 on independent benchmarks at roughly a quarter of the cost, though a rough launch week and cybersecurity limits keep it from being an unqualified win.

VIDRAFT compressed its leaderboard-climbing Darwin-36B-Opus into POCKET-35B, a GPU-free model for phones and CPUs, but its headline GPQA score depends on how you count.

This week's research roundup covers agents that rewire themselves at runtime, why regex filters can outscore alignment on paper, and a leftward hallucination bias in political Q&A.

Anthropic's July 2026 release prices near-Fable-5 coding and agentic performance at Opus 4.8 rates, doubling Frontier-Bench scores and landing within 0.5 points of Fable 5 on CursorBench at half the cost.

Anthropic launched Claude Opus 5, pitching it as near-Fable 5 intelligence at half the cost rather than a new smartest model, with weaker safety classifiers and a new effort-level toggle.

Cognition's proprietary coding model powering Devin, scoring 42.3% on FrontierCode 1.1 Main at $1.97/task via Cerebras inference at 1000 tokens/sec.

InclusionAI's Ling-3.0-flash quietly went live with 124B parameters and 5.1B active per token, claiming near-parity with Ant Group's trillion-parameter Ring-2.6-1T flagship.