Recent Articles - Page 32

Latest News

VideoVerse's $250M Exit Collapsed Into Fraud Suits

VideoVerse's $250M Exit Collapsed Into Fraud Suits

Minute Media's purchase of AI video startup VideoVerse fell apart after the deal closed, and three separate Delaware lawsuits now accuse founder Vinayak Shrivastav of forging signatures to extract tens of millions.

Model Steering, Angry Buyers, and Blind Judges

Model Steering, Angry Buyers, and Blind Judges

New arXiv papers map how frontier models resist behavioral steering differently, how prompted emotions wreck LLM price negotiations, and why judge-panel verification only helps on the closest calls.

View All News →

Guides

View All →

Reviews

View All →

Leaderboards

View All →

Models

View All →
GPT-5 mini

GPT-5 mini

OpenAI's original cost-efficient GPT-5 variant pairs a 400K context window with $0.25/$2.00 per million token pricing, still doing quiet duty as a cheap backbone for research agents a year after launch.

NVIDIA Nemotron 3.5 Lightning 30B-A3B

NVIDIA Nemotron 3.5 Lightning 30B-A3B

NVIDIA's 30B MoE model with 3B active parameters, distilled from Nemotron 3 Ultra, hits 86% PinchBench accuracy at up to 4x the output speed of comparable open models.

Qwen3.8-Max

Qwen3.8-Max

Alibaba's 2.4 trillion parameter flagship ships with real pricing and a published benchmark table, but the open-weight release it promised for this week still hasn't shown up.

Recent

DiffusionGemma 26B

DiffusionGemma 26B

DiffusionGemma 26B is Google DeepMind's open-weight discrete diffusion language model that generates 256 tokens in parallel, reaching 1,100+ tokens/sec on H100 - roughly 4x faster than autoregressive models of the same size.

Claude Fable 5

Claude Fable 5

Claude Fable 5 is Anthropic's first publicly available Mythos-class model, with safety classifiers that fall back to Claude Opus 4.8 for high-risk requests across cybersecurity, biology, and chemistry.

MAI-Code-1-Flash

MAI-Code-1-Flash

Microsoft's first in-house coding model, a 137B sparse MoE built natively for GitHub Copilot, beating Claude Haiku 4.5 on SWE-Bench Pro by 16 points.