Recent Articles - Page 72

Latest News

VideoVerse's $250M Exit Collapsed Into Fraud Suits

VideoVerse's $250M Exit Collapsed Into Fraud Suits

Minute Media's purchase of AI video startup VideoVerse fell apart after the deal closed, and three separate Delaware lawsuits now accuse founder Vinayak Shrivastav of forging signatures to extract tens of millions.

Model Steering, Angry Buyers, and Blind Judges

Model Steering, Angry Buyers, and Blind Judges

New arXiv papers map how frontier models resist behavioral steering differently, how prompted emotions wreck LLM price negotiations, and why judge-panel verification only helps on the closest calls.

View All News →

Guides

View All →

Reviews

View All →

Leaderboards

View All →

Models

View All →
GPT-5 mini

GPT-5 mini

OpenAI's original cost-efficient GPT-5 variant pairs a 400K context window with $0.25/$2.00 per million token pricing, still doing quiet duty as a cheap backbone for research agents a year after launch.

NVIDIA Nemotron 3.5 Lightning 30B-A3B

NVIDIA Nemotron 3.5 Lightning 30B-A3B

NVIDIA's 30B MoE model with 3B active parameters, distilled from Nemotron 3 Ultra, hits 86% PinchBench accuracy at up to 4x the output speed of comparable open models.

Qwen3.8-Max

Qwen3.8-Max

Alibaba's 2.4 trillion parameter flagship ships with real pricing and a published benchmark table, but the open-weight release it promised for this week still hasn't shown up.

Recent

EXAONE 4.5

EXAONE 4.5

LG AI Research's first open-weight vision-language model packs 33B parameters, 262K context, and STEM scores above GPT-5-mini - but ships under a non-commercial license.

Qwen3.5-Omni

Qwen3.5-Omni

Alibaba's Qwen3.5-Omni takes text, images, audio, and video as input and streams both text and speech output in a single end-to-end model with a 256K context window.

GPT-5.4-Cyber

GPT-5.4-Cyber

OpenAI's GPT-5.4-Cyber is a cyber-permissive fine-tune of GPT-5.4 Thinking with binary reverse engineering, 88.23% on professional CTFs, and access gated through the Trusted Access for Cyber program.

Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS

Google's Gemini 3.1 Flash TTS ships in preview with 30 voices, 70-plus languages, 200-plus inline audio tags, and Elo 1,211 on the Artificial Analysis TTS Arena.

Veo 3.1

Veo 3.1

Google DeepMind's Veo 3.1 generates 4K video with native audio and is now free for every Google account at 10 clips per month via Google Vids.

Qwen3.6-Max-Preview

Qwen3.6-Max-Preview

Alibaba's first closed-weights flagship Qwen ships with a 256K context window, tops six agentic coding benchmarks, and ranks third on the Artificial Analysis Intelligence Index.

GPT-Rosalind

GPT-Rosalind

OpenAI's first domain-specific reasoning model for biology and drug discovery, launched April 16 2026 as a US-only research preview with a 0.751 BixBench score.

Kimi K2.6

Kimi K2.6

Moonshot AI's Kimi K2.6 is a 1T-parameter MoE with 32B active per token, 256K context, a 300-agent swarm running 4,000 coordinated steps, and the top SWE-Bench Pro score among open-weight models at 58.6%.

The Claw Security Ledger - 10 Products in the Dock

The Claw Security Ledger - 10 Products in the Dock

We audited ten AI agent products sold under the Claw name. The ledger shows 11 live CVEs, 130 published advisories, 1,184 malicious marketplace skills, and one leaked SSL private key - concentrated almost entirely in a single vendor.

AI Labs Are Losing Billions - Here's Who Really Pays

AI Labs Are Losing Billions - Here's Who Really Pays

OpenAI burned $2.5B in cash on $4.3B of revenue in the first half of 2025. Anthropic cut its gross margin forecast from 50% to 40%. Here's the compute subsidy math behind every AI subscription, and who's actually paying for it.