Articles Tagged "Benchmarks"

Kimi K2.7-Code

Kimi K2.7-Code

Moonshot AI's Kimi K2.7-Code is a 1T-parameter open-weight MoE coding model with mandatory thinking mode, 256K context, and 30% fewer reasoning tokens than K2.6.

MAI-Thinking-1

MAI-Thinking-1

Microsoft's first in-house reasoning model, a 35B-active sparse MoE with 256K context, 97% on AIME 2025, and no distillation from third-party labs.

Best AI Models for RAG - June 2026

Best AI Models for RAG - June 2026

Gemini 2.5 Flash still leads LIT-RAGBench English RAG accuracy at 87.0%, but the full benchmark data reveals two overlooked entries: GPT-4.1-mini at 84.1% and o4-mini at 83.9%.

Claude Fable 5

Claude Fable 5

Claude Fable 5 is Anthropic's first publicly available Mythos-class model, with safety classifiers that fall back to Claude Opus 4.8 for high-risk requests across cybersecurity, biology, and chemistry.