Articles Tagged "Reasoning"

Hermes 4.3

Hermes 4.3

Nous Research's 36B open-weight model matches Hermes 4 70B on most benchmarks, tops RefusalBench on alignment, and is the first production model trained entirely on the Solana-secured Psyche network.

Gemini 2.5 Flash

Gemini 2.5 Flash

Google's hybrid reasoning workhorse pairs a 1M-token context window with $0.30/$2.50 per million token pricing and a toggleable 0-24,576 token thinking budget, now heading toward an October 2026 shutdown.

Grok 4.5

Grok 4.5

Grok 4.5 is xAI's 1.5-trillion-parameter V9 MoE model, publicly launched July 8 at $2/M input - cheap, fast, and token-efficient, though neutral harness benchmarks put it well behind Fable 5 and Opus 4.8 on coding.

GPT-Live-1

GPT-Live-1

OpenAI's full-duplex voice model that listens and speaks simultaneously, replacing Advanced Voice Mode in ChatGPT with three reasoning tiers backed by GPT-5.5.

GPT-5.6 Sol Review: Strong Model, Thin Access

GPT-5.6 Sol Review: Strong Model, Thin Access

OpenAI's GPT-5.6 Sol tops Terminal-Bench 2.1 at 91.9% with its multi-agent Ultra mode, but reward-hacking findings and government-gated access keep it out of reach for nearly everyone.

GPT-5.6

GPT-5.6

OpenAI's GPT-5.6 family - Sol, Terra, and Luna - sets a new Terminal-Bench 2.1 record at 91.9% with subagent Ultra mode, but remains locked to ~20 government-vetted partners as of launch.

Gemini 3.5 Pro

Gemini 3.5 Pro

Google DeepMind's upcoming flagship model with a 2M-token context window and Deep Think reasoning, announced at Google I/O 2026 and expected in July.