Articles Tagged "Inference"

Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite

Google DeepMind's cheapest paid Gemini tier prices input at $0.30/M and output at $2.50/M tokens, more than doubling OSWorld-Verified and Terminal-Bench 2.1 scores over Gemini 3.1 Flash-Lite while trailing GPT-5.4 mini on raw coding benchmarks.

Gemini 3.6 Flash

Gemini 3.6 Flash

Google DeepMind's workhorse Flash model cuts output tokens 17% versus Gemini 3.5 Flash, drops output pricing to $7.50/M, and cuts DeepSWE task tokens by 65% while trailing GPT-5.6 Luna and Grok 4.5 on raw coding scores.

Grok 4.1 Fast

Grok 4.1 Fast

Grok 4.1 Fast is xAI's agent-optimized model with a 2M-token context window, #1 ranking on tau-bench Telecom, and one of the lowest input prices among frontier-adjacent APIs at $0.20/M tokens.