Articles Tagged "Benchmarks"

EXAONE 4.5: LG's Open VLM Beats GPT-5-mini on STEM

EXAONE 4.5: LG's Open VLM Beats GPT-5-mini on STEM

LG AI Research released EXAONE 4.5, a 33B open-weight vision-language model that posts higher STEM scores than GPT-5-mini and Claude 4.5 Sonnet - but a non-commercial license caps its real-world reach.

Muse Spark

Muse Spark

Meta's first closed-source frontier model scores 52 on the Artificial Analysis Intelligence Index, leads on HealthBench Hard, and ships free at meta.ai - but has no public API yet.

Claude Mythos Preview Finds Thousands of Zero-Days

Claude Mythos Preview Finds Thousands of Zero-Days

Anthropic's restricted Claude Mythos Preview model autonomously discovered thousands of high-severity vulnerabilities across every major OS and browser, including bugs hiding in plain sight for 27 years.