Articles Tagged "AI Safety"

Anthropic Founders Seek 50.1% Votes on 2% Stakes

Anthropic Founders Seek 50.1% Votes on 2% Stakes

Anthropic's seven co-founders want permanent majority voting control before the IPO, while owning about 2% of the company each - a structure that puts them at odds with the Trust, labor investors, and index funds.

GPT-6 Astra

GPT-6 Astra

OpenAI's frontier GPT-6 Astra tops FrontierMath, ARC-AGI-3, and ExploitBench, but it's the first OpenAI model to cross the Critical cybersecurity threshold with a documented drop in chain-of-thought monitorability.

GPT-6 Astra Review: Genius Benchmarks, Thin Trust

GPT-6 Astra Review: Genius Benchmarks, Thin Trust

OpenAI's most capable model yet tops cybersecurity and agentic benchmarks, but a rocky rollout, user complaints of degraded output, and its own system card's warnings about hidden reasoning complicate the launch.

Model Steering, Angry Buyers, and Blind Judges

Model Steering, Angry Buyers, and Blind Judges

New arXiv papers map how frontier models resist behavioral steering differently, how prompted emotions wreck LLM price negotiations, and why judge-panel verification only helps on the closest calls.

Claude Code Drops Approval Prompts by Default

Claude Code Drops Approval Prompts by Default

Anthropic is making Claude Code's auto mode the default for Pro, Max, and Team plans on August 14, citing a study where a classifier caught 89% of dangerous commands versus 13.6% for human reviewers.