
Gamed Benchmarks, Context Anxiety, and LoRA's Limits
This week's research roundup covers agent benchmarks that reward exploits over real capability, reasoning models that give up despite having the answer, and why LoRA can't internalize multi-step procedures.










