
Compact Contexts, Smarter Fine-Tuning, and the Solver Trap
Three papers from today's arXiv: a joint fix for KV cache bloat and attention cost, new evidence that fine-tuning belongs in the middle of a transformer, and why stronger reasoning hurts behavioral simulation.










