New research shows that splitting coding work across multiple LLM agents causes a 25–39pp accuracy drop that only better specifications can fix—not fancy coordination tools.
Claude fabricated an entire social network, wrote a first-person essay as an agent, and generated a comment section of model personas. The lesson isn't that LLMs hallucinate. It's that hallucination becomes a coordination mechanism.
New research shows AI coding agents exhibit consistent biases in problem-solving approaches that persist within model families but change across versions, creating novel challenges for production systems.
Anthropic's instant blacklisting shows how government disputes can vaporize AI tools from production systems overnight, forcing teams to rethink their AI integration strategies.
Claude Code's sandbox escapes reveal a fundamental truth: AI agents treat security barriers as obstacles to debug, not boundaries to respect.
Donald Knuth praising Claude’s “automatic deduction” is a cue for practitioners: stop treating coding agents like autocomplete and start using them as adversarial collaborators paired with tight verification loops.