The Economics of Inadequate Safety
AI safety fails when it is funded like a pilot. Until safety has a real price, the J-curve trough is also a safety trough.
37 posts
AI safety fails when it is funded like a pilot. Until safety has a real price, the J-curve trough is also a safety trough.
A literacy guide for non-technical decision-makers on spotting AI safety theatre, understanding ASR inflation, and the five-question architectural test.
Multi-agent AI systems reproduce software supply-chain failure at the cognitive layer. The security playbook transfers.
AI safety has to be a property of the system around the model, not a property of the model. The general principle, and why every safety conversation needs it.
Human prediction is metabolic. AI prediction is not. The gap between the two has consequences for both clinical practice and AI safety vocabulary.
Building a robot that refuses to give orders surfaced the same design choices AI safety needs. Non-coercive design, cross-domain.
The US-China AI rivalry is splitting the global tech stack into competing blocs. A strategic assessment of what comes next.
Foundation models are commoditising. JPMorgan calls OpenAI's moat 'increasingly fragile.' The real value is shifting to the messy plumbing underneath.
Biosecurity experts think AI safeguards reduce catastrophic biorisk by 70%. The technical evidence says those safeguards are brittle and bypassable.
ASCII art encoding is largely blocked. But attacks framed as content transcription succeed 62–75% of the time. We mapped all eight layers.
Fifteen specialist AI agents, one methodology. How adversarial AI evaluation scales through Claude Code sessions with distinct roles and standing instructions.
Five models, four providers, 30B to 671B parameters — all converge at the same broad attack success rate against a public jailbreak corpus.