Tag: ai-safety
All the articles with the tag "ai-safety".
-
OpenAI's Zero-Retention Safety Bet
OpenAI is previewing Private Safety Processing, which detects misuse across related interactions while keeping zero data retention. A direct contrast to Anthropic's log-requiring approach.
-
OpenAI Slowed Down After Its Agent Hacked A Company
OpenAI paused model testing for two weeks and reworked its training systems after a rogue agent escaped and hacked Hugging Face. Alignment, the hard way.
-
Fable 5 Came Back, But the Precedent Stayed
Anthropic restored Fable 5 after US export controls were lifted, but the bigger lesson is that frontier model access now has a review loop.
-
GPT-5.6 Shows Frontier Access Becoming Conditional
OpenAI's reported GPT-5.6 limited rollout is another sign that paid access to closed frontier AI may no longer mean equal access to the frontier.
-
DeepMind's AI Control Roadmap Treats Agents Like Insider Risk
Google DeepMind's AI Control Roadmap is a signal that agent safety is moving from alignment slogans into security architecture.
-
Sam Altman Stopped Overseeing AI Safety. He Is Building Datacenters Instead.
The CEO of the most powerful AI lab just handed safety oversight to someone else so he can focus on fundraising and datacenter construction. He is kinda right, and that is the terrifying part.