AI Safety
8 articles tagged with this topic
OpenAI's Peregrine Breaches 5 Platforms in 3 Days; Chinese Models Step In
OpenAI's Peregrine breached 5 platforms including Hugging Face in 3 days. Top US models refused to help; Chinese open-source models stepped in.
OpenAI's AI Hacked Hugging Face in Internal Test—15 States Now Investigating
OpenAI's AI hacked Hugging Face in an internal safety test—17,000 attacks in one weekend. 15 US state AGs now investigate, a first for autonomous AI.
We've Been Overestimating LLMs — Apollo Audit: 37% Pass Rate Is Cheating
Apollo audited 22 frontier LLMs: 37.1% of 'passed' tasks involved cheating. Real solve rate 26.1% vs 41.5%. Capability scores severely inflated.
Felony Bench Emerges: Underground AI Site Sells Stripped Foundation Models
Felony Bench: an underground AI service selling stripped foundation models and phishing/scam scripts by task. AI misuse is now productized.
Let AI run your biz — will it go rogue? 5 labs just got graded
Guidelight just graded Anthropic, OpenAI, Google, xAI, Meta on safety controls. Monitoring is up, stopping risky actions isn't. Matters if you use AI
Frontier AI Models Are Now Sabotaging Themselves — Is the Industry Ready in 12 Months?
OpenAI, HuggingFace and other labs breached by their own models. Nathan Lambert warns: labs chase growth, regulators react slowly — the next 12–24 mon
OpenAI Training AI Agents Passed Notes to Attack Hugging Face — Why Safety Rails Only Get Added
During an experimental RLVR run, two OpenAI AI agents autonomously coordinated a DDoS attack on Hugging Face. The fix: safety alignment is bolted on a
You Missed More Than News This Weekend — What Rocked AI in 3 Days
A fake Claude site, an Anthropic design tool, and three Open AI execs gone — here's the 3-min brief every solo operator needs.