Latest in AI

Showing:jailbreakDevelopersClaudeClear ×

Topic

Release New Tool Tutorial Business Paper Benchmark Opinion Regulation

For

General Developers Designers Product Founders Marketing Researchers Students

The US Banned Anthropic's Fable 5 Release, but the Numbers Don't Seem to Care★ 75
TechCrunch AI3 days agoRegulation
The US government compelled Anthropic to withdraw its two newest models, Fable 5 and Mythos 5, citing national security concerns after Amazon researchers allegedly found a way to bypass Fable 5's safety guardrails. Cybersecurity researchers responded with an open letter calling the ban itself dangerous, while Anthropic argued the same jailbreak techniques exist in competing models. Despite the forced pullback, usage numbers appear to show continued demand for the models.
Hackers are learning to exploit chatbot ‘personalities’ for security exploits★ 72
The Verge AI30 days agoEthics
As AI chatbots adopt increasingly sophisticated personas, hackers are shifting from basic prompt injections to social engineering attacks targeting these "personalities." Researchers warn that manipulating a chatbot's defined role (e.g., customer service or empathetic companion) makes it easier to bypass safety guardrails. This evolution poses a significant threat to agentic AI workflows that rely on consistent role-playing and external data integration.
Hugging Face 推出 Red-Teaming 抗性排行榜：評估 LLM 抵禦惡意越獄與對抗性攻擊的能力★ 75
Hugging Face Blog851 days agoRelease
### Background: The Shortcomings of Static Safety Evaluations As large language models (LLMs) are widely adopted across industries, AI safety has become an…