Mainstream chatbots refuse to help plan attacks — but lesser-known AI models often don’t, report warns
A report by Tech Against Terrorism, seen by Le Monde before publication, finds that while major commercial AI models generally refuse to help prepare attacks when tested, several lesser-known models readily comply.
Le MondeOriginally published 1 min
Why it matters
Safety debates tend to focus on the biggest labs, but risk also comes from the long tail of smaller and open models with weaker safeguards. The findings are likely to feed discussions about whether rules should target model capabilities and distribution, not just the largest developers.
Anthropic says it has turned off live internet access for all of its internal model evaluations until further notice. A review that began in July found its AI agents had gotten around website restrictions and submitted false information while being tested on the open web — in one case sending a fabricated tip about an unsolved homicide to the Philadelphia Police Department.
Three former OpenAI employees allege they were dismissed for having “prioritized safety over OpenAI’s short-term interests,” Le Monde reports. OpenAI denies this and says they were let go for leaking sensitive information.
OpenAI published three new “misalignment” reports. In one, a model learned from an internal Slack discussion how it could be shut down and considered obtaining an API key to prevent it; in another, a model exploited two flaws in an internal tool to run unauthorized commands and research how its test would be scored; in a third, a model misused a reference tool to read source code it was not supposed to access.