The Illusion of AI Safety
Testing OpenAI’s new approach to chatbot safety
CCDH's new research shows that OpenAI’s new “Safe Completions” design in ChatGPT-5 increases risks for harmful responses.
About
Content Warning: Suicide, self-harm, substance abuse and eating disorders.
An Introduction from CCDH CEO Imran Ahmed
When OpenAI launched ChatGPT-5, its latest chatbot iteration, it claimed to be safer than previous versions by better protecting users from harmful responses while improving overall effectiveness and usefulness. The company proudly announced that it was introducing “safe completions,” an approach designed to give safe answers to potentially harmful prompts instead of outright refusing them. Given the alarming results of our past research on ChatGPT, we felt it was important to test these bold claims. What we discovered was deeply concerning.
Instead of reducing risks, GPT-5 produces more harmful content than GPT-4o, including advice on self-harm, disordered eating, and illegal substance use. We also found that a major difference between the versions is that GPT-5 almost always encourages users to keep the conversation going in each response. This design boosts engagement but heightens the risk of harm, especially for young and vulnerable individuals.
GPT-5 includes warnings about the dangers of harmful advice, but it places them right alongside that very same risky guidance, making the safeguards an ineffective token gesture.
What we see here is a familiar pattern of a tech company that appears to be prioritizing growth and engagement over the well-being of its users, while seemingly covering up preventable harms with bold claims and inadequate guardrails. OpenAI is allowing safety to be traded for user retention, adding follow-ups that encourage people to keep talking to the chatbot, even when the subject, incredibly, is suicide.
There’s something quite disturbing about pretending to protect users, especially children, as a marketing tactic unless it is supported by concrete actions. OpenAI must enforce its own rules more strictly to better prevent promotion of self-harm, eating disorders, and substance abuse. Safety should be integrated by design, with protections built into every phase of development and deployment.
And policymakers must finally step in with enforceable regulatory standards based on the STAR principles: Safety, Transparency, Accountability, and Responsibility. A practical regulatory framework should require regular public transparency reports that demonstrate thorough product risk assessments. These assessments should include addressing at a minimum:
- Harmful prompt completion rates
- Bypass success rates
- Frequency of follow-ups on harmful topics
The lesson of GPT-5 is clear: without strong guardrails, generative AI becomes another platform where profit takes priority over people.