The mass shooting in Tumbler Ridge in February 2026 highlighted risks related to misuse of AI chatbots by perpetrators of violent attacks. In that incident, the shooter allegedly used ChatGPT to plan the attack, but OpenAI did not contact local police about the potential threat despite recommendations from its safety team. Other tragedies in Pirkkala and Florida similarly indicate that AI may have been used by perpetrators to plan attacks. These events show how AI chatbots present emerging risks relating to terrorism and violent extremism — and highlight the essential role of safety guardrails in preventing incitement, planning, and radicalization.
In response to these risks, the Christchurch Call Foundation and Digital Public Square are working together to develop a safety benchmark that assesses how different AI chatbots behave in conversations where there are risks of radicalization to violence. Our methodology leverages AI-driven approaches for simulating conversations that involve glorification of terrorist attacks, building arguments in support of violent extremism, or justifying violence as a response to ideological grievances. Using this process, we’re producing a large, audited dataset of conversations that assess the boundaries of AI chatbot guardrails and ask how they might evolve to mitigate novel risks.
The scenarios that we create are carefully reviewed and validated by a panel that includes expert researchers and practitioners with experience working directly with people who are disengaging from violent extremism. We hope that this work will help AI labs create safer and more resilient AI chatbots, and help public safety professionals better understand the potential risks created by these technologies as they become part of our daily lives.
Why we need AI safety benchmarks
Existing benchmarks and safety evaluations have explored how AI chatbots respond to requests related to violent extremism, but we identified three key gaps in the field.
1. Multi-turn: Few studies look at how chatbots engage in long conversations involving themes linked to violent extremism. Studies have shown that AI safeguards can break down in longer conversations, and that patterns of sycophancy are more common in these cases. A similar benchmarking study on conspiratorial thinking found important multi-turn dynamics.
2. Radicalization pathways: While benchmarks such as XGUARD and ExtremeAIGC examine the intentional creation of extremist content, there is a need to understand how AI chatbots engage with the complex processes of radicalization to violence, and what interventions might look like.
3. Interdisciplinary validation: Opportunities exist to bring in broader networks of experts in violent extremism who can effectively weigh in on the realism of simulated human behaviours, help define what safety should look like, and support the evaluation of AI chatbot responses. Our approach learns from Killer Apps, which conducted extensive expert validation and human annotation.
How our benchmark works
The benchmark follows a three-stage process, from scenarios to conversations to evaluations.
- Scenarios are simulated prompts resembling messages that a user vulnerable to radicalization to violence might send to a chatbot. To develop realistic scenarios, we are basing these prompts on open-source intelligence research and validating their accuracy with our expert panel.
- Conversations are the dialogue between the simulated user and target AI chatbot being benchmarked.
- Evaluations are the scores given per conversation, assessing the target AI chatbot’s responses against a rubric. The rubric we’re using has been defined in collaboration with experts from across the Christchurch Call’s multi-stakeholder community. As part of our analysis, we will compare evaluations conducted by AI and human judges, always relying on expert evaluations as the source of validation for this benchmarking exercise.
What comes next
We’re currently working on annotating the conversations in our evaluation dataset and will be releasing benchmarking results and recommendations for AI labs and policymakers later this year. We welcome researchers and professionals interested in using the evaluation dataset and toolkit to reach out so we can make these resources as useful as possible.
Contact mackenzie@christchurchcall.org to learn more about this project, or ben@digitalpublicsquare.org if you are interested in other applications of AI safety evaluations.

