OpenAI Launches AI Safety Bug Bounty Program
As artificial intelligence systems grow more powerful, so do the risks of misuse. OpenAI has taken a proactive step by launching a new AI safety bug bounty program focused on identifying and mitigating AI-specific abuse risks. This initiative builds on OpenAI’s existing security bounty program, expanding its scope to address emerging threats in the AI landscape.
What Is OpenAI’s AI Safety Bug Bounty Program?
Announced in March 2026, the program rewards researchers who report design or implementation flaws in OpenAI products that could lead to material harm. Unlike traditional security bounties, this program prioritizes AI-specific risks such as prompt injection attacks, data exfiltration, and harmful actions by agentic AI tools like ChatGPT and Atlas Browser.
Scope of the Program
- Third-party prompt injection: Exploits that manipulate AI outputs through external inputs.
- Data exfiltration: Unauthorized data extraction via AI tools.
- Agentic AI misuse: Harmful actions performed by AI assistants on OpenAI’s platform.
- Proprietary data exposure: Vulnerabilities in account or platform integrity.
How It Works
Researchers submit reports through Bugcrowd, OpenAI’s bounty platform. Submissions are reviewed by dedicated safety and security teams. High-severity issues with clear remediation steps may earn rewards up to $7,500. OpenAI emphasizes that reward decisions are case-specific, prioritizing impact over rigid criteria.
Why This Matters for AI Safety
OpenAI’s program reflects a growing industry trend: proactive AI risk management. By incentivizing researchers to uncover flaws in agentic AI systems—tools that act autonomously on user behalf—the company aims to prevent real-world harm before it occurs.
Key Innovations
The program covers vulnerabilities in connectors and integrations (e.g., MCP integrators) that could enable large-scale abuse. For example, flaws in Codex or Operator tools that allow unauthorized data access or harmful actions are now in scope. This approach addresses risks unique to AI’s operational model, where systems interact dynamically with user inputs and external data.
Comparing AI and Traditional Security Bounties
While traditional bounties focus on code-level vulnerabilities (e.g., SQL injection, XSS), OpenAI’s program expands the definition of “security” to include behavioral risks. For instance, a prompt injection that bypasses content filters to generate harmful outputs qualifies as a reportable issue—even if no code execution is involved.
Industry Context
- Google: Paid $17 million in bug bounties in 2025.
- Microsoft: Expanded its bounty program to third-party code.
- OpenAI: Previously launched Codex security scanners to detect vulnerabilities.
What Researchers Should Know
OpenAI provides clear guidelines for submissions, emphasizing reproducibility and actionable fixes. Researchers are encouraged to test tools like Atlas Browser, Codex, and ChatGPT connectors. However, the program explicitly excludes issues already covered by OpenAI’s standard security bounty.
Getting Started
- Review OpenAI’s bounty scope and rules on Bugcrowd.
- Identify AI-specific risks in agentic tools or connectors.
- Submit detailed reports with reproduction steps and mitigation suggestions.
Conclusion: A New Frontier in AI Risk Mitigation
OpenAI’s AI safety bug bounty program sets a precedent for how the industry can address risks inherent to autonomous systems. By bridging the gap between traditional security and AI-specific threats, the program offers a blueprint for responsible innovation. As AI continues to evolve, such initiatives will be critical to ensuring safety keeps pace with capability.
Ready to explore AI safety research? Visit OpenAI’s Bugcrowd page to learn how you can contribute to securing the next generation of AI tools.








