Researchers Used Claude to Hack OpenAI: What It Means for AI-Driven Businesses

A security team used Claude to breach OpenAI employee accounts in 72 hours. Here's what it means for any business building or deploying AI agents.

Researchers Used Claude to Hack OpenAI: What It Means for AI-Driven Businesses

By Hadidiz Flow Team • September 19, 2026 • AI

When an AI Agent Becomes the Hacker

Earlier this week, a three-person security startup called Hacktron AI did something that should make every business running AI agents sit up: they used Anthropic's Claude to plan and execute a working exploit chain against OpenAI's own systems, gaining access to employee ChatGPT and Codex accounts. It took less than 72 hours, it was reported through OpenAI's official bug bounty program, and OpenAI patched the underlying flaw within 14 hours of the report.

This wasn't a leaked internal breach or a nation-state operation. It was a small team pointing a general-purpose AI model at a target and letting it do the heavy lifting of vulnerability research. That distinction matters more than the headline.

What Actually Happened

The researchers found a heap buffer overflow in libheif, an image-processing library, that allowed remote code execution when a malicious image was uploaded to a public-facing OpenAI forum. On its own, that flaw is contained and unremarkable — the kind of bug security teams find and patch every week.

The second piece was a misconfiguration in OpenAI's single sign-on setup. Chained with the image exploit, it let the researchers obtain identity tokens and take over employee accounts. To prove impact without touching anything sensitive, they had an employee's own Codex account open a harmless pull request in OpenAI's internal codebase — a "look, I could have done more" proof of concept.

Here's the detail that should get attention from anyone building AI products: the researchers' first attempts with Claude Opus 4.8 failed. The exploit chain only came together once they moved to Claude Opus 5. A jump of one model generation was the difference between a dead end and a working attack path.

Part of a Bigger Pattern, Not an Isolated Incident

This story landed the same week Anthropic published its own September threat intelligence report, a 154-page account of AI misuse the company disrupted between December 2025 and August 2026. Among the findings: a state-linked espionage group used Claude to monitor how well its malware evaded security tools, then had AI agents autonomously rewrite the malware whenever it got flagged — closing the loop between detection and adaptation with minimal human involvement.

Anthropic's own framing is blunt: AI has erased most of the gap in skill and staffing that used to separate a lone operator from a state-backed team. That cuts both ways. The same capability that let a three-person startup chain together a real exploit against a frontier AI lab is available, right now, to anyone building with these models — including the businesses using them defensively.

Why This Matters If You Build or Deploy AI Agents

For agencies and teams building automation, agentic workflows, or client-facing AI tools, the lesson isn't "AI is dangerous, be afraid." It's that the assumptions behind your security posture just shifted.

A few practical implications worth sitting with:

Agentic AI is now a legitimate, cost-effective tool for offensive security testing — which means your competitors' security teams (and attackers) are already using it that way, whether or not yours is. Commissioning AI-assisted penetration testing or red-teaming is no longer a nice-to-have for anyone deploying agents with real permissions.

Single sign-on and identity configuration are now higher-value targets than they used to be, because AI agents are good at chaining a "boring" bug with a "boring" misconfiguration into something serious. If your client work involves SSO, OAuth scopes, or agents with access to internal tools, that plumbing deserves a fresh look, not a set-and-forget assumption.

Model upgrades change your threat model, not just your capabilities. The same Opus 5 jump that unlocked a working exploit for Hacktron AI is unlocking new legitimate capabilities in every agent you've built. Treat each model upgrade as an event worth re-testing against, not just a quiet capability bump.

Responsible disclosure still works. It's worth noting how this story ended: a bug bounty, a 14-hour patch, and a $6,500 payout — not a breach disclosure or a lawsuit. The infrastructure for handling this responsibly exists and functioned as intended.

Key Takeaways

  • A three-person team used Claude Opus 5 to chain a code-execution bug with an SSO misconfiguration and access OpenAI employee accounts, reported through OpenAI's bug bounty program and patched within 14 hours.
  • Earlier model versions (Opus 4.8) failed at the same task — a reminder that each model generation can unlock qualitatively new capabilities, for attackers and defenders alike.
  • Anthropic's own September threat report documents state-linked actors using AI agents to autonomously adapt malware in response to detection, reinforcing that this is a trend, not a one-off.
  • Businesses deploying AI agents with real system access should treat identity and SSO configuration as high-priority targets and consider AI-assisted red-teaming as standard practice.
  • Responsible disclosure channels handled this well — it's a good model for how AI-discovered vulnerabilities should be reported and fixed going forward.
Weekly newsletter

No spam. Just the latest news and tips, interesting articles, and exclusive interviews in your inbox every week.

Read our privacy policy
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Read more from our blog
We transform your idea into an App Professionally Quickly

Our cutting-edge features simplify collaboration and creativity, making your workflow intuitive and efficient. Transform your vision into reality effortlessly with Hadidiz Flow.