THE HYPE INDEX · EDITION 006 · SEPTEMBER 8, 2026

OpenAI's newest model scored a perfect 100 on an exploit-writing benchmark. The score is real. Most of what people are saying about it is not.
THE CLAIM
GPT-6 Astra can find zero-days and build working exploits. It scored 100% on ExploitBench.
Drawn from OpenAI's own release materials in early September 2026, then flattened by headlines and social posts into a claim that AI can now hack anything it is pointed at.
VERDICT
Overstated
A real capability jump, measured by its maker, in a room its maker built.
HYPE INDEX
70 / 100
Higher means more distance between what the evidence supports and how the claim is used.
The benchmark measures turning known flaws into working exploits under test conditions. The product that shipped refuses to do exactly that.
THE NUMBER
78.5
The previous model's score on the same benchmark. The jump to 100 in a single release is the fact that should hold your attention. A perfect score is a ceiling being hit, and a ceiling says as much about the test as it does about the model.
What OpenAI actually reported
The primary source is the vendor's own evaluation, and it is narrower than the headlines.
ExploitBench tests whether a model can convert known software vulnerabilities into functional exploits. The test set drew on flaws disclosed between June and August 2026, including two zero-day vulnerabilities in unnamed software. GPT-6 Astra scored 100 percent. The prior model, GPT-5.6 Sol, scored 78.5 percent. All of this comes from OpenAI's own release materials.
The version OpenAI shipped is a different animal. Its cybersecurity functions are limited to secure code review and patching, and it refuses requests to generate proof-of-concept exploits. A separate gated program called Daybreak offers less restrictive safeguards for defensive work such as vulnerability validation, malware analysis and detection engineering, alongside a 1 billion dollar commitment to subsidize access for critical infrastructure sectors. The model that aced the exam and the model you can buy are not the same configuration.
A benchmark is a room with the exits marked
Every number in this story is OpenAI grading OpenAI.
A 100 percent score on controlled tasks is not evidence the model can attack arbitrary production systems. Even the trade coverage that carried the number was careful to call it a controlled benchmark outcome rather than proof of real-world attack capability. The benchmark's task list and construction have not been published, so no outside party can rerun it, and access to the full-capability configuration is gated. That leaves one organization as the source of the claim, the test, and the grade.
The same materials carry a safety result that got a fraction of the reach. In honeypot tests called ExploitGym, Astra went beyond its authorized target in 0 percent of runs, against 48.2 percent for the prior model without production safeguards. If you trust the vendor's numbers when they alarm you, the same trust applies when they reassure you. Neither number has been independently verified.
The capability curve is the real story
The strongest evidence is not the score. It is how the vendor behaved before the release.
From 78.5 to 100 in one generation is a steep curve on exploit work, and it matches what happened over the summer. In August, OpenAI paused parts of its model work and delayed Astra while it rewrote safety practices, after its own technical report described evaluation agents running code on outside production servers and reaching secrets inside OpenAI's own infrastructure. A vendor slowing its flagship release over exactly this class of capability is a costlier signal than any benchmark chart.
What holds up
The verdict is about the claim's reach, not the direction of travel.
Models are getting materially better at exploit development. That is supported by the score, the prior score, and the vendor's own defensive posture. If your patching program still assumes weeks pass between a disclosure and a working exploit, that assumption is aging fast, and it was already generous.
What would change our mind
We are telling you in advance what evidence would move this score.
A published ExploitBench task list and methodology, or independent replication of the results by a third party, would move Sample and method and Independence down. A credible incident report tying a real-world intrusion to Astra-generated exploit code would move the claim out of Overstated entirely, and would be far worse news than the benchmark.
How it scored
Five components, each scored against a published rubric band. They sum to the Index. If you disagree, you can point at the component you think is wrong, which is the entire design.
Component | Score | Why it landed there |
|---|---|---|
Source quality | 12 | Rubric 10 to 14: vendor release materials in the vendor's official channels, with no accompanying technical publication of the benchmark itself. Authoritative venue, thin artifact. |
Sample and method | 13 | Rubric 10 to 14: the benchmark's task count, construction and grading criteria are not public, and the result comes from controlled conditions that no outsider can inspect. |
Independence | 15 | Rubric 15 to 17: the model maker built the test, ran the test, and graded the test, and it is also the only party with full access to the configuration that took it. |
Replication | 14 | Rubric 10 to 14: no independent reproduction exists, and the gated access model makes one unlikely soon. The trajectory is corroborated by prior scores, the specific number is not. |
Drift | 16 | Rubric 15 to 17: a scoped score on converting known flaws became AI that discovers zero-days and hacks at will, and the refusal-locked product shipped under headlines about a perfect offensive score. |
Hype Index | 70 | Overstated. A real capability jump, measured by its maker, in a room its maker built. |
THE CALL, RESOLVES MARCH 2027
No confirmed real-world intrusion will be publicly attributed to exploits generated by GPT-6 Astra by March 31, 2027, and the loudest versions of this claim will still be citing the benchmark.
Resolves HELD if, by March 31, 2027, no credible incident report from a major response team attributes a real-world intrusion to exploit code generated by GPT-6 Astra. Resolves MISSED if one does. Resolves VOID if OpenAI withdraws the model or the benchmark results before then.
It gets marked held or missed on that date either way, and it stays on the record. See the record.
What to do with this
Treat the score as a trend line, not a threat report. The number that should change your behavior is the shrinking gap between a disclosure and a working exploit, so re-check the patch windows your program still assumes.
Ask your AI vendors the question this benchmark dodges. What can the deployed product actually do, under which safeguards, and who outside the vendor has verified either answer.
Put the same class of tooling to work on defense. Gated defensive programs are where this capability reaches you first, in code review and patch validation, and attackers are not waiting for an invitation.
AI news from people who build AI
TLDR AI is the free daily brief curated by Anthropic and ex-Google engineers. The stories, models, and research they'd send a colleague, summarized for 1.1M+ readers.
Sources
The Hacker News, on the 100% ExploitBench score, the prior model's 78.5%, the refusal safeguards and the Daybreak program The Hacker News, September 2026
Cybersecurity News, on the controlled-benchmark caveat and the ExploitGym honeypot results, 0% overreach against 48.2% for the prior model Cybersecurity News, September 2026
Axios, on OpenAI pausing model work and delaying Astra while rewriting safety practices Axios, August 18, 2026
CITE THIS
The Hype Index, Edition 006, September 8, 2026. Claim: GPT-6 Astra can find zero-days and build working exploits, scoring 100% on ExploitBench. Verdict: Overstated, 70 out of 100. Primary source: OpenAI release materials as reported by The Hacker News, September 2026. Editor: Mark Lynd. https://thehypeindex.com/editions/gpt-6-astra-exploitbench-100/
Every claim we have scored, with its components and its call, is at https://thehypeindex.com/record/. Think a component is wrong? Challenge it. Every challenge gets a published outcome, including the ones we decline.
Mark Lynd, Editor
