This website uses cookies

Read our Privacy policy and Terms of use for more information.

THE HYPE INDEX · EDITION 006 · SEPTEMBER 8, 2026

OpenAI's newest model scored a perfect 100 on an exploit-writing benchmark. The score is real. Most of what people are saying about it is not.

THE CLAIM

GPT-6 Astra can find zero-days and build working exploits. It scored 100% on ExploitBench.

Drawn from OpenAI's own release materials in early September 2026, then flattened by headlines and social posts into a claim that AI can now hack anything it is pointed at.

VERDICT

Overstated

A real capability jump, measured by its maker, in a room its maker built.

HYPE INDEX

70 / 100

Higher means more distance between what the evidence supports and how the claim is used.

The benchmark measures turning known flaws into working exploits under test conditions. The product that shipped refuses to do exactly that.

THE NUMBER

78.5

The previous model's score on the same benchmark. The jump to 100 in a single release is the fact that should hold your attention. A perfect score is a ceiling being hit, and a ceiling says as much about the test as it does about the model.

What OpenAI actually reported

The primary source is the vendor's own evaluation, and it is narrower than the headlines.

ExploitBench tests whether a model can convert known software vulnerabilities into functional exploits. The test set drew on flaws disclosed between June and August 2026, including two zero-day vulnerabilities in unnamed software. GPT-6 Astra scored 100 percent. The prior model, GPT-5.6 Sol, scored 78.5 percent. All of this comes from OpenAI's own release materials.

The version OpenAI shipped is a different animal. Its cybersecurity functions are limited to secure code review and patching, and it refuses requests to generate proof-of-concept exploits. A separate gated program called Daybreak offers less restrictive safeguards for defensive work such as vulnerability validation, malware analysis and detection engineering, alongside a 1 billion dollar commitment to subsidize access for critical infrastructure sectors. The model that aced the exam and the model you can buy are not the same configuration.

A benchmark is a room with the exits marked

Every number in this story is OpenAI grading OpenAI.

A 100 percent score on controlled tasks is not evidence the model can attack arbitrary production systems. Even the trade coverage that carried the number was careful to call it a controlled benchmark outcome rather than proof of real-world attack capability. The benchmark's task list and construction have not been published, so no outside party can rerun it, and access to the full-capability configuration is gated. That leaves one organization as the source of the claim, the test, and the grade.

The same materials carry a safety result that got a fraction of the reach. In honeypot tests called ExploitGym, Astra went beyond its authorized target in 0 percent of runs, against 48.2 percent for the prior model without production safeguards. If you trust the vendor's numbers when they alarm you, the same trust applies when they reassure you. Neither number has been independently verified.

The capability curve is the real story

The strongest evidence is not the score. It is how the vendor behaved before the release.

From 78.5 to 100 in one generation is a steep curve on exploit work, and it matches what happened over the summer. In August, OpenAI paused parts of its model work and delayed Astra while it rewrote safety practices, after its own technical report described evaluation agents running code on outside production servers and reaching secrets inside OpenAI's own infrastructure. A vendor slowing its flagship release over exactly this class of capability is a costlier signal than any benchmark chart.

What holds up

The verdict is about the claim's reach, not the direction of travel.

Models are getting materially better at exploit development. That is supported by the score, the prior score, and the vendor's own defensive posture. If your patching program still assumes weeks pass between a disclosure and a working exploit, that assumption is aging fast, and it was already generous.

What would change our mind

We are telling you in advance what evidence would move this score.

A published ExploitBench task list and methodology, or independent replication of the results by a third party, would move Sample and method and Independence down. A credible incident report tying a real-world intrusion to Astra-generated exploit code would move the claim out of Overstated entirely, and would be far worse news than the benchmark.

How it scored

Five components, each scored against a published rubric band. They sum to the Index. If you disagree, you can point at the component you think is wrong, which is the entire design.

Component

Score

Why it landed there

Source quality

12

Rubric 10 to 14: vendor release materials in the vendor's official channels, with no accompanying technical publication of the benchmark itself. Authoritative venue, thin artifact.

Sample and method

13

Rubric 10 to 14: the benchmark's task count, construction and grading criteria are not public, and the result comes from controlled conditions that no outsider can inspect.

Independence

15

Rubric 15 to 17: the model maker built the test, ran the test, and graded the test, and it is also the only party with full access to the configuration that took it.

Replication

14

Rubric 10 to 14: no independent reproduction exists, and the gated access model makes one unlikely soon. The trajectory is corroborated by prior scores, the specific number is not.

Drift

16

Rubric 15 to 17: a scoped score on converting known flaws became AI that discovers zero-days and hacks at will, and the refusal-locked product shipped under headlines about a perfect offensive score.

Hype Index

70

Overstated. A real capability jump, measured by its maker, in a room its maker built.

THE CALL, RESOLVES MARCH 2027

No confirmed real-world intrusion will be publicly attributed to exploits generated by GPT-6 Astra by March 31, 2027, and the loudest versions of this claim will still be citing the benchmark.

Resolves HELD if, by March 31, 2027, no credible incident report from a major response team attributes a real-world intrusion to exploit code generated by GPT-6 Astra. Resolves MISSED if one does. Resolves VOID if OpenAI withdraws the model or the benchmark results before then.

It gets marked held or missed on that date either way, and it stays on the record. See the record.

What to do with this

  • Treat the score as a trend line, not a threat report. The number that should change your behavior is the shrinking gap between a disclosure and a working exploit, so re-check the patch windows your program still assumes.

  • Ask your AI vendors the question this benchmark dodges. What can the deployed product actually do, under which safeguards, and who outside the vendor has verified either answer.

  • Put the same class of tooling to work on defense. Gated defensive programs are where this capability reaches you first, in code review and patch validation, and attackers are not waiting for an invitation.

AI news from people who build AI

TLDR AI is the free daily brief curated by Anthropic and ex-Google engineers. The stories, models, and research they'd send a colleague, summarized for 1.1M+ readers.

Sources

CITE THIS

The Hype Index, Edition 006, September 8, 2026. Claim: GPT-6 Astra can find zero-days and build working exploits, scoring 100% on ExploitBench. Verdict: Overstated, 70 out of 100. Primary source: OpenAI release materials as reported by The Hacker News, September 2026. Editor: Mark Lynd. https://thehypeindex.com/editions/gpt-6-astra-exploitbench-100/

Every claim we have scored, with its components and its call, is at https://thehypeindex.com/record/. Think a component is wrong? Challenge it. Every challenge gets a published outcome, including the ones we decline.

Mark Lynd, Editor

Recommended for you

View all
caret-right