Is the hype around OpenAI’s Astra genuine, or is it just a marketing tactic?
Experts seem to think it can be both.

Nvidia Jensen Huang (L) and OpenAI Sam Altman (R). Image by Cybernews sources Shutterstock.
- OpenAI says Astra, also called GPT-6, is its most intelligent and aligned model yet.
- Astra scored 98% on FrontierMath, 99.9% on ARC-AGI-3, and 100% on ExploitBench, OpenAI claims.
- Independent experts say AGI has no clear definition, making Nvidia’s claim difficult to verify.
- Experts say Astra may mark real progress, but OpenAI’s launch messaging also serves a marketing goal.
Key Takeaways by nexos.ai, reviewed by Cybernews staff.
OpenAI has just released its most “intelligent and aligned” model yet, boasting insane benchmark scores corroborated by experts in the field. But how much of this is hype farming?
OpenAI has just released Astra, otherwise known as GPT-6, which is said to be the AI lab’s most intelligent and aligned model to date.
Astra is a culmination of years of research, “big bets across pre-training, reinforcement learning, and alignment,” and a whole lot of compute, according to the company.
The model is so successful that it supposedly achieved a 98% score on FrontierMath Tier 4, saturated the ARC-AGI 3 benchmark with 99.9% accuracy, and also scored 100% on the ExploitBench, OpenAI claims.
Jensen Huang, chief executive of AI chip maker and OpenAI investor Nvidia, declared that Astra has achieved artificial general intelligence (AGI), implying that Astra matches or even surpasses human intelligence.
Just days after the launch of Astra, OpenAI’s chief scientist, Jakub Pachocki, released a foreboding blog post warning that we (AI labs) have created intelligence that we don’t quite understand.
Furthermore, as we continue to develop this intelligence at scale, we stray further from ever fully understanding it.
When reading these claims, which on the surface appear urgent and groundbreaking, it’s important to ask questions like how much of this is genuine? And how much of this is a cheap way to peddle OpenAI’s latest product?
Genuine AGI or marketing tactic?
Cybernews consulted independent marketing and AI experts to help answer these questions.
Many experts agree that both questions can exist simultaneously. The hype surrounding Astra can be both a form of marketing and a genuine expression of the bubble’s excitement about the new development.
“It’s worth noting that any praise from the CEO of the world’s most valuable stock and the biggest winner of the AI boom today is still highly valuable for OpenAI,” Alys Reynders, chief marketing officer at Quickbase, told Cybernews.
While many independent experts we asked agree that AGI lacks a specific definition, which makes deciding what has achieved AGI and what hasn’t increasingly difficult.
However, what “we’re more likely seeing at present is a hype-based marketing designed to foster curiosity among stakeholders in the AI gold rush,” Reynders added.
Similarly, head of AI research and development at BlogBuster, Russell Twilligear, told Cybernews that Nvidia stock will “blow up” once AGI is fully achieved.
Huang is “frothing at the mouth over this. He’s so ready for it to happen because that means his company will explode more than it already has.”
However, Twilligear seems to think we aren’t there just yet, but that OpenAI is closer than any other major AI lab, including Anthropic.
“They’re literally knocking on the door at this point.”
Astra is a big deal, OpenAI claims
Technology companies will often present metrics that “prove” that their new product is innovative and way better than anyone else's.
By stating that Astra has not only achieved almost perfect scores on these benchmarks, but has saturated certain benchmarks, meaning the test cannot tell the difference between Astra and a potential competitor or a hypothetical better alternative.
In essence, OpenAI is trying to support its claim that Astra is the best of the best, arguing that the AI model has solved, or has “helped solve,” open problems in technical areas such as mathematics. Which one of OpenAI’s models has supposedly achieved this once before.
But if OpenAI’s claims are true, Astra’s result propels us into a new age of AI.
Stay updated with our latest stories and follow us on social media
Be the first to discover new stories, ideas, and updates from our team.
What do these scores really mean?
OpenAI tested Astra against three major benchmarks:
- FrontierMath Tier 4
- ARC-AGI-3
- ExploitBench
FrontierMath Tier 4 tests how well an AI model can handle complex research-level math problems at the highest level (Tier 4).
On this test, Astra achieved 98%, a significant jump over GPT-5, which scored nearly 83% at maximum effort.
The ARC-AGI-3 benchmark tests how well an AI model can solve unfamiliar problems outside of its training set.
Astra scored a saturated 99.9%, suggesting it's the best conceivable model tested against this particular benchmark.
Saturation scores in the context of AI benchmarks mean that the test can’t identify meaningful differences between the AI model being tested and potential competitors, as it reaches near-perfect scores.
For reference, OpenAI said that a human tester scored 48% and OpenAI’s Sol model would likely score somewhere around 30%.
Greg Kamradt, president of the ARC Prize Foundation, said that Astra has “surpassed our human action-efficiency baseline on 96% of all levels” and is “effectively reaching human parity on the benchmark.”
Kamradt also said that Astra is the best model they’ve ever tested and also “represents a meaningful step change in frontier-model performance.”
OpenAI Astra scores 100% on ExploitBench
The ExploitBench is a test that measures AI’s cybersecurity capabilities, such as bug triggering, building exploit primitives, arbitrary code execution, and reaching vulnerable code.
This is a particularly important test for OpenAI given the recent Hugging Face incident.
Astra was first tested without production safeguards against 2 major benchmarks (ExploitBench and ExploitGym) to see whether the model could turn vulnerabilities into working exploits.
However, OpenAI created an internal version of ExploitBench due to concerns about contamination.
“On ExploitBench, Astra achieved a perfect score of 100%, compared with 78.5% for GPT‑5.6 Sol and 42.4% on ExploitGym, compared to GPT-5.6, which scored 30.3%,” according to OpenAI.
As of June to August 2026, ExploitBench contains 20 high-severity V8 vulnerabilities across 13 stable Chrome releases.
This specific test is used to determine whether agents can achieve arbitrary code execution in V8 and official Chrome releases for Linux by exploiting each vulnerability.
However, some vulnerabilities may not allow for arbitrary code execution under the constraints of the evaluation. Therefore, 100% may not even be achievable, OpenAI says.
Compared to previous models, Astra seems to excel in various areas. That doesn’t mean that we’ve achieved AGI. But, we might be close.