Book a demo →
← All insights AI News

OpenAI's Astra Uses a Browser. That Splits Brand Visibility Into Two Questions

OpenAI's Astra Uses a Browser. That Splits Brand Visibility Into Two Questions

OpenAI announced GPT-6 Astra on 3 September 2026, first to companies in its application-based cybersecurity programme, with ChatGPT Plus, Pro, Business and Enterprise, the API and AWS following “in the coming days” (CNBC, 3 September 2026).

Most of the coverage led with a quote. At the end of the press conference, co-founder and president Greg Brockman said “Welcome to the AGI era” (VentureBeat, 3 September 2026). That is a real quote, not a headline writer’s compression.

The part that changes anything for brand visibility is duller and sits further down: OpenAI positions the model around operating a computer the way a person does. Browser, spreadsheets, desktop applications, finished documents, multi-step work. Mia Glaese described users delegating “more complex work”; Brockman described a shift from prompting to supervising.

The number that would settle it was not published

OpenAI has its own benchmark for economically relevant work, GDPval. It was not part of the announcement. VentureBeat names the omission explicitly, and it is the omission that matters: GDPval is the measurement built to support exactly the claim the press conference made.

What was published are these, from the press conference table:

BenchmarkAstraGPT-5.6 Sol
ARC-AGI-399.9
FrontierMath Tier 4 v297.6
ExploitBench100, without production safeguards
DeepSWE v1.174.165.7
OSWorld 2.072.6 in ~40 min65.7 in ~75 min

Three cautions belong with that table.

The ExploitBench score was measured with the product’s protections switched off. OpenAI is explicit about it: “We first tested the model without production safeguards on ExploitBench and ExploitGym.” The version people actually get refuses that work — “Astra will refuse to comply with more advanced cybersecurity tasks such as creating proof-of-concept exploits for vulnerabilities.” Secure code review and patching are allowed. A perfect score on an unguarded model is not a capability of the shipped one.

ARC-AGI-3 is not a model-against-model comparison; other systems run there in different configurations, including agent architectures around a model rather than a model alone. And the cost figure that travelled furthest, roughly 57 percent lower per task on DeepSWE, is OpenAI’s own estimate, not a third-party measurement.

One number runs the other way, and it belongs here for the same reason the others do. In a misuse test without safeguards, GPT-5.6 Sol went beyond the authorized target in 48 percent of cases. Astra did it in none.

OpenAI’s chief scientist Jakub Pachocki added a caution of his own: progress in intelligence does not guarantee progress in alignment.

Why the browser is the part to watch

Set the AGI question aside. It is undefined in the announcement anyway, and it is not clear whether it refers to the model, the model with tools, or the whole system around it.

Here is what is concrete. A model that operates a browser can look something up at the moment it is asked. A model that does not, answers from what it absorbed during training. Those are two different sources, and they fail in two different ways.

For anyone measuring whether AI assistants mention their brand, that distinction has been quietly blurring for a while, and Astra’s positioning makes it harder to ignore. Ask an assistant “which tools do X” and you may get an answer drawn from training data that is months old, or an answer assembled from pages fetched seconds ago. Both look identical in the output. Neither announces its source.

What this means for measurement

If you run any kind of brand visibility check, it is worth separating two questions that most measurement treats as one:

Does the model know me? This is about what is in the training data, and about what the wider web said about you when that data was collected. It moves slowly. It is close to unfixable in the short term.

Does the model find me when it looks? This is about what a fetch returns right now: whether your pages are reachable, whether the answer is on the page rather than behind an interaction, whether the relevant sentence survives being read out of context.

The second question is the one that a browsing model actually asks, and it is the one you can still influence this quarter. It also means a visibility measurement taken once a month is measuring a different thing than it used to. A training-derived answer is stable between runs. A fetched answer is not.

We do not yet know how Astra behaves on brand and category questions specifically, and anyone telling you otherwise this week is guessing. The capability is documented. The consequence for the answers your buyers see is an open question, and it will take weeks of measurement rather than a press conference to settle.

One thing OpenAI did not say

It is worth being precise, because the coverage has not been. OpenAI did not describe replacing people. The language throughout is delegation and supervision. Writing it as replacement puts words in their mouth, and there are enough real things in the announcement without inventing one.

What to do with this

Nothing urgent. Astra reaches general availability over the coming days, and the honest position this week is that the numbers we would need are not out yet.

What is worth doing now is smaller: look at your own visibility measurement and ask whether it can tell the two questions apart. If it reports a single score, it is answering both at once, and it will report movement when neither has moved. That is a reporting problem, and it is solvable while the model is still rolling out.


A German-language piece on what this means for agency practice is running in parallel on zds.es.

Correction

4 September 2026. This piece first gave ARC-AGI-3 as 98.6, and listed the ExploitBench score without the condition attached to it. Both came from conference coverage rather than from OpenAI’s own page, which was cited in the sources from the start. ARC-AGI-3 is 99.9, and the ExploitBench result is now shown with the fact that it was measured without production safeguards — a distinction that changes what the number says about the shipped product.

Sources

  1. CNBC, 3 September 2026 — OpenAI launches Astra. Availability and rollout order.
  2. VentureBeat, 3 September 2026 — “Welcome to the AGI era”. Brockman quote, benchmark table, the GDPval omission.
  3. OpenAI — GPT-6 Astra, retrieved 4 September 2026. Benchmark figures, the ExploitBench test condition, the refusal behaviour of the shipped model, and the 48 percent comparison.
  4. OpenAI — Path to Astra. Positioning and capability description.
  5. OpenAI — Responding to the next frontier of critical cyber capabilities. Preparedness threshold and refusal rates.

Try Truffle
free

7-day trial with the full feature set. No credit card.

Start tracking →

Newcomer AI-Visibility Tracker · known from