OpenAI’s GPT-6 Astra Ultrafast is available through its developer API and to eligible ChatGPT Work and Codex users, NVIDIA confirmed on October 1, 2026. NVIDIA says the mode generates tokens—the units of model output—up to eight times faster than Astra Standard. The pitch is shorter waits inside workflows where an AI system repeatedly writes code, calls tools and checks results.
NVIDIA’s October 1 explanation adds a concrete detail to that rollout: OpenAI used its own internal models to optimize the software running inference on NVIDIA GPUs. Inference is the process of running a trained model to produce responses. The acceleration therefore involves both Blackwell hardware and changes to the software that serves Astra, rather than a hardware announcement alone.
Models help tune the software serving them
OpenAI compute chief technology officer Uday Ruddarraju credited NVIDIA’s programmability with helping deliver Ultrafast’s acceleration. NVIDIA describes an ongoing process: OpenAI’s models help test and implement improvements to inference software after deployment. The company says those refinements can make responses faster and infrastructure more productive over time. That is a claim about continued optimization, not a fixed endpoint reached when the model first ships.
Philippe Tillet, OpenAI’s inference lead, pointed to another part of that work. He said NVIDIA’s tooling and documentation helped OpenAI make its models good at programming Blackwell and Rubin GPUs. In his description, Astra turns that knowledge into high-performance kernels: the GPU programs involved in running computations. Tillet linked those improvements to latency, throughput and cost—how long responses take, how much work the infrastructure handles and what serving it costs.
We used our internal models to optimize inference on NVIDIA GPUs, and NVIDIA’s programmability helped us deliver the acceleration behind Astra Ultrafast.
Uday Ruddarraju, chief technology officer of compute at OpenAI, quoted by NVIDIA
The waiting repeats between tool calls
NVIDIA’s case for faster generation focuses on repeated steps, not just a single answer appearing sooner. A coding agent can write an edit, run a tool, inspect the result and decide its next move. Each turn creates another opportunity to wait for model output. NVIDIA says Ultrafast can shorten edit-test-debug cycles and reduce response-generation time between tool calls, while making interactive applications feel more responsive.
The wording of the speed claim matters. NVIDIA’s eightfold figure compares token generation with Astra Standard; it is not stated as an eightfold reduction in the time required to finish an entire coding task. Its workflow examples explain where faster output could help, but the stated measurement remains generation speed. “Up to” also makes the figure a maximum claim rather than a promised rate for every request.
Bedrock adds another route, with a different speed claim
AWS announced Astra Ultrafast on Amazon Bedrock on September 30. Its announcement attributes to OpenAI a claim of up to six times faster API inference, with up to 300 tokens per second. That is a different stated comparison from NVIDIA’s eightfold token-generation claim. The two figures should not be collapsed into one universal speed guarantee across every access route.
Confirmed access routes
- OpenAI API: NVIDIA says developers can use Astra Ultrafast now.
- ChatGPT Work and Codex: NVIDIA confirms availability for eligible users, not unrestricted access for every account.
- Amazon Bedrock: AWS offers access through its console or supported Bedrock APIs.
AWS positions its offering for real-time coding assistants, interactive agents and customer-facing applications. It also points to existing AWS controls for securing workloads, governing access and auditing model invocations. Those operational features accompany the speed tier; they are distinct from its performance claim. AWS directs customers to Bedrock documentation for supported regions, endpoints, features and pricing.
Reader comments
Newest comments first. Replies stay oldest first.