
How NVIDIA GPUs help accelerate OpenAI's GPT-6 Astra Ultrafast
GPT-6 Astra Ultrafast, running on NVIDIA Blackwell GPUs, is available now in the OpenAI API and to eligible ChatGPT Work and Codex users. According to NVIDIA's blog, it offers up to 8x faster token generation than the Astra Standard mode.
GPT-6 Astra Ultrafast runs on NVIDIA Blackwell GPUs and is available now in the OpenAI API and to eligible ChatGPT Work and Codex users, according to NVIDIA's blog.
The mode offers up to 8x faster token generation than Astra Standard, powered by inference optimizations that tap into the capabilities of the Blackwell architecture, the post says.
Why it matters for agentic loops
Faster generation can shorten coding agents' edit-test-debug cycles, reduce the time spent generating responses between tool calls and make interactive applications feel more responsive. That matters most when the same work repeats: an agent writes code, uses a tool, checks the result and decides what to do next. Ultrafast is aimed at exactly these time-sensitive loops, and NVIDIA says its AI infrastructure helps OpenAI serve more useful model outputs when developers need them.
"NVIDIA's deep investment in tooling and documentation has enabled us to make our models exceptionally good at programming Blackwell and Rubin GPUs," said Philippe Tillet, inference lead at OpenAI. "Astra can turn that knowledge into high-performance kernels that make NVIDIA hardware compelling across the full frontier of latency, throughput and cost. With Astra Ultrafast, that means faster model responses as agents write code, use tools and work through complex tasks."
Performance keeps improving after deployment
Performance gains don't stop when a model is deployed, per the post. OpenAI is using its own models to help refine the inference software running on NVIDIA GPUs, taking advantage of the platform's programmability to test and implement improvements. That ongoing work can make responses faster and deployed infrastructure more productive over time.
"Our work with NVIDIA is helping us make AI faster and more useful," said Uday Ruddarraju, chief technology officer of compute at OpenAI. "We used our internal models to optimize inference on NVIDIA GPUs, and NVIDIA's programmability helped us deliver the acceleration behind Astra Ultrafast."
One platform across workloads
NVIDIA says a programmable platform lets developers and researchers reuse the same infrastructure across training, inference and reinforcement learning as models evolve. That flexibility helps teams repurpose compute as demand shifts, improving utilization and avoiding overprovisioning for each workload.
Developers can use GPT-6 Astra Ultrafast through the API today; the Ultrafast guide covers access, pricing and implementation details.
SiTech — AI-powered web development
We build fast, modern websites and bring AI into real business workflows. Have a project or a question? We'd love to help.