StartupNews · Breaking News
OPENAI UNVEILS "ULTRAFAST" MODE FOR GPT-5.6 SOL: A MODEL THAT ANSWERS 14 TIMES FASTER, WITHOUT GETTING DUMBER
Powered by Chipmaker Cerebras's Wafer-Sized Processors, the New Preview Tier Signals That the Next Battle in AI Won't Just Be Fought Over Intelligence — It Will Be Fought Over Milliseconds
By StartupNews · Published · Updated

For the past three years, the story of the AI race has largely been a story about intelligence: which model can reason more deeply, pass harder exams, write cleaner code, or hold a longer conversation without losing the thread. This week, OpenAI made the case that the next chapter of that race will be fought on a different battlefield entirely — speed. On August 13, 2026, the company introduced Ultrafast, a new service tier for its flagship GPT-5.6 Sol model that OpenAI says can generate responses up to 14 times faster than its standard processing tier, without sacrificing the model's underlying intelligence.
The headline number is striking on its own: up to 750 output tokens per second, a pace that would let the model produce a page of text in roughly the time it currently takes to read a single sentence. But the more interesting story sits underneath that number — in the unusual chip architecture powering it, the specific kinds of businesses OpenAI is targeting with it, and what the move signals about how the company believes AI will actually get used once speed stops being the bottleneck.
QUICK-LOOK SNAPSHOT
Company : OpenAI Announcement Date : August 13, 2026 Product : Ultrafast, a new service tier for GPT-5.6 Sol Speed Claim : Up to 14 times faster than Standard processing Peak Throughput : Up to 750 output tokens per second Availability : Limited preview, OpenAI API only, select customers Hardware Partner : Cerebras Systems (Wafer-Scale Engine, WSE-3) Underlying Model : GPT-5.6 Sol (part of the GPT-5.6 family, launched June 2026) Broader OpenAI-Cerebras Deal : 750 megawatts of wafer-scale compute capacity, rolling out 2026–2028, reportedly valued above $10 billion Target Use Cases : Voice AI, customer support, commerce, developer/coding agents, financial research, security incident response
WHAT ULTRAFAST ACTUALLY IS
Ultrafast is not a new model in the way GPT-5.6 Sol itself was a new model when OpenAI launched it in June 2026 alongside two siblings — Terra, positioned as a balanced, general-purpose option, and Luna, a smaller, speed-oriented model. Rather, Ultrafast is a new way of running the exact same GPT-5.6 Sol — OpenAI's most capable model in the 5.6 family — on specialised hardware engineered specifically to minimise the time between a user's request and the model's response.
That distinction matters enormously, and it is the detail OpenAI has leaned on hardest in its own public messaging. Historically, the tradeoff for developers building real-time AI products has been binary: choose a smaller, faster, less capable model to hit acceptable response times, or choose a larger, smarter, slower model and accept that users will wait. OpenAI's pitch for Ultrafast is that this tradeoff, at least for a specific set of workloads, no longer has to exist — that it is now possible to get frontier-level intelligence at speeds previously reserved for much smaller, lighter-weight systems.
HOW MUCH FASTER IS "14X"? A SIMPLE ILLUSTRATION
Standard GPT-5.6 Sol ██░░░░░░░░░░░░░░░░░░░░░░░░░░░░ ~54 tokens/sec (approx.) Ultrafast GPT-5.6 Sol ██████████████████████████████ up to 750 tokens/sec
(Illustrative comparison based on OpenAI's disclosed peak throughput figure for Ultrafast; standard-tier throughput varies by workload and is not precisely disclosed by OpenAI, so the bar above approximates the ratio implied by the company's "up to 14x" claim.)
To put 750 tokens per second in more human terms: a token is roughly three-quarters of a word, so 750 tokens per second works out to somewhere in the neighbourhood of 500 to 560 words every second. A task that might take a standard AI model half a minute to fully respond to — drafting a detailed email, summarising a long document, writing a block of code — could, in principle, complete in two to three seconds on Ultrafast.
THE HARDWARE BEHIND THE HEADLINE: HOW CEREBRAS MAKES THIS POSSIBLE
The technical story behind Ultrafast is arguably more interesting than the speed number itself, because it involves a genuinely unusual approach to chip design that stands apart from the graphics processing units, or GPUs, that have powered the vast majority of AI training and inference over the past decade.
Ultrafast runs on hardware built by Cerebras Systems, a Sunnyvale, California-based chipmaker that has spent close to a decade pursuing an idea most of the semiconductor industry considered impractical: instead of cutting a silicon wafer into dozens or hundreds of individual chips the way virtually every other chipmaker does, Cerebras builds a single, enormous chip out of an entire wafer. Its latest generation, the WSE-3, is roughly the size of a dinner plate — many times larger than any conventional processor — and packs a staggering 44 gigabytes of on-chip SRAM memory directly alongside its compute cores.
WHY WAFER-SCALE CHIPS ARE FASTER FOR THIS SPECIFIC JOB
Traditional GPU setup: [Compute cores] <--- data bottleneck ---> [Memory, off-chip] Cerebras wafer-scale: [Compute cores + 44GB memory, all on one wafer, no bottleneck]
That architectural difference addresses what engineers call the "memory bandwidth bottleneck" — the delay caused every time a chip's processing cores have to reach outside themselves to fetch data (in this case, a model's parameters, essentially its learned "knowledge") from separate memory chips sitting nearby on a circuit board. In a typical GPU-based system, this back-and-forth between compute and memory happens constantly during inference, and it is one of the primary reasons response generation has speed limits even on extremely powerful hardware. By keeping the model's data physically on the same piece of silicon as the processing cores, Cerebras's wafer-scale design largely eliminates that back-and-forth, which is the central technical reason the company can claim such a dramatic speed advantage. Cerebras has separately stated that its systems can produce responses up to 15 times faster than comparable GPU-based systems on certain workloads — a figure closely in line with the 14x speedup OpenAI is now advertising specifically for GPT-5.6 Sol.
This is not a brand-new relationship. OpenAI and Cerebras have reportedly been in dialogue since as far back as 2017, when both companies were relatively early in their respective, ambitious missions — OpenAI pursuing advanced general-purpose AI, Cerebras pursuing a fundamentally different approach to chip architecture that defied the conventional wisdom of the semiconductor industry. That long relationship culminated, earlier in 2026, in a considerably larger commercial agreement: a multi-year deal for OpenAI to deploy 750 megawatts of Cerebras's wafer-scale computing capacity, rolling out in phases from 2026 through 2028, in a partnership reported to be worth more than $10 billion — making it, by several industry accounts, the largest high-speed AI inference deployment announced anywhere in the world to date. Ultrafast is the first major, customer-facing product to emerge from that broader infrastructure partnership.
WHO GETS TO USE IT, AND FOR WHAT
For now, Ultrafast is being made available only through the OpenAI API, and only to a limited, hand-selected group of customers as part of a preview program — it is not yet available inside the consumer ChatGPT app, and OpenAI has been explicit that broader access will expand gradually "as capacity grows," a reflection of the fact that Cerebras's specialised wafer-scale hardware is, by its nature, far more expensive and complex to manufacture and deploy at scale than conventional GPU infrastructure.
OpenAI has identified a fairly specific cluster of use cases where it believes the speed increase changes what is practically possible, rather than simply making an existing experience marginally more pleasant:
VOICE AND CONVERSATIONAL AI, where natural, human-paced back-and-forth conversation depends on response latency low enough that pauses don't feel awkward or robotic — a threshold that has historically been difficult to hit with frontier-scale models.
CUSTOMER SUPPORT AND COMMERCE, where a slow AI response during a live chat or checkout flow can directly translate into an abandoned interaction or lost sale.
DEVELOPER AND CODING AGENTS, where AI tools increasingly chain together many sequential steps — reading code, running tests, revising, re-testing — and where the cumulative delay across dozens of these steps compounds quickly if each individual step is slow.
FINANCIAL RESEARCH, where trading and analysis workflows are often explicitly time-sensitive, and where being seconds faster than a competitor's analysis can carry direct commercial value.
SECURITY INCIDENT RESPONSE, a use case OpenAI has highlighted from its own internal operations: reading system logs, analysing traces, synthesising what happened during an active incident, and helping validate a fix — all tasks where a security team's ability to respond quickly can materially affect how much damage an incident causes.
OpenAI has said its own internal teams are already using Ultrafast for exactly this kind of incident-response work, as well as for research workflows involving searches across internal knowledge bases, data queries and information synthesis — tasks the company says it is exploring compressing from what used to be overnight, batch-processed jobs into work that can instead be done interactively, in real time, during the working day.
Early external feedback echoes that framing. John Crepezzi, who works on AI assistants at the quantitative trading firm Jane Street — one of the initial preview customers — was quoted describing the speed increase brought by Cerebras as enabling genuinely different ways of using the models, making it practical for developers to work in a more focused, productive rhythm alongside AI tools rather than waiting on them.
HOW FAST IS FAST? BENCHMARK COMPARISONS
Beyond the headline 14x and 750-tokens-per-second figures, independent testing has offered a few more concrete points of comparison that help illustrate just how large a gap Ultrafast opens up relative to standard AI inference.
On Humanity's Last Exam, a notoriously difficult benchmark spanning 2,500 expert-level questions across dozens of academic disciplines, GPT-5.6 Sol running on Ultrafast reportedly completed the entire exam set in 11 hours and 11 minutes. For comparison, a competing frontier model — Claude Fable 5, run under standard processing conditions — took 78 hours and 27 minutes to complete the same benchmark, a roughly sevenfold difference in total completion time. It's worth noting this comparison reflects overall task-completion time on a benchmark run under specific testing conditions, not necessarily a direct, apples-to-apples measure of raw per-token generation speed across identical workloads — but it nonetheless offers a vivid, real-world illustration of how throughput differences compound across long, complex tasks.
Separately, Cerebras has reported a 5.6x speedup on GDP-Val, a benchmark designed to measure AI model performance on real-world, economically valuable tasks, with the company stating this speed increase came without any corresponding drop in output quality — an important claim, since the entire value proposition of Ultrafast depends on speed and intelligence not being traded off against one another.
WHY COMPOUNDING SPEED MATTERS MOST FOR MULTI-STEP TASKS
Single AI response: Modest time saved 10-step agent workflow: Time savings compound significantly 100-step agent workflow: Time savings compound dramatically
Illustrative logic: an AI agent completing a 30-second task on standard infrastructure could, at a roughly 14x speedup, complete the same task in around 2 seconds — and workflows that chain together many such steps see those savings multiply with each additional step in the chain.
This compounding effect is precisely why OpenAI has pointed to agentic workflows — AI systems that autonomously chain together many sequential actions, such as coding agents that write code, run it, read the results, and revise it repeatedly — as one of the categories that stands to benefit most disproportionately from a throughput increase of this scale. A delay that feels tolerable once, when a human is waiting for a single answer, becomes a serious drag on productivity when it's multiplied across dozens or hundreds of automated steps within a single agent task.
THE STRATEGIC CONTEXT: WHY SPEED, WHY NOW
OpenAI's own public framing of Ultrafast is notably explicit about the strategic thinking behind it. In the company's announcement, it argued that until now, achieving real-time speed with AI models has typically required developers to choose a smaller or more specialised model, accepting reduced intelligence as the price of faster responses. Ultrafast, in OpenAI's telling, represents progress in a different direction entirely — not a smaller, faster model, but more useful work completed per second from the same, full-strength frontier model. The company's stated view is that once speed stops requiring a tradeoff against intelligence, AI can move into the most time-sensitive parts of a business, unlocking categories of application that were previously impractical regardless of how capable the underlying model was.
That framing lands at a particular moment in the broader AI industry, one where the competitive conversation has visibly begun shifting beyond pure benchmark performance toward questions of cost, latency and deployment efficiency — the practical, unglamorous factors that determine whether a capable model can actually be built into a real product that millions of people use every day, rather than remaining an impressive but slow demonstration. OpenAI has acknowledged this shift directly, noting that rivals including Anthropic have similarly introduced accelerated versions of their own models, a sign that the entire frontier AI industry increasingly views inference speed as a genuine competitive battleground in its own right, not merely a secondary engineering concern behind the pursuit of raw intelligence.
THE COMPETITIVE LANDSCAPE: A CROWDED RACE FOR THE FASTEST TOKEN
OpenAI's move into ultra-fast inference places it in an increasingly crowded field of chipmakers and AI labs all racing to solve the same underlying problem, and the competitive dynamics among the hardware providers themselves are almost as interesting as the AI labs building on top of them.
SELECTED PLAYERS IN THE HIGH-SPEED AI INFERENCE RACE
Company Approach Notable Position ──────────────────────────────────────────────────────────────────────────── Cerebras Wafer-scale chips (WSE-3), 44GB Powers OpenAI's on-chip SRAM, no off-chip memory Ultrafast tier; bottleneck went public on Nasdaq in May 2026 Groq Purpose-built "LPU" (Language Acquired by Nvidia; Processing Unit) chips designed Nvidia has said it specifically for low-latency, will allocate roughly sequential token generation a quarter of its total data-centre capacity to Groq- based deployments Nvidia Dominant GPU maker; pairing its Positions Groq as a Blackwell/Vera Rubin architecture premium, ultra-fast with Groq's acquired technology tier alongside its for a blended high-speed offering broader GPU fleet Google Custom TPU v7 ("Ironwood") chips, Controls a majority claiming substantially better share of the custom performance-per-watt than cloud AI accelerator comparable GPU hardware market AWS Custom Trainium3 chips; also the First hyperscaler to first major cloud provider to offer Cerebras chips deploy Cerebras hardware directly through its own inside its own data centres cloud platform
The scale of the OpenAI-Cerebras partnership — 750 megawatts of dedicated capacity, reportedly worth more than $10 billion over its multi-year term — has been described by industry analysts as a pivotal validation moment for Cerebras specifically, arriving just weeks before the chipmaker's own initial public offering in May 2026. Landing a customer as prominent as OpenAI for a deployment of this scale gave Cerebras a powerful reference case to present to public-market investors evaluating whether wafer-scale computing represents a genuine, durable alternative to GPU-dominated infrastructure, or a narrower niche technology suited only to specific workloads.
At the same time, competitive responses from larger players have moved quickly. Nvidia's decision to acquire Groq and integrate its low-latency chip technology alongside its own dominant GPU lineup has been widely read as a direct competitive answer to the inference-speed threat that companies like Cerebras represent — with Nvidia's own chief executive publicly stating the company would allocate a meaningful share of its total data-centre capacity toward Groq-style, ultra-fast inference deployments going forward. Cerebras executives have, notably, pointed to that very move by Nvidia as validation of their own long-held thesis: that the real economic battleground in AI increasingly lies in ultra-fast inference, where the fastest tokens command the highest commercial value, and where memory-rich, specialised architectures are uniquely positioned to win that fight.
THE GPT-5.6 FAMILY: WHERE SOL FITS
To understand why OpenAI chose GPT-5.6 Sol specifically as the model to accelerate, it helps to understand how the broader GPT-5.6 family is structured. When OpenAI introduced the family in June 2026, it launched three distinct models simultaneously, each tuned for a different point on the tradeoff between raw capability and operating efficiency:
GPT-5.6 SOL — the most capable, "frontier" model in the family, designed for the hardest reasoning, coding and analysis tasks GPT-5.6 TERRA — positioned as a balanced, general-purpose option for everyday use cases GPT-5.6 LUNA — a smaller, speed-and-efficiency-focused model for latency-sensitive or cost-sensitive applications
The full family became broadly available across ChatGPT, OpenAI's Codex coding tool, and the API by July 2026. In that original three-model structure, Luna was effectively the "fast" option, while Sol represented the ceiling of what the family could do intellectually, at the cost of being the slowest and most computationally expensive to run. Ultrafast effectively collapses that tradeoff for a specific slice of customers: rather than stepping down to Luna to get real-time speed, qualifying developers can now, in principle, get Sol-level intelligence delivered at speeds that rival or exceed what a much smaller model would typically achieve — provided their workload fits within the constraints of the current limited preview.
WHAT COMES NEXT
OpenAI has been candid that Ultrafast remains an early-stage preview rather than a fully mature, generally available product. Access is currently restricted to a small group of customers, chosen deliberately so that OpenAI can study, in live production environments, exactly which categories of business workflow benefit most from an order-of-magnitude increase in inference speed — insight the company says will directly inform how it builds and prices the product as it scales toward wider release. The company has said explicitly that access will expand as capacity grows, a dependency that ties Ultrafast's rollout timeline directly to the broader, multi-year buildout of Cerebras's wafer-scale infrastructure through 2028.
For the wider AI industry, the launch adds fresh momentum to what is quickly becoming one of 2026's defining infrastructure storylines: a genuine, multi-front contest over who can deliver frontier-level AI intelligence at the lowest possible latency, fought not just between AI labs like OpenAI and Anthropic, but equally between the chipmakers — Nvidia, Cerebras, Google, AWS and others — racing to build the physical hardware underneath them. If OpenAI's bet proves out, and speed increases of this magnitude can be sustained and scaled without compromising model quality, the practical effect may be less about any single flashy new feature and more about a quiet expansion of where AI can be usefully deployed at all — moving frontier intelligence out of the realm of "impressive but slow chatbot" and into the time-critical center of live business operations, from trading floors to security operations centers to the checkout flow of an online store.
A GLOSSARY FOR READERS NEW TO AI INFRASTRUCTURE
For readers less familiar with the technical vocabulary that comes up repeatedly in stories about AI speed and infrastructure, a few terms are worth defining plainly:
TOKEN — the basic unit of text an AI language model processes and generates, roughly equivalent to three-quarters of a word on average. Model speed is typically measured in tokens generated per second, since that unit is more precise than counting whole words.
INFERENCE — the process of actually running a trained AI model to generate a response to a user's request, as distinct from "training," the earlier, much more computationally expensive process of teaching the model in the first place. Ultrafast is entirely about speeding up inference, not training.
LATENCY — the delay between when a request is sent to an AI model and when its response begins arriving; low latency is especially critical for real-time applications like voice conversations, where even small delays can make an interaction feel unnatural.
GPU (GRAPHICS PROCESSING UNIT) — the type of chip that has powered the overwhelming majority of AI training and inference over the past decade, originally designed for rendering graphics but repurposed for AI due to its ability to perform many calculations in parallel. Nvidia is the dominant maker of AI-focused GPUs.
WAFER-SCALE CHIP — Cerebras's distinctive approach to chip design, in which an entire silicon wafer is used to build a single, enormous processor, rather than being cut into many smaller individual chips as is standard practice across the rest of the semiconductor industry.
SRAM (STATIC RANDOM-ACCESS MEMORY) — a fast, but expensive and space-intensive, type of computer memory. Cerebras's chips place an unusually large amount of SRAM directly on the same silicon as their processing cores, which is the central architectural choice behind their speed advantage.
MEMORY BANDWIDTH BOTTLENECK — the delay caused when a chip's processing cores must retrieve data from memory located physically elsewhere on a circuit board, rather than on the same chip; this bottleneck is one of the primary limits on how fast conventional GPU-based AI inference can run.
SERVICE TIER — a pricing and performance category offered by an AI provider (such as OpenAI's "Standard" versus "Ultrafast" tiers), typically reflecting different tradeoffs between cost, speed, and access to particular hardware.
AGENTIC WORKFLOW — a task in which an AI system autonomously chains together multiple sequential actions — for example, writing code, running it, reading the results, and revising it — often with limited or no human intervention at each individual step.
FREQUENTLY ASKED QUESTIONS
WHAT EXACTLY DID OPENAI ANNOUNCE? A new service tier called Ultrafast for its GPT-5.6 Sol model, which the company says can generate responses up to 14 times faster than its Standard processing tier, reaching peak throughput of up to 750 output tokens per second.
IS THIS A NEW, DIFFERENT AI MODEL? No. Ultrafast runs the exact same GPT-5.6 Sol model that already exists — it is a different way of running that model on specialised hardware, not a smaller or altered version of it. OpenAI has emphasised that the model's underlying intelligence is unchanged; only the speed of generating responses is different.
WHAT HARDWARE MAKES THIS POSSIBLE? Ultrafast runs on wafer-scale chips built by Cerebras Systems, which places an unusually large amount of on-chip memory directly alongside its processing cores, largely eliminating a data-retrieval delay that limits speed on conventional GPU-based systems.
CAN I USE ULTRAFAST IN CHATGPT TODAY? Not yet. Ultrafast is currently available only through the OpenAI API, and only to a limited group of preview customers selected by OpenAI. It is not available in the consumer ChatGPT app at this time.
WHY IS OPENAI LIMITING ACCESS? Cerebras's wafer-scale hardware is more complex and expensive to manufacture and deploy at scale than conventional GPU infrastructure, and OpenAI has said it wants to study which business workflows benefit most from the speed increase during this preview period before expanding access more broadly.
WHAT KINDS OF TASKS BENEFIT MOST FROM ULTRAFAST? OpenAI has pointed to real-time or near-real-time workflows including voice conversations, live customer support, e-commerce checkout flows, developer and coding agents that chain together many sequential steps, financial research, and security incident response.
HOW DOES THIS COMPARE TO WHAT COMPETITORS ARE DOING? OpenAI has acknowledged that rival AI labs, including Anthropic, have introduced their own accelerated model tiers. On the hardware side, Cerebras competes with a range of alternative approaches to fast inference, including Groq's purpose-built chips (now owned by Nvidia), Google's custom TPU processors, and Amazon Web Services' custom Trainium chips.
WILL ULTRAFAST COST MORE THAN STANDARD PROCESSING? OpenAI has not disclosed specific pricing for the Ultrafast tier as of this preview announcement. Given the specialised, more expensive nature of the underlying wafer-scale hardware, it is reasonable to expect some premium relative to standard processing, though the company has not confirmed exact figures publicly.
THE BUSINESS CASE: WHY COMPANIES MIGHT PAY MORE FOR SPEED
For any business evaluating whether to adopt a faster but likely more expensive AI service tier, the underlying calculation typically comes down to a simple question: does the value created by a faster response exceed the additional cost of generating it? For a handful of the use cases OpenAI is targeting with Ultrafast, that calculation appears to favour speed fairly clearly.
Consider financial trading and research, where firms have historically paid substantial premiums for infrastructure that shaves even single-digit milliseconds off decision-making and execution times, on the logic that being first to act on new information carries direct, quantifiable value. A firm like Jane Street, among the initial Ultrafast preview customers, operates in exactly this kind of environment, where the difference between an AI system that takes thirty seconds to synthesise a piece of research and one that takes two seconds is not merely a matter of convenience but potentially a matter of competitive advantage measured in real trading outcomes.
Customer support and e-commerce present a related but distinct case: here, the value of speed is less about split-second competitive advantage and more about conversion and retention. Research across digital commerce has repeatedly shown that even small increases in page-load or response delay correlate with measurably higher rates of users abandoning a task, whether that's a purchase, a support conversation, or a search query. An AI customer-support agent that responds in under a second, at a pace closer to human conversational rhythm, is likely to keep more users engaged through to a resolved outcome than one that introduces multi-second pauses at every exchange.
For coding and developer agents, the calculation is different again, and arguably the most straightforward of the group: because these workflows chain together many discrete steps — write code, execute it, read results, revise, re-execute — the delay at each individual step compounds across the full task. A fourteen-fold speedup applied once might save a developer a few seconds; applied across fifty sequential steps in an autonomous coding session, it can be the difference between a task that completes in under a minute and one that takes the better part of an hour, fundamentally changing how developers choose to use — or not use — such tools in their daily workflow.
A DECADE-LONG PARTNERSHIP, VALIDATED
One detail in this story worth dwelling on is just how long the relationship between OpenAI and Cerebras has actually been developing before reaching this point. According to statements from both companies, engineering teams at OpenAI and Cerebras have been in regular dialogue since 2017 — a full six years before ChatGPT's public launch reshaped the entire technology industry's understanding of what large language models could do. At that time, both companies were pursuing ideas widely considered speculative, even fringe, within their respective corners of the technology world: OpenAI's founding mission to build software capable of general intelligence, and Cerebras's decision to defy decades of conventional semiconductor manufacturing wisdom by building single chips out of entire silicon wafers rather than the small, individually cut dies used by the rest of the industry.
Sachin Katti, who leads infrastructure strategy at OpenAI, has described the company's broader hardware approach as building a deliberately diversified portfolio of computing systems, matching different types of chips to different types of workloads rather than relying on any single vendor or architecture. In that framing, Cerebras's wafer-scale systems serve a specific, high-value niche within OpenAI's much larger computing footprint — dedicated specifically to the subset of workloads where minimising latency matters more than any other single factor, sitting alongside OpenAI's far larger investments in traditional GPU infrastructure for training and general-purpose inference.
For Cerebras, landing this deal — and, more specifically, seeing it translate into a highly visible, named consumer-facing product like Ultrafast rather than remaining a behind-the-scenes infrastructure arrangement — represents a significant business milestone. The company's own public statements have described 2026 as an extraordinary year on the back of this partnership, one that Cerebras believes will bring its wafer-scale technology to hundreds of millions, and eventually billions, of end users indirectly, through products built by OpenAI and, potentially, other future partners. The timing also mattered for Cerebras's own path to public markets: the company completed its initial public offering on the Nasdaq in May 2026, and having a marquee reference customer like OpenAI already committed to a multi-billion-dollar, multi-year deployment gave prospective public-market investors a concrete, high-profile proof point for evaluating whether wafer-scale computing represents a durable, investable alternative to the GPU-dominated status quo, or a narrower technology suited only to a limited set of specialised use cases.
KEY FACTS AT A GLANCE
• Product: Ultrafast, a new service tier for GPT-5.6 Sol • Announced: August 13, 2026 • Speed claim: Up to 14x faster than Standard processing • Peak throughput: Up to 750 output tokens per second • Hardware: Cerebras Systems' wafer-scale WSE-3 chips, with 44GB of on-chip SRAM memory • Broader partnership: 750 megawatts of Cerebras compute capacity for OpenAI, deploying 2026–2028, reportedly worth over $10 billion • Access: Limited preview, OpenAI API only, select customers; not yet available in consumer ChatGPT • Target workloads: Voice AI, customer support, commerce, coding/developer agents, financial research, security incident response • Benchmark data point: GPT-5.6 Sol on Ultrafast completed Humanity's Last Exam (2,500 questions) in roughly 11 hours, versus 78-plus hours for a competing frontier model under standard processing
THE BOTTOM LINE
OpenAI's Ultrafast preview is, in one sense, a fairly narrow product announcement: a new, restricted-access API tier available to a handful of customers for now. But the underlying signal is considerably larger than the immediate rollout. By pairing its most capable model with genuinely novel, purpose-built hardware rather than simply asking customers to trade intelligence for speed, OpenAI is making an explicit bet that the next phase of the AI race will be decided as much by how fast a frontier model can respond as by how smart it is — and that the two no longer have to be in tension. Whether that bet pays off at scale will depend on questions that remain unresolved today: how quickly Cerebras's specialised, expensive-to-manufacture wafer-scale hardware can be produced and deployed at the volume OpenAI's full customer base would require, how aggressively competitors like Nvidia's newly acquired Groq technology can close the speed gap, and whether the specific business workflows OpenAI is targeting — voice, commerce, security, finance, coding agents — actually generate enough commercial value from faster responses to justify what wafer-scale inference costs to run. For now, though, OpenAI has made its opening move unmistakably clear: in the next chapter of the AI race, milliseconds may matter every bit as much as IQ points.
DISCLAIMER: This article is intended for general informational and news purposes only. Performance figures and speed comparisons cited above are drawn from statements made by OpenAI, Cerebras Systems, and third-party technology reporting current as of publication, and reflect claims made by the companies involved and their preview customers rather than independently, universally reproduced benchmark testing. Actual performance may vary by workload, and businesses evaluating AI infrastructure providers should conduct their own testing before making procurement decisions.