OpenAI announced a new processing feature for GPT-5.6 Sol on Thursday (13), named Ultrafast. This new mode has the potential to increase response speed by up to fourteen times compared to the established standard, making it ideal for scenarios requiring immediate decisions and feedback.
Initially, this service is accessible via the API and utilizes infrastructure provided by Cerebras, which supports an output rate of up to 750 tokens per second. Currently, access is limited to a select group of clients while the company monitors its use in real environments and plans a gradual capacity expansion.
OpenAI's strategy aims to eliminate the need to choose between highly intelligent models and agile responses. With Ultrafast, the organization intends to apply GPT-5.6 Sol in various areas, including customer service, financial analysis, incident response, commerce, and interactive research.
OpenAI characterizes Ultrafast as a new class of service intended for applications where the time between posing a question and receiving an answer can directly impact the utility of artificial intelligence. The company claims that the efficiency advancements of GPT-5.6 have enabled a closer convergence between reasoning capability and operational speed.
One example of application presented is system failure response. In this context, the tool can analyze application reports, recent code modifications, and engineer reports to help identify the cause of a problem and prepare a temporary fix while the outage persists.
Additionally, applications in the financial and security sectors are cited, covering market indicator analysis, transaction evaluation, and detection of unusual patterns in rapidly changing contexts.
In customer support, the extra speed allows systems to manage more complex issues during the interaction itself. The proposal also covers voice functionalities, where a delay in response can harm conversation fluidity.
In e-commerce, the model can be used to answer product queries, confirm stock, suggest personalized items, and assist in resolving difficulties during the purchase process, aiming to act before slowness causes consumer abandonment.
Another field explored is research and experimentation. According to the company, tasks that previously required long execution periods can now become interactive cycles, allowing researchers to test a hypothesis, observe results, make adjustments, and repeat the experiment within the same working period.
The initial implementation phase includes companies in the programming, commerce, finance, and support segments. Among the clients mentioned by OpenAI are Jane Street, Podium, Basis, and Rogo.
Jane Street, through John Crepezzi, an AI assistant specialist, assessed that the speed increase offered by Cerebras modifies the usage possibilities of the models, promoting a more focused and productive relationship between developers and AI.
OpenAI decided to start with corporate applications to observe the technology in production operations. The goal is to determine in which activities a significant jump in speed generates the greatest benefit and how the products behave when closely tracking user pace.
OpenAI itself has also subjected Ultrafast to internal tests, with a special focus on reliability incidents, where teams need to quickly interpret constantly changing information during an event.
In these cases, the system is used to consolidate log and trace data, organize dialogues, propose checks, and assist in drafting or validating fixes, although the final responsibility for decision and implementation remains with the engineers.
In the research area, internal teams use the mode to query knowledge bases, access data, and compile, structure, and synthesize information through integrated tools. The distinction lies in the rhythm of the work cycle: instead of leaving sequences of experiments to run overnight and analyze them the next day, Ultrafast enables multiple rounds of investigation during business hours.
The infrastructure supporting this new feature results from the collaboration between OpenAI and Cerebras, responsible for providing the low-latency inference technology used in the service. With this architecture, GPT-5.6 Sol can achieve the announced rate of up to 750 tokens per second, which OpenAI relates to the ability to create more reactive products and integrate AI into processes where delay between phases constitutes a limitation.
The company emphasizes that Ultrafast is a preliminary version, available only to a limited number of clients, and the expansion of access will depend on increased capacity. Those interested in following this evolution can sign up to receive availability notifications, as the information collected during the testing phase will guide the next stages of implementation.



