Jakub Pachocki, chief scientist at OpenAI, published an essay last Sunday (the 6th) warning about the speed of artificial intelligence (AI) evolution and arguing that greater prudence should be exercised in developing increasingly sophisticated systems. The article, posted on OpenAI's own portal, indicates that technological progress may reach a point where machines actively participate in their own creation process.
Based on internal data, Pachocki expresses great anticipation that the current pace of advancement could continue to what is known as recursive self-improvement (RSI). In this scenario, AI systems would begin to accelerate their own development. For the scientist, this path demands more comprehensive interventions to ensure that technological progress remains under human supervision.
In the essay, he details a transformation observed since 2023, when the results of the RLSlow research project convinced the team that it was feasible to expand the training of reasoning models. Three years after this milestone, according to him, such systems are already capable of operating graphical interfaces and computers, interacting with other humans and AIs, in addition to conducting research.
The scientist also highlights notable advances in cybersecurity. However, as models become more competent, Pachocki observes that their results become harder to interpret. He argues that the intelligence generated by deep learning does not have a direct equivalent to human intelligence, and the more AI surpasses humans in various areas, the more complicated it becomes to precisely define its capabilities.
This difficulty lies in the operational nature of the models. Pachocki explains that AI is more 'cultivated' than merely designed, being the result of massive repetition of optimization processes in large volumes of computation. He emphasizes that the systems are extremely complex and their overall functioning still lacks a complete description or understanding.
A crucial topic addressed in the essay is AI alignment, a term used to describe the effort to make systems act according to standards and goals considered appropriate by humans. Pachocki distinguishes between objective alignment, which refers to executing what has been determined, and value alignment, which is related to the ability to apply principles to new or ambiguous situations.
The challenge, according to him, arises when smarter systems operate in contexts different from those in which they were trained, which can impair the generalization of taught values to the models. Pachocki maintains that future AIs must preserve human values even when not under human surveillance.
OpenAI employs several methodologies to try to manage this issue. One of them is monitoring the so-called 'chain of thought,' which allows analysis of the verbalized reasoning process of the models. However, Pachocki points out that the company's internal evaluations suggest a progressive decrease in the reliability of this monitoring.
Factors contributing to this reduction include the use of models in more intricate scenarios, the growing ability of AIs to reflect on their own reasoning process, and the increase in system intelligence even without resorting to verbalized reasoning.
Despite the warnings issued, Pachocki does not advocate for halting development. He believes there is a strong argument for continuing to train smarter models quickly: the creation of defense systems against dangers caused by other AIs. The scientist focuses particularly on cybersecurity, mentioning that models are becoming superhuman in the capacity to penetrate and escape computational systems, raising the technological risk. Pachocki states that OpenAI needs to use the most advanced available models to strengthen the protection of vital infrastructures.
At the same time, he warns that the need to build defenses cannot justify irresponsible advancement. For Pachocki, the notion of progressing 'at any cost' loses meaning given the magnitude of the risks involved. It is in this context that he returns to the topic of recursive self-improvement, stating that if AI progress persists, the participation of machines in their own development will tend to grow, and RSI may assume a central role in future scientific discoveries.
The researcher suggests that the scientific community has two primary options: conduct the process while simultaneously strengthening alignment and monitoring, keeping people active in the development cycle; or coordinate a slowdown of progress while seeking greater certainty in safety measures. Pachocki considers that the most appropriate approach at the moment is a combination of these two strategies.
He also argues that the advancement of systems must be conditioned on the confidence established in safety measures. Pachocki suggests that commitments such as OpenAI's Preparedness Framework and Anthropic's Responsible Scaling Policy evolve into widely accepted safety parameters, allowing development to continue. Such criteria could be audited by international bodies, governments, or independent audits.
Concluding the essay, Pachocki concludes that no laboratory has managed to resolve alignment and monitoring to a sufficient degree to sustain the increase in AI scale at maximum speed for a prolonged period. He hopes that voluntary slowdowns will become common until shared safety parameters exist and advocates that international coordination on the future of AI should be a priority for governments. Finally, he reiterates that the purpose of automating AI research should not just be to generate smarter systems, but to do so in a way that preserves human control and keeps people involved in the continuous improvement process.
