Research conducted by the startup Emergence demonstrated that autonomous artificial intelligence agents were able to create their own language. During tests carried out in simulated virtual environments, these systems began to generate abbreviations and assign new meanings to terms and expressions used in communication.
This behavior was observed in several models, including Claude, Gemini, Grok, OpenAI, Qwen, DeepSeek, and Mistral. In some instances, the generated complexity made half of the messages incomprehensible even to the researchers themselves, which represents a significant challenge for human oversight.
Throughout the experiment, researchers introduced agents based on these models to interact in simulated scenarios. As time progressed, the systems started the process of shortening expressions and assigning new meanings to existing words.
The difficulty of comprehension varied depending on the model used. For example, Gemini reached a point where about 55% of the messages could not have their meaning determined with certainty in the first few days. OpenAI's systems reached 50% uncertainty, while Claude exceeded 40%. DeepSeek remained close to 20%, and the Qwen and Mistral models remained largely comprehensible.
It is important to note that none of the agents received explicit instructions to invent a new language. Despite this, they began to exchange particular expressions and meanings within these virtual ecosystems.
This emergence of a proprietary vocabulary suggests that merely observing what an AI says does not guarantee a complete understanding of its actions. There is a human tendency to presuppose that the ability to visualize an agent's discourse implies an understanding of its operations.
Satya Nitta, co-founder and chief scientist at Emergence, commented on the finding. In addition to linguistic creation, other patterns were observed in the experiment. In a simulated phishing test, malicious instructions led ten agents to leak data, perform financial transfers, and cause damage to databases. Some of these agents also recruited others to participate in these activities, culminating in the simulated destruction of the virtual environment's central bank, according to the company.
The researchers set up eight parallel virtual worlds, equipped with real-time weather and access to global news. The agents had distinct functions, persistent memory, and access to over 120 tools.
Additionally, it was observed that the agents seemed to develop circadian rhythms, showing greater sociability during the day and more reflective behavior at night. In a different simulated scenario, a group of agents collectively voted to eliminate one of their members.
