After a long period of hidden development, deepfake technology has made a qualitative leap that surprised even experts. Artificial intelligence has reached such a level of realism in generating faces, voices, and movements that it can now easily deceive people without special training or even automated verification systems.
Transition to Real-Time Synthesis
By 2026, a radical change is expected: the end of the era of pre-recorded clips and the beginning of the era of real-time synthesis. This means that fake content will be able to react and interact with a person during live communication.
From Static Videos to Digital Actors
Previously, a deepfake was a static video that could be analyzed later. However, new generation models ensure temporal consistency: movements look logical from frame to frame, facial features do not flicker or distort around the eyes or jaw, and the personality remains stable even when the 'performance' changes. This allows for the free combination of movement characteristics and identity, for example, applying one gesture to different faces or giving multiple expressions to one fake personality.
The practical consequence of this is that the deepfake ceases to be just a file and begins to function as a 'digital actor' capable of responding in real time to a video call or conversation. This significantly complicates manual verification, as many visual signs of manipulation, such as flickering, artifacts, or lighting inconsistencies, disappear.
Voice as the First Broken Element
Similar to the development of video, voice cloning has also overcome a barrier. Only a few seconds of audio sample are needed to create a convincing voice with natural intonation, rhythm, pauses, and even breathing, which can mislead the human ear in most everyday situations.
This progress already has direct consequences: large retailers report receiving over 1000 fraudulent calls generated by AI daily. Such a volume was unimaginable recently and demonstrates how synthetic voice fraud has turned from a theoretical risk into an industrial operation.
Factors for Rapid Scaling
The rapid spread of this phenomenon is explained by three factors. Firstly, technical realism: the most modern video models are designed specifically to maintain consistency between frames, eliminating visual errors that previously served as forensic evidence. Secondly, indistinguishable voice: voice cloning has exceeded what researchers call the 'threshold of indistinguishability,' making it unreliable to rely solely on hearing. Thirdly, almost zero technical barrier: the availability of generation tools and the automation of the process using AI agents allow almost anyone to create coherent and large-scale fake content without needing advanced technical knowledge.
According to security sector estimates, the growth rate is impressive: the number of detected deepfakes is projected to grow from approximately 500 thousand in 2023 to nearly eight million in 2025, corresponding to an annual increase of about 900%.
Detection and Defense Challenges
In light of this situation, the traditional defense method—analyzing pixels for visual errors—is losing its effectiveness. Researchers point out that current defense is shifting from human judgment to infrastructure protection, including:
- Source integrity: cryptographic signing of media files at the source, following standards such as C2PA (Coalition for Content Provenance and Authenticity).
- Multimodal forensic tools: systems that simultaneously analyze video, audio, and metadata instead of relying solely on visual inspection.
- Open-set detection models: academic initiatives, such as those developed in Brazil, that use genuinely human images as a standard for skin texture, light, and shadow, rather than trying to recognize only cataloged deepfake patterns, since it is impossible to predict all future forms of manipulation.
Practical Implications for Society
For businesses, the immediate risk is fraudulent attacks via voice or video, ranging from 'digital kidnapping' to a fictitious executive authorizing fund transfers. For public discourse, the threat lies in real-time disinformation: simulations of interviews, bulletins, or polls that appear authentic enough to spread and influence before any verification can take place.
The central point for 2026 remains that the line between 'real' and 'synthetic' is no longer a matter of technical quality but a matter of trust in infrastructure—here, the source signature, authentication, and digital literacy will carry more weight than a person's ability to 'feel' that something is fake.