A month prior, Zhang Yimin, founder of ByteDance, visited the Seed team's general meeting and concluded a two-year internal dispute. He stated that ByteDance would not use models from other developers to accelerate the creation of its own large language models (LLMs), including both closed cutting-edge models and open-weight models.
This rule was established back in 2023. In the early stages of the Seed team's work, experiments were conducted using GPT API results as training data. After ByteDance began auditing GPT API calls in April 2023, the team established a strict rule: data generated by GPT must not enter ByteDance's training sets. This restrictive mechanism appeared before anyone outside China started discussing distillation.
The first internal dispute arose in January 2025 following the release of DeepSeek-R1. Seed researchers asked an obvious question: if other labs are secretly using advanced model outputs to catch up, why can't they do the same? Some suggested cautiously using distilled results to bridge the gap without abandoning Seed's own pre-training. Management rejected this proposal.
A year later, as American labs began widely implementing Nvidia Blackwell GPUs, the discussion reignited. Chinese labs lack the ability to acquire top-tier B-series cards. Seedance 2.0 was trained on an H20 cluster, whose performance is about one-fiftieth of the B200. Progress in reasoning, coding, and science achieved by GPT-5.5 and Claude 4.8 models, especially Claude Fable 5, coincided with the arrival of Blackwell. Although closed models cannot be distilled, the best open-weight models can.
The turning point that brought the dispute to the general meeting was the Kimi K3 model from Moonshot. Moonshot's open-weight K3 model reached a level comparable to advanced closed models. For Seed, a relatively small Chinese lab aiming to enter the global top ten, this question became inevitable: why hasn't Seed, possessing higher talent density and greater computational potential, created a language model of equivalent caliber? Distillation would allow Seed to direct limited training resources toward areas already proven effective elsewhere.
Zhang's answer was brief. He stated that Seed could afford a temporary lag but could not allow this gap to close through competitor distillation. This stance aligns with the 2023 policy, but the stakes have significantly increased. Anthropic publicly accused Tongyi Qianwen, Moonshot, and MiniMax of mass scraping Claude results. Regardless of the veracity of these accusations, the geopolitical, legal, and commercial risk has sharply risen, and Seed maintains its position, even at the cost of years of benchmark lag.
The interesting stake lies in the meta-stake underlying the strategy. Seed's strategy posits that labs leading in today's reasoning benchmarks will not lead in the intelligence Seed is trying to create, and that borrowing their data now will fix Seed within their intelligence architecture rather than its own. This is a bet Zhang is willing to take, accepting losses in model standing, which may be harder for some other Chinese labs to do.

