AdvertisementArtificial intelligenceTechTech Trends

China faces new AI bottleneck as it runs out of Chinese-language training data

The global supply of high-quality, publicly available human-generated text could be fully exhausted within the next six years

3-MIN READ3-MIN Listen

A visitor looks at a large Chinese language dictionary at an exhibition in the West Kowloon Cultural District in Hong Kong on May 12, 2025. Photo: Jelly Tse

Ben Jiangin BeijingPublished: 10:00am, 8 Aug 2026Updated: 10:04am, 8 Aug 2026

China’s high-stakes race to build next-generation artificial intelligence models is entering a critical new phase, where a less visible yet far more existential threat is coming into view: a severe shortage of high-quality training data.

While the US chokehold on advanced computing chips has dominated headlines, Chinese AI experts increasingly warn that running out of quality data could prove to be the next major bottleneck to the nation’s tech ambitions – and one that hardware workarounds cannot easily solve.

It is a challenge confronting AI giants on both sides of the Pacific – and some US companies are already resorting to aggressive measures to stay ahead.

The global supply of high-quality, publicly available human-generated text could be fully exhausted within the next six years, according to US-based research institute Epoch AI.

OpenAI co-founder Andrej Karpathy has also warned of a looming “data wall” by the end of this decade, beyond which model capabilities could hit a plateau unless they were fed fresh, reliable information.

Top American labs are spending lavishly to mine offline human knowledge, igniting a fierce ethical debate in the process.

AdvertisementSelect VoiceSelect Speed00:0000:001.00x

Leave a Reply

Your email address will not be published. Required fields are marked *