Even trillion-parameter models can’t salvage dirty data, domestic large language models are starting to redo pre-training.
31 minutes ago
Beating AI Insight News Flash. Leiphone reports that over the past six months, multiple domestic large language model (LLM) vendors have begun redoing their pre-training processes, with all issues traced to data quality. One vendor spent nearly a year training a trillion-parameter model, only to have its performance outperformed by a small model with just around 1% of the parameter count. The team ultimately found the root cause lay in the training data: web spam, duplicate corpora, and low-quality annotations had been mixed into the training set, rendering even the largest models ineffective. The case comes from an anonymous source, and the specific company has not been disclosed. Tencent Hunyuan has publicly announced a full redo of its pre-training work. Since February this year, the team has rebuilt its pre-training and reinforcement learning infrastructure. Previous media reports revealed that the old version of Hunyuan had issues including ranking manipulation data being mixed into the training set and inconsistent annotation rules. The new team redefined data standards, cleaned up existing corpora, and Hy3 listed data quality and diversity as key improvement areas. Alibaba and Baidu are also stepping up data governance efforts. Alibaba’s Qwen3 uses the Qwen2.5 series models to clean documents, and has significantly supplemented synthetic math and code data. Baidu’s Ernie 4.5 has added deduplication, low-quality filtering, data mapping, and manual review processes. Neither company has publicly stated that they had to redo work due to dirty data, but both have invested more efforts in data processing for their new-generation models.
NEAR has been deployed to the Hyperliquid spot market.
11 minutes ago
Binance lists three bStocks tokenized securities trading pairs.
11 minutes ago
US Treasury May Deploy Nearly $1 Trillion in Fiscal Cash to Stabilize Money Markets
11 minutes ago
Binance adds new trading pairs ARB/USDT and ENA/USD1 to its full-position margin trading.
11 minutes ago
Genius Terminal co-founder teases new gameplay for Genius.fun, token foundation page expected to launch.
11 minutes ago
Anthropic前员工做「自我改进AI」,3个月估值从10亿涨到50亿美元
11 minutes ago
Hot feeds
A trader profits $448K by monitoring #Binance's new listings!
2024.12.13 17:37:29
Last week, funds have flowed into #Bitcoin, #Ethereum, and #Hyperliquid.
2024.12.16 14:48:36
A $PEPE whale that had been dormant for 600 days transferred all 2.1T $PEPE($52M) to a new address.
2024.12.14 10:35:27
When Elon Musk tweeted about Moltbook, the meme coin MOLT experienced a short-term 30% price surge, hitting a new all-time high of $114 million.
2026.01.31 18:37:29
A smart #AI coin trader made $17.6M on $GOAT, $ai16z, $Fartcoin,$arc.
2025.01.05 16:05:18
A sniper earned 2,277 $ETH ($8.3M) trading $SHIRO within 18 hours!
2024.12.03 23:09:08
MoreHot Articles

How did I turn $1,000 into $30,000 with smart money?
2024.12.09

10 promising AI Agent cryptos
2024.12.05

The 30-Year-Old Entrepreneur Behind Virtual, a Multi-Million Dollar AI Agent Society
2025.01.22

10 smart traders specializing in MEMEcoin trading on Solana
2024.12.09

A trader lost $73.9K trading memecoins in just 3 minutes — a lesson for us all!
2024.12.13

What is $SPORE? Let us take you through the on-chain records to show you how it works.
2024.12.25

