1.5TB reduced to 214GB: Tencent releases extreme quantized version of Hy4 preview
4 days ago
Beating AI News: Just after the Hy4 preview went open-source, Tencent has released an extreme quantized version of its Hunyuan model. The original model weights are nearly 1.5TB, while the new GGUF version is only around 214GB, drastically lowering the local deployment barrier for this 770B MoE model. Tencent did not uniformly quantize the entire model to 1.25-bit; instead, it applied different quantization levels based on each layer’s sensitivity to precision: non-critical layers are compressed to as low as ~1.31-bit, while sensitive layers retain 2-bit or higher precision, resulting in an average of ~2.38 bits per weight (bpw). In four benchmarks provided by Tencent, the quantized version only dropped 0.2 to 1.6 points compared to the BF16 original. After compression, Tencent also tested heterogeneous device joint inference with prima.cpp. A setup consisting of an RTX 4090 laptop and a 4-A4000 server, with only 80GB of total VRAM and 64GB of RAM, achieved an inference speed of 1.02 tokens per second—roughly 6 times faster than running the model offloaded on the laptop alone. Multiple devices with different configurations can also jointly share the model inference workload.
Trump renews pressure on the Federal Reserve to cut interest rates, threatening to cut off trade with trade deficit countries if it fails to do so.
5 hours ago
AI cloud computing firm Nscale plans to raise $3.5 billion in pre-IPO financing.
5 hours ago
First Week Review After Apple's Leadership Change: iPhone 18 Pro Color Scheme Revealed, Foldable iPhone Ultra May Face Supply Constraints, MacBook Ultra Expected to Debut
5 hours ago
U.S. National Sheriffs' Association withdraws its opposition to the CLARITY Act, shifting to a neutral stance.
5 hours ago
Viewpoint: Bitcoin's rally recaptures market focus as companies accelerate accumulation of BTC and ETH
5 hours ago
Uniswap co-founder: AMC CEO attempted overreach in law enforcement, and stock tokenization was carefully structured in a legal manner.
5 hours ago
Hot feeds
A trader profits $448K by monitoring #Binance's new listings!
2024.12.13 17:37:29
Last week, funds have flowed into #Bitcoin, #Ethereum, and #Hyperliquid.
2024.12.16 14:48:36
A $PEPE whale that had been dormant for 600 days transferred all 2.1T $PEPE($52M) to a new address.
2024.12.14 10:35:27
When Elon Musk tweeted about Moltbook, the meme coin MOLT experienced a short-term 30% price surge, hitting a new all-time high of $114 million.
2026.01.31 18:37:29
A smart #AI coin trader made $17.6M on $GOAT, $ai16z, $Fartcoin,$arc.
2025.01.05 16:05:18
A sniper earned 2,277 $ETH ($8.3M) trading $SHIRO within 18 hours!
2024.12.03 23:09:08
MoreHot Articles

How did I turn $1,000 into $30,000 with smart money?
2024.12.09

10 promising AI Agent cryptos
2024.12.05

The 30-Year-Old Entrepreneur Behind Virtual, a Multi-Million Dollar AI Agent Society
2025.01.22

10 smart traders specializing in MEMEcoin trading on Solana
2024.12.09

A trader lost $73.9K trading memecoins in just 3 minutes — a lesson for us all!
2024.12.13

What is $SPORE? Let us take you through the on-chain records to show you how it works.
2024.12.25

