FreeToken open sourced: Running DeepSeek-V4-Flash on a single RTX 5090 delivers a maximum speed of 25 tokens per second.
1 hours ago
Beating AI News: Researchers from institutions including UC Berkeley and MIT have open-sourced FreeToken, a Mixture of Experts (MoE) inference engine built specifically for local devices, designed to solve the problem of oversized models failing to fit in video memory (VRAM). FreeToken stores most of a model’s expert weights in system memory, caching only a subset in VRAM. When an expert not present in VRAM is required, it dynamically decides whether to transfer the weights to the GPU or offload computation directly to the CPU based on PCIe and CPU memory bandwidth, enabling collaborative inference across CPU, GPU, system memory and PCIe. According to paper benchmarks, a gaming desktop equipped with an RTX 5090 runs the 284B-parameter DeepSeek-V4-Flash at 22–25 tokens per second. The model only activates approximately 13B parameters per generated token, making it particularly well-suited for this CPU-GPU hybrid inference approach. Additional test results show FreeToken runs a 35B model at 39.3 tokens per second on an RTX 4060 laptop with 8GB VRAM; a single 96GB RTX PRO 6000 can run the 753B-parameter GLM-5.2 at roughly twice the speed of llama.cpp. The project currently supports over 20 MoE models and is open-sourced under the Apache 2.0 license.
Binance will list MARSCOINUSDT perpetual contracts with up to 20x leverage.
4 minutes ago
Runway has launched an App World Model that enables users to generate full app interfaces directly without any coding.
4 minutes ago
Hong Kong stocks closed, with the Hang Seng Tech Index down nearly 1.5% and MINIMAX falling nearly 4.2%.
4 minutes ago
Analysis: In August, only large wallets added approximately 60,000 BTC, while small holders continued to reduce their holdings.
4 minutes ago
Bitcoin fell below $78,000, with a daily decline of 0.77%.
4 minutes ago
Upbit has designated Injective (INJ) as a trading alert asset.
4 minutes ago
Hot feeds
A trader profits $448K by monitoring #Binance's new listings!
2024.12.13 17:37:29
Last week, funds have flowed into #Bitcoin, #Ethereum, and #Hyperliquid.
2024.12.16 14:48:36
A $PEPE whale that had been dormant for 600 days transferred all 2.1T $PEPE($52M) to a new address.
2024.12.14 10:35:27
When Elon Musk tweeted about Moltbook, the meme coin MOLT experienced a short-term 30% price surge, hitting a new all-time high of $114 million.
2026.01.31 18:37:29
A smart #AI coin trader made $17.6M on $GOAT, $ai16z, $Fartcoin,$arc.
2025.01.05 16:05:18
A sniper earned 2,277 $ETH ($8.3M) trading $SHIRO within 18 hours!
2024.12.03 23:09:08
MoreHot Articles

How did I turn $1,000 into $30,000 with smart money?
2024.12.09

10 promising AI Agent cryptos
2024.12.05

The 30-Year-Old Entrepreneur Behind Virtual, a Multi-Million Dollar AI Agent Society
2025.01.22

10 smart traders specializing in MEMEcoin trading on Solana
2024.12.09

A trader lost $73.9K trading memecoins in just 3 minutes — a lesson for us all!
2024.12.13

What is $SPORE? Let us take you through the on-chain records to show you how it works.
2024.12.25

