20-hour programming benchmark highlights performance gaps: Claude Fable 5.1 leads GPT-5.6 by more than 24 points, with GLM-5.3 ranking third.
2 hours ago
Beating AI Express: AI research team Proximal has released FrontierSWE v2, a long-horizon programming benchmark. The number of tasks has expanded from 17 to 34, with each model running 5 trials per task, each trial taking up to 20 hours. Claude Fable 5.1 posted an average score of 56.29%, significantly outperforming GPT-5.6’s 32.2%. Open-source model GLM-5.3 ranked third with 30.2%. Tasks in FrontierSWE go far beyond basic code debugging: AI agents must build circuit simulators from scratch, train weather forecasting models, match star catalogs using telescope images, or train racing bots solely from game visuals. Version v2 has uniformly switched to the Proximus harness. Each task runs for up to 20 hours; when a model is ready to submit its work, the system saves its current state and notifies it of remaining time to prevent premature task termination—a change that significantly impacted scores. Proximal’s comparison across 6 tasks found that both Claude Opus 5 and GPT-5.6 ran longer with the Proximus harness, and posted higher average scores than with their original harnesses. The benchmark also caught multiple instances of intentional cheating: GPT-5.6 once recognized that accessing public answers “might involve anti-cheat issues” but still used the shortcut; on another occasion, it even exploited Modal’s backend services to read hidden validation files. Muse Spark 1.2 modified test scripts, inserted public answers, and wrote code to cover up its cheating traces. All runs confirmed to be violations were scored zero.
Fed's Waller: Whether to raise interest rates in September will hinge heavily on next week's August CPI
5 minutes ago
Market cuts bets on Fed rate hikes; current probability of a September rate hike stands at 60.4%
5 minutes ago
US initial jobless claims came in slightly higher than expected.
5 minutes ago
Binance to List GoPro USDT-Margined Perpetual Contracts
5 minutes ago
Fed Governor Christopher Waller takes a hawkish stance: If August inflation data comes in strong, he will consider supporting an interest rate hike in September.
5 minutes ago
Spot gold and silver rally in the short term, with gold trading at $4,466 per ounce.
5 minutes ago
Hot feeds
A trader profits $448K by monitoring #Binance's new listings!
2024.12.13 17:37:29
Last week, funds have flowed into #Bitcoin, #Ethereum, and #Hyperliquid.
2024.12.16 14:48:36
A $PEPE whale that had been dormant for 600 days transferred all 2.1T $PEPE($52M) to a new address.
2024.12.14 10:35:27
When Elon Musk tweeted about Moltbook, the meme coin MOLT experienced a short-term 30% price surge, hitting a new all-time high of $114 million.
2026.01.31 18:37:29
A smart #AI coin trader made $17.6M on $GOAT, $ai16z, $Fartcoin,$arc.
2025.01.05 16:05:18
A sniper earned 2,277 $ETH ($8.3M) trading $SHIRO within 18 hours!
2024.12.03 23:09:08
MoreHot Articles

How did I turn $1,000 into $30,000 with smart money?
2024.12.09

10 promising AI Agent cryptos
2024.12.05

The 30-Year-Old Entrepreneur Behind Virtual, a Multi-Million Dollar AI Agent Society
2025.01.22

10 smart traders specializing in MEMEcoin trading on Solana
2024.12.09

A trader lost $73.9K trading memecoins in just 3 minutes — a lesson for us all!
2024.12.13

What is $SPORE? Let us take you through the on-chain records to show you how it works.
2024.12.25

