OpenAI Accused of Quietly Adjusting GPT-6 Astra Evaluation Data, Some Metrics Make Competitors Appear 'Worse'
59 minutes ago
Beating AI Express: Since OpenAI launched GPT-6 Astra on September 3, multiple model benchmark evaluation metrics have been continuously adjusted. Some changes have boosted Astra’s performance, while scores of some competing models have dropped, sparking external skepticism about AI "benchmark manipulation" and evaluation transparency. Specifically, Astra’s hallucination rate was once lowered from 4.2% to 2%, while GPT-5.6 Sol’s rate fell from 12.2% to 9.4%, before both rates returned to 4.2% and 12.2% respectively. In math evaluations, Anthropic’s Fable 5.1 score also dropped from 87.8% to 78%, and has since rebounded to 83%; GPT-5.6 Sol’s score fell from 83% to 80.5%, then returned to 83%. Additionally, Astra’s score in the ARC-AGI-3 benchmark rose from 98.6% in its pre-release draft to 99.99% on the final page, while its programming evaluation score was also slightly adjusted upward from 57.7% to 57.9%. OpenAI stated that evaluation results are influenced by factors including model version, tool configuration, inference level, and test runs, adding that the adjustments were made to ensure the data more accurately reflects the model’s optimal performance. However, Stanford University researchers argue that frequent re-running of evaluations may involve so-called "benchmaxxing" – the practice of adjusting test conditions to maximize benchmark scores. Industry insiders note that as competition among AI models intensifies, evaluation data has become a key tool for measuring model capabilities and competing for market share, with improving the transparency and reproducibility of benchmark tests drawing growing attention.
Microsoft’s MAI-Image-2.6-Flash variant offers strong cost-performance: 2.8x faster, with 1,000 images priced under $20.
7 minutes ago
GitHub Copilot launches multi-model teaming: HydraFusion cuts costs by up to 67%
7 minutes ago
U.S. spot Bitcoin ETFs saw a net inflow of $174.6 million yesterday.
7 minutes ago
Anthropic’s IPO may be delayed until mid-October, as the firm aims to hit a $2 trillion valuation.
7 minutes ago
Tesla's "Golden Era" Gets Off to a Chilly Start: Cybercab Faces Regulatory Headwinds Right After Launch
7 minutes ago
More than 30 suspected insider addresses took early positions in SLINK, reaping over $4.7 million in quick profits.
7 minutes ago
Hot feeds
A trader profits $448K by monitoring #Binance's new listings!
2024.12.13 17:37:29
Last week, funds have flowed into #Bitcoin, #Ethereum, and #Hyperliquid.
2024.12.16 14:48:36
A $PEPE whale that had been dormant for 600 days transferred all 2.1T $PEPE($52M) to a new address.
2024.12.14 10:35:27
When Elon Musk tweeted about Moltbook, the meme coin MOLT experienced a short-term 30% price surge, hitting a new all-time high of $114 million.
2026.01.31 18:37:29
A smart #AI coin trader made $17.6M on $GOAT, $ai16z, $Fartcoin,$arc.
2025.01.05 16:05:18
A sniper earned 2,277 $ETH ($8.3M) trading $SHIRO within 18 hours!
2024.12.03 23:09:08
MoreHot Articles

How did I turn $1,000 into $30,000 with smart money?
2024.12.09

10 promising AI Agent cryptos
2024.12.05

The 30-Year-Old Entrepreneur Behind Virtual, a Multi-Million Dollar AI Agent Society
2025.01.22

10 smart traders specializing in MEMEcoin trading on Solana
2024.12.09

A trader lost $73.9K trading memecoins in just 3 minutes — a lesson for us all!
2024.12.13

What is $SPORE? Let us take you through the on-chain records to show you how it works.
2024.12.25

