25 AI models collectively lose to humans: AI still struggles to understand much everyday common sense
1 hours ago
Beating AI News: Scale Labs, the research arm of Scale AI, has partnered with Elorian to launch a new benchmark called Humanity’s Sixth Sense (HSS), designed to test whether AI can infer implicit information from images and videos. The benchmark includes 522 open-ended questions covering 288 images and 234 videos, such as whether two vehicles can pass each other, why a woman suddenly slows down while chasing a bus, and who holds more sway in a given scenario. The research team evaluated 25 multimodal models and had 20 human participants complete the tasks. Humans achieved an accuracy rate of 93.1%, while the top-performing model, GPT-6 Astra, scored only 53.6% even when using maximum inference effort. GPT-6.1 Sol and Claude Opus 5.5 followed with scores of 46.6% and 44.6% respectively, while the median score across all models was just 30.9%. Researchers analyzed 8,573 model failures and found that 94% were linked to missing key clues, misidentifying objects, or failing to infer implicit relationships in visuals. Only around 5% of errors were categorized as logical reasoning mistakes. Social comprehension proved particularly challenging: 21 out of the 25 models performed worst on these types of questions. Video-based tasks were also generally more difficult than image-based ones. Increasing inference effort does not always help. Models consumed an average of around 4,000 inference tokens per question, and for some tasks, longer reasoning times correlated with lower performance. The research team also tested letting agents zoom in, crop, and re-examine visuals, which led to improved results, though they still lagged far behind human performance.
US initial jobless claims for the week ended October 3 came in at 197,000, against an expectation of 200,000.
8 minutes ago
GlobalFoundries and TSMC Sign $2 Billion Silicon Interposer Agreement
8 minutes ago
Google's US shares reversed pre-market losses, now up 1%.
8 minutes ago
Jev tops Ramp Software’s Dark Horse Chart, with its enterprise adoption rate rising by 1 percentage point in one month.
8 minutes ago
STRK breaks above $0.06, rallying 18.1% in 24 hours.
8 minutes ago
Perplexity open-sources a new retrieval model: its 0.6B small model can directly query indexes built by the 9B model.
8 minutes ago
Hot feeds
A trader profits $448K by monitoring #Binance's new listings!
2024.12.13 17:37:29
Last week, funds have flowed into #Bitcoin, #Ethereum, and #Hyperliquid.
2024.12.16 14:48:36
A $PEPE whale that had been dormant for 600 days transferred all 2.1T $PEPE($52M) to a new address.
2024.12.14 10:35:27
When Elon Musk tweeted about Moltbook, the meme coin MOLT experienced a short-term 30% price surge, hitting a new all-time high of $114 million.
2026.01.31 18:37:29
A smart #AI coin trader made $17.6M on $GOAT, $ai16z, $Fartcoin,$arc.
2025.01.05 16:05:18
A sniper earned 2,277 $ETH ($8.3M) trading $SHIRO within 18 hours!
2024.12.03 23:09:08
MoreHot Articles

How did I turn $1,000 into $30,000 with smart money?
2024.12.09

10 promising AI Agent cryptos
2024.12.05

The 30-Year-Old Entrepreneur Behind Virtual, a Multi-Million Dollar AI Agent Society
2025.01.22

10 smart traders specializing in MEMEcoin trading on Solana
2024.12.09

A trader lost $73.9K trading memecoins in just 3 minutes — a lesson for us all!
2024.12.13

What is $SPORE? Let us take you through the on-chain records to show you how it works.
2024.12.25

