Anthropic trains an AI to detect AI: Without providing legitimate cases, its false positive rate surged to 97%.
58 minutes ago
Beating AI Express: Anthropic researchers are attempting to train an agent dedicated to auditing other AI models. They first secretly implant hidden problematic behaviors—including bias, deception, and fake citations—into target models, then let Haiku 4.5 conduct investigations independently. Haiku has no prior knowledge of the correct answers, so it must adjust its questions, modify scenarios, and run repeated tests within a maximum of 15 rounds. During training, researchers do not directly tell Haiku whether it correctly detected anomalies. After each investigation, the system compares Haiku’s results against an investigation record from an older Haiku version, then has Sonnet 4—who already knows the correct answers—evaluate which investigation is more accurate. The old record is updated every 50 training steps, effectively letting Haiku continuously challenge its own past performance. However, if all training tasks involve problematic models, Haiku will quickly develop flawed patterns. To secure rewards, it starts assuming all models are problematic; in some setups, its false positive rate against normal models reaches nearly 97%. Later, researchers swapped half the training tasks to use completely normal models, requiring Haiku to correctly identify a target model as normal after investigation to earn high rewards. The trained Haiku indeed changed its approach: it proactively adjusts scenarios, controls variables, and compares responses across iterations, rather than immediately pushing models to admit faults. Its final comprehensive audit score rose from 44.2 to 48.7, nearly matching Opus 4.6’s 48.4. When tested on the more challenging AuditBench—where models actively hide anomalies—its detection rate improved from 11.5% to a maximum of 28.1%.
Trader "Maji" added nearly $5 million in positions in the past hour, currently ranking first on Hyperliquid's daily profit leaderboard.
5 minutes ago
Fed's Schmid: Midterm elections will not affect the Federal Reserve's October meeting decision.
5 minutes ago
Optical module and storage stocks rise broadly in US pre-market trading.
5 minutes ago
Coincheck has completed registration as a Japanese electronic payment service provider, and will enter the stablecoin and on-chain finance sectors.
5 minutes ago
CryptoQuant Founder: Peak of This Bull Market Cycle May Be Driven by Institutional Capital and ETFs Outside the US
5 minutes ago
Arga Labs secures $10 million in funding to build AI agent clones of Slack and GitHub
5 minutes ago
Hot feeds
A trader profits $448K by monitoring #Binance's new listings!
2024.12.13 17:37:29
Last week, funds have flowed into #Bitcoin, #Ethereum, and #Hyperliquid.
2024.12.16 14:48:36
A $PEPE whale that had been dormant for 600 days transferred all 2.1T $PEPE($52M) to a new address.
2024.12.14 10:35:27
When Elon Musk tweeted about Moltbook, the meme coin MOLT experienced a short-term 30% price surge, hitting a new all-time high of $114 million.
2026.01.31 18:37:29
A smart #AI coin trader made $17.6M on $GOAT, $ai16z, $Fartcoin,$arc.
2025.01.05 16:05:18
A sniper earned 2,277 $ETH ($8.3M) trading $SHIRO within 18 hours!
2024.12.03 23:09:08
MoreHot Articles

How did I turn $1,000 into $30,000 with smart money?
2024.12.09

10 promising AI Agent cryptos
2024.12.05

The 30-Year-Old Entrepreneur Behind Virtual, a Multi-Million Dollar AI Agent Society
2025.01.22

10 smart traders specializing in MEMEcoin trading on Solana
2024.12.09

A trader lost $73.9K trading memecoins in just 3 minutes — a lesson for us all!
2024.12.13

What is $SPORE? Let us take you through the on-chain records to show you how it works.
2024.12.25

