Lookonchain APP

App Store

Laminar reduces agent debugging costs to just 1/23 of GPT-6-sol’s

1 hours ago

Beating AI Express: AI Agent observability platform Laminar has launched flow-1, a model purpose-built to inspect Agent execution traces. The model reads model calls, tool calls, and return results to pinpoint where an Agent errs and the root cause of the mistake. flow-1 operates on Laminar’s Signals agent, capable of searching the full Agent execution trace and drilling down into specific steps as needed. For training, Laminar first used synthetic survey data for supervised fine-tuning, then further trained the model on tool calls and complex trace analysis via reinforcement learning. In Laminar’s self-built benchmark of 523 challenging traces, flow-1 posted an error detection F1 score of 0.835, outperforming GPT-6-sol’s 0.816. GPT-6-sol has a higher recall rate, catching more actual errors, while flow-1 delivers higher precision with fewer false positives. For traces under 100,000 LLM tokens, flow-1 costs an average of ~$0.0011 per analysis, compared to ~$0.026 for GPT-6-sol. Per Laminar’s calculations, $1 can analyze roughly 888 traces with flow-1 versus just 38 traces with GPT-6-sol, a cost difference of around 23 times. flow-1 is also approximately 25% cheaper than GPT-6-luna. All current results are from Laminar’s in-house benchmark, with no third-party verification available. flow-1’s training data is primarily synthetic workflows, around 48% of which are software engineering tasks. Some community members have questioned whether the model can maintain its performance when facing novel, unforeseen failures in real production environments.

Relevant content

US government transfers 264.863 BTC seized in Bitfinex hack case to a new address

According to on-chain monitoring, the US government has just transferred funds seized from the Bitfinex hack, with roughly 264.863 BTC (valued at approximately $22.87 million) moved to a new address.

9 minutes ago

The U.S. government transferred 264.863 BTC seized in the Bitfinex hack case to a new address.

According to on-chain monitoring, the U.S. government has just transferred funds seized from the Bitfinex hack, with approximately 264.863 BTC (worth around $22.87 million) moved to a new address.

9 minutes ago

Winklevoss has filed a listing application for a Zcash (ZEC) spot ETF with the U.S. Securities and Exchange Commission (SEC), under the ticker symbol WINK.

Winklevoss has filed a listing application for a ZEC spot ETF with the U.S. SEC, under the ticker symbol WINK.

9 minutes ago

Bitcoin ETFs see $91.72M outflow, Ethereum ETFs down $215.37M over 7 days

Oct 6 Update: #Bitcoin ETFs: 1D NetFlow: -1,059 $BTC(-$91.72M)?? 7D NetFlow: +1,074 $BTC(+$93M)?? #Ethereum ETFs: 1D NetFlow: -21,432 $ETH(-$58.29M)?? 7D NetFlow: -79,193 $ETH(-$215.37M)??

9 minutes ago

DeFi Development has been authorized to launch the CHAD preferred stock repurchase plan.

Nasdaq-listed DeFi Development Corp. has authorized the launch of a CHAD preferred stock repurchase program, which allows for the repurchase of all outstanding CHAD preferred shares, including any future shares the company may issue. The program enables the company to conduct repurchases flexibly when CHAD’s share price falls below its $10 par value per share. The Solana treasury-focused company stated it has no immediate plans to repurchase CHAD, noting it first wants the security to build market liquidity and stabilize at or near its $10 par value. This indefinite authorization adds a potential capital allocation tool to DeFi Development (DFDV)’s existing $300 million CHAD ATM issuance program. The company previously said it plans to raise funds via this program to purchase additional SOL, and paid CHAD’s first dividend on October 1.

9 minutes ago

Banbury Road trains 32 models together, with the models independently developing distinct areas of expertise.

Beating AI Express News: AI lab Banbury Road has released Kardashev-0.7, a system composed of 32 distinct models. The team trained these models together via reinforcement learning, without predefining each model’s specific tasks, enabling them to develop unique specialties and complementary capabilities during training. This methodology is called RL for Population Scaling (RLPS). Prior RLPS experiments only reached up to 16 models. In the 8-model trial, the jointly trained model population scored 81.70%, outperforming 71.04% for 8 independently trained models and 72.65% for the same model sampled 8 times. In the 16-model experiment, 22 questions were answered correctly by only one model each, verifying that different models did learn distinct capabilities. Kardashev-0.7 expands this to 32 models. Banbury Road claims it delivers state-of-the-art performance at 0.007 to 0.02x inference cost and 0.03x memory usage. Its API beta is currently on a waitlist. However, the population scores from the earlier 8 and 16-model trials were derived post-hoc by picking the best result from multiple model outputs using standard answers. Since real-world applications lack standard answers, Banbury Road has not yet disclosed a full solution for the system to automatically select correct responses.

9 minutes ago

Popular tokens

BitcoinEthereumHyperliquidSolanaTRONBNBTetherAaveXRPPepeFartcoinOndoJupiterUniswapBonkPendleEthenaArbitrumAvalancheLidoChainlinkPolygonDogecoinCardano