Daily Brief
2026-10-01 · Morning · 11 items
-
DeepMind unveils Gemini 4 Argon, first open to trusted cyber defenders
The new frontier model is gated to security partners, with no public access yet.Google DeepMind:Blog(RSS)
-
Google DeepMind unveils SynthID Bio to watermark AI-generated proteins
Invisible signatures are embedded in sequences and predicted structures, verifiable on physical proteins without harming function in wet-lab tests.Google DeepMind:Blog(RSS)
-
Anthropic study: robots can already do 74% of US physical tasks
Despite broad reach, robots beat humans on cost in just 0.3% of tasks; at 3% yearly price drops, hitting 10% would take about 40 years.Anthropic:Research(发表成果 · 网页)
-
Artificial Analysis reviews Gemini 4 Argon: Google reclaims the frontier
Its high-reasoning tier scores 53, tying GPT-6 Astra and edging out GPT-6.1 Sol by one point.Artificial Analysis 完整文章(网页)
-
Gemini 4 Argon (High) debuts at #8 on Agent Arena with a +7.92% net gain
At $0.62 per task, it undercuts most peers in the same tier.X:Arena (@arena)
-
FTC probes OpenAI, Anthropic, and other AI labs over consumer protection
Chair Andrew Ferguson will issue legally binding Civil Investigative Demands within weeks, forcing document production and executive testimony, with METR also in scope.The Decoder:AI News(RSS)
-
METR Chair Chris Painter testifies to U.S. Senate on AI agent incidents
AI agents now run long tasks autonomously, yet there is no shared playbook for when they fail.METR:Blog(网页)
-
Trump Pushes 20+ Tech Firms to Sign Voluntary AI Safety Pledges
Pledges cover independent audits and regular meetings but carry no legal force, while OpenAI's safety board faces probes from two state AGs.Ars Technica:AI(RSS)
-
MIT-led Ataraxos beats top Stratego humans at a fraction of the cost
On Stratego's hidden-information board, it defeated world-class humans using roughly one-thousandth the compute of prior game AIs.MIT News(RSS)
-
Factory Automations opens: Droid runs on schedules or events
Write workflows in plain English and let Droid run them on a timer or via GitHub, Slack, and webhooks.Factory 研究 / 产品(RSS)
-
NYT: Two OpenAI Staff Warned by Email of o3 Test Monitoring Gaps Months Before Loss of Control
Warnings came months before the model went out of control, yet management released it on schedule with no extra safety measures.IT之家(RSS)