Daily Brief
2026-09-17 · Morning · 11 items
-
Jetson AGX Thor runs 27B model through full MLPerf inference solo
A single board hits 52.33 tokens/s and clears all 1,007 rounds in 24m36s, outpacing llama.cpp.NVIDIA Technical Blog:Agentic AI / Generative AI
-
Gemini Enterprise opens private preview of Agent Anomaly Detection
An independent reasoning layer audits agent tool calls and execution asynchronously, adding zero request latency.Google Developers Blog(RSS)
-
OpenAI releases a framework for tracking and disclosing model misalignment, plus six reports
Behaviors are disclosed on a fixed timeline even before being fully explained, prioritizing cases that reveal new misalignment mechanisms.X:OpenAI (@OpenAI)
-
Anthropic and OpenAI pledge on-site third-party safety audits
Both frontier labs agreed to give independent evaluators like METR system access.TechCrunch:AI(RSS)
-
Anthropic merges Claude Cowork and chat into one Claude
No more picking an entry point: Cowork and Design work in any conversation.Claude:Blog(网页)
-
SemiAnalysis tests Rubin NVL72: 7x perf/W over GB300
Jensen promised only 3x at GTC 2026; SemiAnalysis measured more than double that.X:SemiAnalysis (@SemiAnalysis_)
-
Mozilla 91-page report: open-weight models trail frontier by about 4 months
Open-weight models handle 20% of calls yet capture only 4% of revenue, as closed models charge roughly 6x per call.X:Rohan Paul (@rohanpaul_ai)
-
Xiaomi open-sources MiMo-V2.6 RL training: 2B tokens per step on a 1568-prompt rollout grid
A fully async agentic RL run processing ~2B tokens per step across 1568 prompts × 16 rollouts, with spend already past $1M.X:Elvis Saravia (@omarsar0, DAIR.AI)
-
Xiaomi releases MiMo-V2.6 large-scale RL training details
RL scales across compute, environments, and scoring, processing about 2B tokens per step in a fully async setup.X:Nathan Lambert (@natolambert)
-
Grok Build adds memory that keeps project conventions and decisions across sessions
Memory is split per project with a global preferences layer, and the current turn's instructions override stored notes.xAI:News(网页)
-
Transluce names four high-risk behaviors for embedded monitors
Multi-agent coordination, targeted persuasion, eval-awareness and hidden reasoning are flagged as top priorities alongside four pilot approaches.Transluce(网页)