Daily Brief
2026-09-24 · Morning · 8 items
-
Google DeepMind rolls out Gemini 3.8 Flash TTS and Flash-Lite
The new TTS models let users design voices from scratch or clone one from a 30-second sample, spanning 100+ languages.Google DeepMind:Blog(RSS)
-
Antigravity SDK adds local-model support for fully offline agents
First up is Gemma 4 26B A4B via Google AI Edge's LiteRT, with full offline agent execution and a recommended 24GB+ memory footprint.Google Developers Blog(RSS)
-
Fireworks launches Ember-1, cutting reasoning tokens by ~40%
Built on Kimi K3, it matches its quality using about 40% fewer tokens.Fireworks AI(网页)
-
Anthropic details a two-week sprint that roughly tripled claude.ai speed
Optimizing four journeys covering 95% of user activity cut the 75th-percentile first-input time from 3.1s to 0.55s.Hacker News:AI 热帖
-
Modal details serving the trillion-token Kimi K2.6 coding model at scale
Per-user latency on a single replica climbed 2.8x while whole-replica throughput rose 5.6x, moving hundreds of billions of tokens per day per service.Modal 官方工程博客(RSS)
-
vLLM ships vllm-metal v0.28.0, bringing V1 support to Apple Silicon
Running on MLX and Metal, the V1 scheduler and paged KV cache now land on Macs.vLLM 官方博客(RSS)
-
Tomer Tunguz: The real AI market is the middle, not the frontier
Frontier-model token share fell from 53% to 45% in two months, exposing price as the decisive factor.Tomer Tunguz 博客(VC 分析)
-
Anthropic engineers share a six-step playbook for AI-driven code modernization
The bottleneck has shifted from writing code to mobilizing the organization, turning multi-year efforts into projects measured in weeks or months.Claude:Blog(网页)