#multimodal
ByteDance Releases SeedRealtime: Simultaneous Audio, Video, and Speech Interaction in Doudoubao
ByteDance's Seed division releases the native full-duplex large model SeedRealtime: a unified architecture integrating audio, video, and text, enabling simultaneous listening, viewing, and speaking, and discerning when to speak in noisy environments. It has been fully launched in the Doudoubao app, and users can experience it directly through the video call entry.
MiniMax H3: A Single Model for Video and Stereo Audio Generation at 2K Resolution for 0.8 Yuan/Second, Now Open Source
MiniMax releases its first open-source multimodal generation model H3: a single model that unifies understanding of text, images, videos, and audio, outputting 15-second videos at 2K resolution with native stereo sound. Priced at 0.8 yuan/second, less than one-third of similar flagship products. Artificial Analysis ranks its video editing capabilities as number one globally. We break down the technology and pricing to help you understand how it can save you money.