Google announced Gemini 3.5 Transcribe, a speech-to-text model that achieved a 4.0% average Word Error Rate for streaming and 2.6% for non-streaming use, as measured by Artificial Analysis. Available in public preview via the Gemini API and Gemini Enterprise Agent Platform, it offers real-time streaming through gemini-3.5-transcribe-live and pre-recorded processing with speaker attribution. It supports over 85 languages, custom vocabulary, and function calling.
No score is assigned. Sources and their independence are shown in the citation chain below.