HN
Today

Gemini-3.5-Transcribe

Google unveiled Gemini 3.5 Transcribe, its newest speech-to-text model, promising intelligent real-time transcription with superior accuracy and features like disfluency cleanup. The announcement sparked developer interest due to its API availability and broad integration across Google's ecosystem, but also led to Hacker News users debating its real-world performance and comparing it against established open-source and commercial alternatives. The perennial question of Google's feature rollout consistency and the quality of its "smart" optimizations was, of course, also hotly discussed.

132
Score
33
Comments
#5
Highest Rank
10h
on Front Page
First Seen
Aug 27, 8:00 PM
Last Seen
Aug 28, 5:00 AM
Rank Over Time
231366566677

The Lowdown

Google has launched Gemini 3.5 Transcribe, its latest and most precise speech-to-text model, designed for intelligent voice interactions and real-time transcription. The model aims to overcome challenges like background noise, complex jargon, and disfluency by converting raw audio directly into accurate, polished, and formatted text.

  • Dual API Access: It offers a "Real-time streaming" API for interactive voice applications with sub-second latency and a "Pre-recorded audio processing" API for transcribing longer audio with speaker attribution and word-level timestamps.
  • Intelligent Features: Key capabilities include smart transcription (handling self-corrections, removing filler words, auto-formatting), function calling to delegate tasks to other Gemini models, and custom vocabulary recognition.
  • Performance Improvements: Google claims significant advancements over its previous Chirp 3 model, with an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming, alongside a 70% improvement in time to final transcription.
  • Broad Integration: Gemini 3.5 Transcribe is integrated into Google products like Gboard on Android ("Rambler" feature), Google Antigravity, the Gemini app on macOS, and is coming soon to Chrome for talk-to-type functionality. It's also available to developers via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.
  • Global Support: The model supports automatic detection and transcription of over 85 languages, accommodating regional accents and dialects.

This release positions Gemini 3.5 Transcribe as a pivotal tool for enhancing voice-driven interfaces and improving transcription accuracy across various applications, both within Google's ecosystem and for third-party developers.

The Gossip

Rollout Realities and Feature Frustrations

Commenters frequently questioned the immediate availability of Gemini 3.5 Transcribe, particularly for the "Rambler" feature on Android, noting Google's typical staggered and device-specific rollouts. Many expressed anticipation but also skepticism regarding the practical implementation of its "smart" features, with some users reporting that the model's text simplification could unintentionally alter intended meaning, leading to more editing work rather than less.

Competitive Comparisons & Hallucination Hazards

The discussion heavily featured comparisons to existing speech-to-text solutions like OpenAI's Whisper, ElevenLabs, and niche models like Voxtral Mini 3b and Parakeet. A significant concern raised was the potential for "hallucinations"—generating incorrect or nonsensical text from silence or noise—a problem some users experienced with previous models like Chirp and even Whisper. While some found Gemini 3.5 Transcribe to be more resilient in this regard, others debated its cost-effectiveness and overall performance relative to established alternatives, especially for highly specialized or multilingual content.

Google's Grasp on Self-Promotion

A subplot emerged when the original poster (OP), also identified as the founder of a startup, used the comment section to enthusiastically promote their company's view on the new model's market impact. This move quickly drew criticism from other Hacker News users who perceived it as inappropriate self-advertising, sparking a mini-debate on the platform's unwritten rules for promotion versus genuine discussion.