ai · · 2 min read

Google Updates Gemini Audio with AI Models That Filter Speech Fillers

By Alex Mercer

Google Updates Gemini Audio with AI Models That Filter Speech Fillers

How the Filler Removal Works in Practice

Google has rolled out an update to its Gemini Audio transcription service, integrating newer Gemini 3.5 AI models to automatically remove verbal fillers like „umsand ”ahsfrom transcribed text. The feature aims to produce cleaner, more readable transcripts for users relying on speech-to-text tools in professional and educational settings.

The update enhances the platform’s ability to distinguish between meaningful speech and common disfluencies, applying real-time editing during transcription. By leveraging advancements in natural language understanding, the system now identifies and omits these filler words without altering the core content or intent of spoken words. This improves clarity in meeting notes, lecture transcripts, and interview summaries where precision and readability matter.

Can Users Control How Much Editing Occurs?

When a user uploads or records audio through Gemini Audio, the updated AI models analyze speech patterns to detect hesitation markers. Rather than transcribing every utterance verbatim, the system applies probabilistic filtering to exclude non-essential vocalizations while preserving sentence structure and context. Google states the models were trained on diverse speech datasets to maintain accuracy across accents, speaking styles, and languages. Early testing shows a reduction in transcript clutter by up to 30% in informal spoken content, with minimal risk of removing meaningful speech.

Yes, Google has included adjustable settings that let users choose the level of disfluency removal applied to their transcripts. Options range from conservative, which retains most filler words for authenticity, to aggressive, which maximizes readability by removing even subtle hesitations. This flexibility allows professionals such as journalists, researchers, and educators to tailor output based on their needs—whether prioritizing verbatim accuracy or clean presentation. The feature is available now for all Gemini Audio users via the latest platform update.

Does removing „umsand ”ahschange the meaning of what was said? No, the AI is designed to only remove non-essential speech fillers that do not contribute to semantic meaning. The core message, tone, and factual content remain intact in the edited transcript.

Frequently Asked Questions

Is this feature available in all languages supported by Gemini Audio? Currently, the filler removal enhancement is optimized for English, with plans to expand to other languages as the models are further trained on multilingual speech data.

Can I turn off the automatic editing if I want a raw transcript? Yes, users can disable the filler removal function entirely through the transcription settings menu, allowing full verbatim output when needed.

More stories:

Content written by Alex Mercer for techbriefe.com editorial team, AI-assisted.

Share:

Leave a comment