Remove filler words
and dead air, in one line.
Say "cut the ums and the long pauses." Blob finds them across the whole recording and removes them, keeping the breaths that make speech sound human.
- 01
Upload the recording
A talking head, a podcast, a screen recording, a webinar, anything with speech.
- 02
Ask for the cut
"Remove the filler words and any pause longer than a second." Blob transcribes, finds every instance, and cuts.
- 03
Tune it
Too aggressive? "Leave the pauses at the end of sentences." Too loose? "Tighter." Each pass is a sentence.
Why this one edit matters more than it sounds
Removing filler is the single highest-leverage cut in most talking-head video, and it is also the most tedious to do by hand. Unscripted recordings carry far more hesitation and dead air than it feels like while you are talking, and on a long take it adds up to minutes. Doing that manually means scrubbing a waveform, finding each gap, and rippling the timeline, hundreds of small operations for one video.
It also compounds. Every later decision gets easier once the pacing is right: captions have less to cover, music has a steadier bed to sit under, and the length you are cutting toward is suddenly within reach. It is the reason experienced editors do it first.
The reason to automate it is not that it is hard. It is that it is the least creative work in the entire process.
The difference between cutting silence and cutting well
Plenty of tools will delete every gap above a threshold. The result usually sounds wrong, and it is worth understanding why.
Speech has two kinds of pause. There is dead air (the gap while someone thinks, checks notes, or loses their thread. And there is structural pause) the beat at the end of a sentence, the breath before an important point, the silence that gives a line weight. Strip both and you get a recording that feels breathless and slightly panicked, like it has been sped up.
Asking in plain language lets you draw the line where you want it:
- "Cut anything over a second, but leave the pauses between sections."
- "Remove the ums but keep the pacing natural."
- "Tighten it, but don't make me sound rushed."
That is a distinction a threshold slider cannot express.
What counts as filler
Beyond the obvious *um* and *uh*, most recordings carry a longer tail:
- Discourse markers: "so", "like", "you know", "I mean", "basically", "right?"
- False starts: the half-sentence you abandoned and restarted.
- Repeats, saying the same clause twice while you find the words.
- Throat-clearing, literal, and also the "so, yeah, anyway" that bridges nothing.
You can be specific about which of these to remove. Some are worth keeping: a presenter with a distinctive rhythm can sound sterile with every marker stripped out. Ask for what you want ("cut the ums and false starts, leave the rest"), rather than accepting one global setting.
Then keep going
A tightened cut is a starting point rather than a finished video. The natural next instructions, in roughly the order that works:
- 01Set the format: "make this vertical for Reels" or "keep it 16:9".
- 02Add captions: most social video is watched muted.
- 03Add music that sits under the voice rather than fighting it.
- 04Fill the gaps with b-roll where the talking head gets visually static.
- 05Cut to length: "get this under two minutes."
Each of those is one sentence. The full walkthrough covers the pacing pass in more depth.
Questions
- Will the cut sound choppy?
- Not by default. Blob keeps the short structural pauses that make speech sound natural and removes the dead air around them. If a cut lands wrong, say so and it adjusts: "leave more room after each sentence" works.
- Does it work on podcasts and long recordings?
- Yes. Long unscripted recordings are where it saves the most time, because the filler is spread across the whole runtime rather than concentrated in a few places.
- Can I remove specific words rather than all filler?
- Yes. Ask for exactly what you want removed ("cut every ‘you know’" or "remove the false starts but keep the ums"), and it applies only that.
- What about multiple speakers?
- Yes. Interviews and conversations work the same way, and you can ask for different treatment for different people, for example "tighten the host but leave the guest alone." Give each person's sections a quick listen afterwards.
- Can I undo it?
- Yes. Each finished render is kept as a version you can go back to, and you can also just say "put that pause back."
Tighten your first recording
Upload a clip, ask for the cut, and watch the dead air disappear. Free to start.
Start free