The video follows.
Heyper Video writes out everything that was said and hands you the text. Delete a sentence and it disappears from the video too. The “ums” and the awkward silences are found for you, captions time themselves to every word, and there is a proper timeline underneath for everything else.
- timed to the frame
- Every wordtimed to the frame
- caption styles
- 12caption styles
- shapes to publish in
- 6shapes to publish in
- exports on your machine
- No queueexports on your machine
Free to start · no card required
Watch the cut land.
Try it — this is the real thing, not a picture of one. Click any word to cut it, close a pause, or clear every “um” at once, and watch the timeline underneath shorten by exactly what you took out.
0:21.9
0:21.9
0w · 0p
words · pauses
1
we weren't sure, so we didn't touch it
Nothing was rendered to work that out. Your edit is simply a list of the bits you kept, which is why every cut is instant and nothing is ever thrown away.
Definite filler. Safe to clear all at once.
Might be filler, might not. We flag it and leave the call to you.
A pause long enough to feel like one. Closing it keeps a natural beat.
gives you
The transcript is the timeline
Find the sentence that did not land, select it, press delete. It is gone from the video too — no dragging a playhead around hunting for a moment you already know is bad.
Fillers and dead air, found for you
Every “um”, and every silence long enough to feel awkward, found and listed for you. Go through them one at a time, or clear the lot in a single click.
Captions worth looking at
Because every word is timed, you get captions that light up as they are spoken — not flat subtitles. Twelve styles, eight fonts, positioned however you like and baked into the file.
A real timeline underneath
Layers, trims, splits and transitions for everything the words cannot say. Landscape, vertical, square or portrait — whatever your footage was shot in.
Restyle the voice, keep the timing
Not a re-recording — the new voice keeps the exact same pacing, so every cut you made and every caption still lands in the right place.
Presenters on camera
Hand a script to a ready-made presenter, or build one from a photograph, and drop the finished take straight onto your timeline.
Record straight into it
Camera, screen, or both at once, recorded right here and sitting on the timeline ready to edit the moment you stop.
Export what you previewed
The finished file, with the colour, the captions and the cuts exactly as they looked on screen. No surprises after the wait.
and this is how they look
No mock-ups: what you see here is what gets burned into your video. Same fonts, same sizes, same pop on the word being spoken. Pick one and watch it run.
- Style
- karaoke
- Font
- anton
- Size
- 6.4% of the frame
- Spoken word grows
- 12% bigger
- Words on screen
- 7
- Sits
- 76% down the frame
at any point
Decide it's going to Reels halfway through and the frame changes around your edit. You don't start again.
A crop is a choice
Reframing lets you move the subject inside the new shape, rather than hoping whatever sat in the middle was the important part. Change the frame at any point in the edit and your cuts, captions and colour come with it.
You keep the quality you started with: a vertical cut of a 1080p video is still a full 1080 tall, not a blown-up slice of one.
without leaving
Five steps, and only one of them costs you anything. Everything after the transcript is instant, because the hard work is already done.
Every cut sounds like an edit, not a glitch
Cut exactly on the word and you clip the start of the next one. We leave a hair of breathing room either side of every cut — never enough to let back in what you removed, always enough to sound natural.
It knows filler from speech
“Um” and “uh” never mean anything, so they are always flagged. “You know” is caught as a phrase. “Like” is only flagged when the pauses around it say it was filler — which is the difference between “it was, like, huge” and “I like it”.
It finds dead air, not quiet moments
It looks at the gaps between words rather than listening for quiet, so a softly spoken line or a held breath is never mistaken for a pause. Close one and a short beat stays behind, so nothing sounds spliced.
Change your mind an hour later
Cutting a word does not delete it. Your original is never rewritten, so anything you take out can be put back — at any point, right up until you export.
One transcript runs the whole edit
The same word timings that place your cuts also break the captions into lines, time the highlight on the spoken word, and put the text on screen at the right moment. Nothing can drift out of sync.
Exporting doesn't wait in a queue
Your video is put together on your own computer — MP4 or WebM, at full size or 1080p, 720p and 480p — with the colour, the captions and the cuts exactly as you previewed them.
who publish weekly
Fast enough to be how your video actually gets made each week — not the thing you brace yourself for the night before it's due.
Turn a webinar into a campaign
Cut the ninety-minute recording down to the argument, caption it for silent autoplay, and export the vertical version in the same session.
Sound like the second take
Strip the fillers and the thinking pauses out of a first take. The delivery tightens without a reshoot, and nobody has to learn an NLE.
Screen recordings that stay current
Record the flow in the browser, cut the wrong turns out by deleting the sentence describing them, and re-export when the UI changes.
Publish the episode, then the clips
Edit the long cut here and hand the same transcript to Heyper Clips to find the moments worth posting.
Restyle the voice, keep the edit
A new voice that keeps the exact same pacing, so every cut and caption still lines up afterwards.
One editor, many formats
16:9 for the site, 9:16 for Reels, 1:1 and 4:5 for feeds — reframed from the same edit rather than rebuilt four times.
Transcription, captions and export are built in. Nothing to install.
plainly
- How you edit
- By editing the words, with a full timeline underneath
- Captions
- 12 styles across 8 fonts, timed to every word
- Shapes
- As shot, plus 16:9, 9:16, 1:1, 4:5 and 4:3
- Export
- MP4 or WebM — canvas, 1080p, 720p or 480p
- Recording
- Camera, screen, or both — no software to install
answered
Every single word comes back with its own start and end time. So when you delete a sentence, we know exactly which piece of video to remove — the cut lands on the frame the word began, not somewhere close to it.
It finds them and shows you what it found — every “um” and “uh”, and every silence long enough to feel awkward. Accept them one by one, or clear them all in a click.
Yes. Switch to a vertical shape and move the picture around inside it. Turning a wide video into a tall one throws away most of the frame, so we let you decide what stays rather than guessing.
Yes — they are part of the video file when you export, exactly as they looked on screen. Same fonts, same positions, same movement.
That is the normal case. Upload it or record it in the browser; nothing has to have been generated here.
with
Your next frame is one prompt away
Open the prompt box, pick a model, and see what comes back. No card, no setup, no onboarding call.