Polish video without an editing pipeline: remove silences and filler words, reframe for any platform, and style captions. Most of it runs with no AI at all.
Cuts are computed from the caption cues, not from a model listening to the audio. A gap of 800ms or more between cues is a silence. A cue whose every word is a filler term is a filler cut. Overlapping cuts merge.
Because it comes from timing rather than judgement, the same recording produces the same cuts every time, and you can see exactly why each one exists.
16:9 for a normal share, 9:16 for vertical, 1:1 and 4:5 for feeds. Intro and outro fades are available on each. The render is a separate output, so one recording can leave in several shapes.
Stitching several recordings into one, and drawing overlays on the video itself, do not exist. Trimming, cuts, reframing, and fades do. Where the rest sits is on the roadmap.
No. Cuts are computed from caption cue timing and a fixed filler-word list, so the same recording produces the same cuts every time. Nothing is sent to a model to decide what to remove.
Silence cuts and reframing are non-destructive: they apply on playback or at render time, and your original file is untouched. Trimming does produce a rendered MP4, so trims carry into downloads and embeds.