AI clip generator: clips chosen by meaning, not by a timer
Zenuro is an AI video clipper for long recordings: podcasts, interviews, streams, webinars, lectures, vlogs. You hand it a file or a link, it transcribes the audio, reads the transcript from start to finish, and cuts the parts that hold up on their own. One upload produces vertical shorts in 9:16 and long horizontal clips, already rendered with captions, a title on screen, a description and hashtags.
The interesting question with any clip generator is not whether it can cut video. It is what decides where the cut goes. Everything below is about that decision, and about what a single run actually hands back.
Generate clips free 30 min/month →How it works, step by step
A file or a link: podcast, interview, stream, webinar, lecture, vlog. Length does not matter, and you do not have to sit and watch the progress bar. Start the run and come back later.
The audio is transcribed with timing for every word. The transcript is cached by the audio itself, so processing the same recording again does not pay for the same transcription twice.
Not the first ten minutes, not a sample: the whole transcript. It looks for fragments that stand up without the rest of the episode, and it marks where each one starts and where it is finished.
One pass per clip: trim, silence removal, 9:16 crop that follows the speaker, burned in captions, a title in the opening seconds, the look you picked.
A title, a description and hashtags for that fragment, not copied from the source video. On the Pro plan each clip also gets a cover image drawn for it.
Selection by meaning versus cutting by a timer
Plenty of tools will slice a file into equal chunks, and a script can do it in one command. The output is technically a set of clips and practically unpublishable. Six things change when the cut points come from what is being said.
Splitting an hour into forty 90 second pieces gives you forty pieces, and most of them open in the middle of a sentence. Zenuro moves the start to where the thought starts and ends the clip where the thought is finished, so the length comes from the fragment and not from a setting.
The usual failure looks like a clip that begins with “and that is exactly why”. The viewer has no idea why. Fragments that only make sense next to the rest of the episode are dropped instead of shipped.
A two hour interview normally holds somewhere between three and eight self contained topics. You get as many clips as there are, not as many as the arithmetic allows. You can also cap the number yourself.
Trims are aligned to speech, so a clip does not start halfway through a word and does not break off mid phrase. Pauses and dead air inside the fragment are cut out, which is usually where the wasted seconds are.
Vertical shorts bring new viewers in, long horizontal clips are what keeps a channel worth subscribing to. Both come out of the same upload, so you are not paying to process the recording twice.
In 9:16 the crop is placed on people rather than on the centre of the original frame, because the centre of a two person interview is usually the wall between them. When the other person starts talking, the frame cuts to them instead of sliding across the gap.
What one run includes
A clip is finished when it can be uploaded without opening an editor. That is the bar here, so the following comes out of the same run and is not a list of add-ons.
Vertical shorts in 9:16 for TikTok, Reels and YouTube Shorts, and long horizontal clips for YouTube, from the same run. Long clips are on the Pro plan.
Faces are detected, the active speaker is tracked, and the frame changes person with a cut rather than a pan.
Burned into the video, timed per word. Four animations: karaoke highlight, word by word, glow, and a word on a plate. The caption font is yours to choose.
A short hook appears within the first 2.6 seconds of the clip, in one of five plate styles. That is the window in which a viewer decides to keep watching.
Pauses, breathing gaps and dead air are removed inside the clip, audio and video staying in sync.
Noir, film and cinema colour treatments, applied to the whole clip instead of being graded by hand.
The same clip can carry translated captions, which is the cheapest way to reach an audience the original language does not.
Written from the content of that specific fragment. A description copied from the parent episode is the fastest way to make a clip invisible in search.
Generated for each clip on the Pro plan, so a long clip arrives with a thumbnail and not with a random frozen frame.
For recordings where the speech is not the point. Selection runs on what happens in the frame rather than on the transcript. Pro plan.
Before downloading you can trim a clip and fix caption text. Corrections survive re-rendering, so a misheard name is a one line fix and not a reason to start over.
Up to 1080p on the free plan, up to 1440p on Starter, up to 4K on Pro. It is a ceiling, not a promise: a 1080p source stays 1080p.
Price, and the free 30 minutes
30 minutes of processing a month, no card required. Clips carry a small watermark. Enough to run one episode through and compare the selection with what you would have cut by hand.
300 minutes, no watermark, clips up to 1440p. In Russia the same plan is 890 ₽.
1200 minutes, long topic clips, cover images for every clip, vlog and action mode, clips up to 4K. In Russia 2190 ₽.
Minutes are counted on the length of the source you process, not on the number of clips it produces, so a two hour interview costs the same whether it yields four clips or twelve. Payment goes through a Russian card or SBP, an international card, or cryptocurrency. Full breakdown on the pricing page.
Frequently asked
It is a tool that takes one long recording and returns short videos that can be published as they are. Zenuro transcribes the audio, reads the transcript, chooses the fragments worth cutting, renders them with captions and a title, and writes a description and hashtags for each one. The part that a person usually spends hours on is finding the fragments, and that is the part being automated.
Splitting by a timer produces pieces of a fixed length with no idea what is inside them, so most of them begin mid sentence and end mid sentence. Zenuro decides where to cut from the transcript: it starts where a thought starts, ends where it is finished, and skips fragments that only make sense in the context of the full episode.
It depends on how much of the recording is worth cutting. An hour long interview usually yields three to six long clips plus a set of shorts, a four hour stream yields more. You can set an upper limit yourself if you only want a handful.
Shorts are the usual vertical length for social feeds. Long clips are set by the topic rather than by a timer, which in practice means roughly two to fifteen minutes. A clip ends where the thought ends.
The median across our runs is around eight minutes per episode, and three quarters finish within fifteen. It runs in the background: start it, close the tab, and the clips are in your account when you come back.
You decide which clips to publish and you publish them. Zenuro prepares the files, the captions, the titles and the descriptions, and you download them. It does not post to your accounts on your behalf.
Yes. 30 minutes of processing per month, no card needed. Clips made on the free plan carry a small watermark, which paid plans remove. One episode is normally enough to see whether the selection matches what you would have picked by hand.
Comparing tools
If you are already testing a specific product, the feature by feature pages are more useful than this one: Zenuro compared with Opus Clip, Zenuro compared with Vizard, and Zenuro compared with Klap. Prices and limits on the pricing page.
Free plan · no card needed · card, SBP or crypto when you upgrade