Podcast to shorts: one episode in, vertical clips out
An hour of conversation usually holds five or ten moments that work on their own. Finding them is the slow part. You scrub the timeline, mark a spot, realise the setup started ninety seconds earlier, move the start, export, and repeat. Zenuro reads the transcript of the episode instead, picks the passages that stand without the rest of the show around them, and gives back finished vertical shorts with captions already burned in. One upload, one run, no timeline work.
Try it free: 30 min a month →How it works, step by step
Paste a link or drop the file. Podcast, interview, stream, webinar. A two hour episode is fine, length is not a problem.
Zenuro transcribes the audio and reads the whole thing, looking for passages with a beginning and an end: a question asked and answered, a story told, a claim made and argued.
Each clip starts where the thought starts. Passages that only make sense with the previous ten minutes attached are moved back to the setup or dropped. Silence between sentences is cut out.
In 9:16 the crop sits on the person talking and switches when the other one takes over. The switch is a cut, not a slow pan across the empty space between two chairs.
Animated captions are burned in word by word, a hook line appears in the first 2.6 seconds, and every clip gets its own title, description and hashtags written from what is said inside it.
Why this is not slicing by a timer
Plenty of tools will chop an hour into sixty second pieces. That is a fast way to get sixty clips and no views, because a clip is not a duration, it is a complete thought. The difference below is where the cut points come from.
Fixed length slicing cuts on the minute, so half the clips open mid word and end mid argument. Here the boundaries follow the shape of the thought: a clip can run 22 seconds or 80, whatever the moment needs.
The classic failure of a hand made clip is an opening like “and that is exactly why”. The viewer has no idea why. Openings that need the rest of the episode are either pushed back to where the setup begins or left out entirely.
Centre cropping a two person interview gives you the gap between the chairs. Face detection plus lip movement decides who is talking, and the vertical frame sits on that person.
Karaoke highlighting, word by word reveal, glow, a single word on a plate: four ways of animating, twelve fonts for Latin captions. Most short form is watched with the sound off, so the captions carry the audio.
The same transcript pass also yields long horizontal topic clips for YouTube. Shorts bring new viewers in, long clips give them a reason to stay on the channel.
Captions can be translated into 14 languages, so an episode recorded in English reaches viewers who would never have searched for it in English.
What one run gives you
You upload once. Everything on this list comes out of that single pass, so there is no second export for the captions, a third for the thumbnails and a fourth for the long version.
Sized for TikTok, Reels and YouTube Shorts, with the frame tracking the active speaker.
Self contained topic segments from the same episode, for the main YouTube channel, on the Pro plan.
Karaoke, word by word, glow or a word on a plate, in the font you pick. Nothing to import into an editor afterwards.
A short line over the opening of the clip, five styles to choose from. It is generated as its own phrase, not trimmed from the title.
Written per clip from the content of that clip, not copied from the episode page.
Pauses and dead air go away. Optional looks: noir, film, cinema.
A generated thumbnail for every clip, three styles, on the Pro plan.
For footage where the speech is not the point: cuts follow what happens on camera instead of what is said.
Who ends up using it
The episode ships once a week, the feeds want something several times a week. One recording covers both, and the clips carry their own titles so they do not compete with each other in search.
Guests share clips of themselves far more readily than a link to a full episode. A run gives you a handful of moments per guest that are worth sending them.
A single answer taken out of a long session works as a reply to one specific question, which is how people search.
The slow part of clipping is not the render, it is watching the episode twice to find the boundaries. That part is what gets taken off the desk here.
Price and the free 30 minutes
Minutes are counted on the length of what you upload, not on the number of clips that come out. A 45 minute episode costs 45 minutes whether it produces three shorts or twelve.
30 minutes a month, no card required. Clips carry a small watermark.
300 minutes, no watermark, clips up to 1440p.
1200 minutes, long topic clips, AI covers, vlog and action mode, clips up to 4K.
Payment goes through an international card, a Russian card or SBP, or cryptocurrency, which activates immediately and works from any country. Full breakdown on the pricing page.
Questions people ask first
Upload the episode to Zenuro by link or as a file. The audio is transcribed, the transcript is read for passages that work on their own, and each one is rendered as a vertical 9:16 clip with the frame on the speaker, animated captions, a hook line, a title, a description and hashtags. You download the files and post them.
As many as there are moments that hold up without the rest of the episode. An hour of conversation usually gives five to ten. You can set the number before the run if you want fewer and better, or leave it open.
Usually between twenty and ninety seconds. The length comes from the passage itself, not from a preset, so a short exchange stays short instead of being padded out to fill a slot.
Yes, that is the normal case for a podcast. Faces are detected, lip movement decides who is currently speaking, and the vertical frame cuts between them. It is a cut rather than a camera move, because sliding the frame across the table shows the wall between the chairs.
Usually no, the clips come out ready to post. If something needs a fix, each clip opens in an editor where you can trim the ends, correct a word in the captions or change the caption style, and it re-renders with that change.
The median across our jobs is about eight minutes per episode and three quarters finish within fifteen. You do not have to wait at the screen: upload the episode, come back later, the clips are in your account.
Thirty minutes of processing a month, no card needed, with a small watermark on the clips. That is roughly half a typical episode, enough to see whether the moment selection matches your taste before you pay anything.
Comparing tools rather than starting from scratch: see how Zenuro lines up as an Opus Clip alternative and as a Vizard alternative, or read the plans and limits in full.
Free plan · no card needed · crypto accepted