
Turning a long interview into shorts is a decision about which spoken answers deserve to survive — and very little of that decision happens on a timeline. Interviews are where much of the paid short-form work actually sits: a creator, a brand or an agency hands you two hours of conversation and wants it cut into self-contained vertical clips. The fastest way to do that well is to read the transcript, choose the complete answers worth keeping, and only then cut. This article gives you that reading method, the honest line between the answers you choose and the cutting a machine can finish, and the framing rules that keep a talking head watchable at 9:16.
An interview becomes shorts by choosing which answers survive
An interview project is rarely one clip. Whether the source is a podcast conversation, a guest interview or a talking-head session, the brief usually reads the same: cut this into shorts that can be posted on their own. That single difference changes the whole job. A long-form edit shapes hours of footage into one continuous story; shorts turn the same footage into several independent videos, and each one must make sense to someone who never saw the other clips and never heard the question behind the answer.
Because of that, the bottleneck of interview clipping is not cutting — it is choosing. For every clip you deliver, the real work is deciding which spoken answer is worth keeping and where that answer starts and ends. Read a long conversation as a series of separate answers and every answer becomes a candidate clip; the craft is picking the answers that stand alone and marking clean boundaries so each short carries a complete thought. That choice happens in the words long before it happens in the frames, which is why everything below starts from the transcript.
The drawing shows the whole method in one glance: take the interview, read it as text, pick the complete answers, then cut, reframe and caption each answer into its own vertical short.
Scrubbing the footage and blunt auto-clipping both fall short
The two 'fast' routes to interview clips fail for opposite reasons, and both fail in the same place — neither builds a map of what was actually said.
Watching and scrubbing is the honest slow way, and it is slow because you are watching instead of deciding. On a two-hour interview you sit through the whole thing once and still have no index: when the client later asks for 'the part where she talks about onboarding', you cannot jump there. You watched it, but you did not read it. The time goes into the screening, not into the selection.
The blunt shortcut — feed the whole interview to an auto-clipper and take whatever comes out — collapses the other way. It can find loud moments, laughter or pauses, but it does not know when a spoken thought is complete, so it hands back fragments that start or end mid-sentence. That is the fastest way to lose a viewer: audience-retention data consistently shows the steepest drop in the opening seconds, and a clip that opens on the tail of an idea is starting mid-thought for everyone who sees it.
Both shortcuts dodge the one step that makes interview clipping fast: reading. The transcript is the map that watching never gives you.
Read the interview and mark complete answers first
Open the transcript and treat it as the source of truth. Editors who cut interviews professionally reach for transcription for exactly this reason — it turns hours of talk into something you can read and search in minutes. PremiumBeat's guide to cutting interview videos walks through transcription and highlighting as the foundation of the whole edit for the same reason. Work from a timecoded transcript, so every answer you mark has a frame to cut at.
Read it looking for complete answers, not for 'interesting moments'. An answer is a thought with a beginning, a middle and an end: the person states what they think or did, explains it, and lands. Mark where a self-contained idea starts and where it finishes. That boundary — not the loudest ten seconds — is your clip.
Then run each candidate through the stranger test. Would it make sense to someone who did not hear the question and did not see the previous minute? If it needs the question restated, or it trails off without landing, either start it earlier, end it later, or drop it. A clip that begins in the middle of an answer and stops before the point is not a short — it is a fragment that only the person who asked the question can follow. The way a clip ends matters as much as the way it starts; ending on a complete thought is what lets it finish instead of just stopping.
Finally, hold to one idea per short. If one answer contains three strong ideas, make it a strong short or split it into three; do not splice unrelated thoughts into a single clip just to fill time. The drawing shows the difference between shipping a fragment and shipping the whole answer.
Frame one speaker: reframing and captions are the repetitive pass
Once the answers are chosen and bounded, each clip still has to become watchable at 9:16, and that work is identical for every clip in the batch: frame whoever is speaking, keep the face clear, and keep captions above the interface.
- Frame the speaker, not the room. Interview footage is usually a wide two-shot. In vertical you cannot show both people at once, so the crop follows who is talking — cut to the active speaker and hold there. (The anatomy drawing below shows the crop over the two-shot and the resulting frame.)
- Keep the face in the upper safe area. Vertical platforms crop a little from the centre and sides of the feed, and the phone UI covers the bottom; the safest composition puts the eyes in the upper third. Each platform's UI also crops the frame differently, so leave margin rather than filling edge to edge.
- Keep captions clear of the bottom. The like bar, the caption and the feed sit over the lower part of the clip; where you place captions on a short decides whether they stay readable. Put them above the zone the interface occupies.
Vertical video was long treated as a framing mistake; the history of the format on Wikipedia is a useful reminder that portrait is a deliberate choice the phone made standard, not a compromise you fight. The point here is narrower: for a talking head the vertical frame is a portrait of one person, and the crop is decided by who speaks.
This pass is where ClipFinish's clip production line earns its place, because it is the part that truly repeats. You read the transcript, mark the answers — the judgement you just did — then set the framing and caption style once, and the tool cuts each marked answer, reframes it to the speaker and captions it in batch. ClipFinish does not decide which answer is worth keeping or where its thought ends; it finishes the mechanical reframing and captioning of the answers you chose. That split — automate the repetitive pass, keep the judgement — is the deeper argument about AI clipping tools.
The anatomy of the crop, and why the caption sits where it does, is in the drawing.
Interview clipping: quick answers
- Which parts of an interview should I clip?
- Read the transcript for complete, self-contained answers — a thought with a start, an explanation and an ending — and mark where each one begins and ends. Prefer answers that make sense to a stranger who never heard the question.
- Should I work from the video or from the transcript?
- Start from a timecoded transcript. It lets you find and bound answers in minutes and search for a specific topic, then watch only the ranges you marked. Watching hours of footage without a text map is the slow way.
- How long should a clip from an interview be?
- As long as the complete answer takes. A self-contained thought that runs forty seconds is a better short than a chopped ten-second fragment. Never shorten an answer into a half-thought just to fit an arbitrary length.
- Can an AI clipper handle interviews for me?
- Partly. A tool can cut, reframe and caption at scale, but it does not know when a spoken answer is complete, and a clip that starts or ends mid-thought loses the viewer. Keep the tool on the repetitive finishing pass and keep the choice of answers to yourself.
- Which speaker do I frame when two people are visible?
- The one talking. Interview footage is usually a wide two-shot, and in vertical you crop to whoever is speaking and hold there.
You turn an interview into shorts by reading it for complete answers and choosing the ones that stand alone — not by cutting harder or faster. Decide each clip's answer and its boundaries in the transcript, frame one speaker at 9:16 and keep captions above the interface; that choice is what makes the clip yours. The mechanical pass that follows — cutting each marked answer, reframing it to the speaker and captioning the batch — is exactly what ClipFinish's clip production line finishes for you, while the decision of which answer is worth keeping stays in your hands. That is the part of an interview that only you can read.