Smart reframing: how Snip decides where to look
Turning a 16:9 source vertical means throwing away two thirds of the width. Seven real extracts, recomposed live in your browser, to see what Snip recognises shot by shot, how the frame follows a face, what it refuses to decide, and how you take a passage back by hand.
Mehdi Seddik17 August 20269 min readSeven shots, seven answers
Every tab is a real extract, carrying the plan the production sidecar computed on it. The vertical on the right is not a video: it is composed in your browser, frame by frame, by the same function that burns the export. Switch to « centre crop » to see the lazy answer on the very same extract.
Two thirds of the picture to throw away
A 16:9 source is 1920 pixels wide, a vertical clip is 1080. Reframing does not shrink the picture: it picks a window inside it, and everything outside disappears. On a three-seat set, that window decides who stays in the clip.
So the work happens in two steps, and the second one is the interesting one: decide what to show, then build a camera move that holds up. Snip cuts the clip into shots, decides a layout on each, and solves a camera path over its whole length.
The frame follows, without shaking
This is the first criterion, and the only one you can see without measuring anything: if the face drifts inside the frame, the crop is wrong even without a single shake.
On the measured clips the frame follows 81 to 87 % of the face travel, with a median gap of 0.011 to 0.074 face heights between the two centres. On a filmed set that is 0.015 face heights, which is 6 pixels on a 400-pixel face.
The 13 to 19 % it does not follow are not a defect, they are the damping: an operator does not render every nod. The chain that produces that compromise has four stages.
- Detect, then stitchFaces are detected frame after frame, then stitched into stable tracks: one track, one person, for the length of the shot. Two fragments that overlap at the same instant are the same head, never two people sitting side by side.
- Smooth the targetA Kalman filter returns a continuous target with a velocity that means something. It bridges short detection gaps, and freezes the frame instead of drifting when a gap lasts.
- Lay a camera pathA linear program turns that target into a piecewise trajectory: held plateaus and clean moves, rather than permanent float.
- Contain, cap, anticipateThe measured face has to stay inside the window, forehead and chin included. Speed is capped at roughly one second to cross half the picture, and the constraint bites a third of a second early so the move spreads out instead of whipping.
Two cases are not moves and must not become one: a speaker change, and a subject jump on the same camera. Letting the solver join the two positions produces a 1 160-pixel displacement spread over four frames, 67 milliseconds: that is neither a clean cut nor a readable pan, it is a whip. Between two speakers an editor cuts, they do not sweep the room. Those instants are therefore treated as cuts, and the two halves are solved separately.
« When in doubt, a simple stable crop always beats a clever wrong one. »
What it recognises, shot by shot
The layout is decided on each shot, never once for the whole clip. The number of framable people decides the shape, and a measured overlay changes the answer.
There is no speaker switch inside a split screen: both people stay visible throughout. The mouth signal is measured and reported in the metadata, it does not decide the crop. Without the audio, nothing lets us verify that the winning face is the one talking.
Keeping the whole shot is a decision
When no layout is certain, the whole shot is kept. That is not an admission: on a landscape, a football pitch, a concert stage, a centred 9:16 throws away two thirds of the width to frame whatever sat in the middle. The director's frame beats a crop that pretends to be a decision.
That leaves the two free bands, a third of the output. They were black, then tinted with a halo. Today they are the scene itself, cropped to fill the vertical, blurred and darkened, with the sharp panel laid on top: the background moves with the picture instead of holding a colour, and the border stops showing.
The blur is never computed at full resolution. The source is reduced to a 120-pixel-tall buffer, blurred there, then stretched back: the cost drops from 430 milliseconds per frame to 0.44.
What it will not decide
Most of the limits are not failures: they are refusals on purpose. The result is almost always the same, the whole shot, but the reason changes.
One decision is also overturned after the fact: the analysis runs on a reduced picture, and a full-resolution counter-check throws away the detections it invented. The first case met in production was a wooden wall split in two for twelve seconds.
Taking it back, span by span
Reframing is not a switch you flip before generation, and that is not an oversight: deciding before seeing the result means guessing. What gets fixed is a clip that already exists, in front of you.
The Framing tab of the editor does not show the reframed clip: it shows the whole source, with the window in use at the current instant drawn on top, and the vertical preview beside it. So you see the gap between what the camera filmed and what the clip keeps. The block below is that gesture, on a still frame.
drag the frame, the wheel tightens it- Pick the mode on the rowThree choices per span: left to the AI, one frame, or two stacked frames. The manual vocabulary is deliberately smaller than the automatic one.
- Place the windowDrag the frame over the source, the wheel tightens it, down to a quarter of the source height. It cannot leave the picture, and its ratio is imposed by its destination.
- Re-cut if a boundary falls badlySplitting cuts at the playhead, including in the middle of a detected shot. Merging gives the span back to its left neighbour, keeping that one's crop.
- SaveOne render, for this clip. Spans left to the AI keep the computed crop; cancelling restores the original plan.
Only the spans you took over go to the worker. In practice one or two scenes are enough, and they are almost always the same ones: the on-screen demo, the object held up to the lens, the wide shot where you would rather see a face.
Frequently asked questions
Can I fix a crop I do not like?
Yes, span by span, in the editor. The clip is pre-cut on the detected shots; each span is either left to the AI or taken over by hand with one frame or two, and can be split anywhere. Only the spans you took over go back into production.Does it work on gameplay?
Yes, as long as there is a facecam: that is what gets measured on the picture, and that is what decides the layout. Gameplay without a webcam comes out as the whole shot, which is the right answer.What if my video is already vertical?
Nothing is reframed. A 9:16 or a square has no width to spare, and taking some would be destructive.Does the frame move for the whole clip?
It moves when the subject moves, and holds its place the rest of the time. The camera path is laid piecewise: held plateaus and clean moves, never a permanent float.Can I turn automatic reframing off?
There is no switch before generation, and that is deliberate: it would ask you to decide without having seen the result. What can be adjusted is adjusted afterwards, on a clip you have in front of you.60 minutes of video a month, no card needed
Reframing, captions and the editor are all in the free tier. The only way to know what it does with your shots is to give it one.
Try Snip