Product

Smart reframing: how Snip decides where to look

Turning a 16:9 source vertical means throwing away two thirds of the width. Seven real extracts, recomposed live in your browser, to see what Snip recognises shot by shot, how the frame follows a face, what it refuses to decide, and how you take a passage back by hand.

Mehdi Seddik17 August 20269 min read
In short
01The frame follows the face continuously, over 81 to 87 % of its travel. The remaining 13 to 19 % are damping on purpose, not lag.02The layout is decided shot by shot: tracking, split screen, triptych, 2x2 grid, facecam over gameplay, document above the camera.03When none of those answers is certain, the whole shot is kept, over a blurred background pulled from the scene. A readable refusal beats an invented crop.

Seven shots, seven answers

Every tab is a real extract, carrying the plan the production sidecar computed on it. The vertical on the right is not a video: it is composed in your browser, frame by frame, by the same function that burns the export. Switch to « centre crop » to see the lazy answer on the very same extract.

Loading the extract…
the 16:9 source
what Snip burns
One person: the frame follows her. Fixed camera, moving subject. The window travels 194 pixels over these nine seconds, without a jolt, and the forehead is never cut. extract — Scilabus

Two thirds of the picture to throw away

A 16:9 source is 1920 pixels wide, a vertical clip is 1080. Reframing does not shrink the picture: it picks a window inside it, and everything outside disappears. On a three-seat set, that window decides who stays in the clip.

So the work happens in two steps, and the second one is the interesting one: decide what to show, then build a camera move that holds up. Snip cuts the clip into shots, decides a layout on each, and solves a camera path over its whole length.

66 %of the source width a 9:16 leaves out81–87 %of the face travel the frame follows0foreheads cut across the measured corpus

The frame follows, without shaking

This is the first criterion, and the only one you can see without measuring anything: if the face drifts inside the frame, the crop is wrong even without a single shake.

On the measured clips the frame follows 81 to 87 % of the face travel, with a median gap of 0.011 to 0.074 face heights between the two centres. On a filmed set that is 0.015 face heights, which is 6 pixels on a 400-pixel face.

The 13 to 19 % it does not follow are not a defect, they are the damping: an operator does not render every nod. The chain that produces that compromise has four stages.

  1. Detect, then stitchFaces are detected frame after frame, then stitched into stable tracks: one track, one person, for the length of the shot. Two fragments that overlap at the same instant are the same head, never two people sitting side by side.
  2. Smooth the targetA Kalman filter returns a continuous target with a velocity that means something. It bridges short detection gaps, and freezes the frame instead of drifting when a gap lasts.
  3. Lay a camera pathA linear program turns that target into a piecewise trajectory: held plateaus and clean moves, rather than permanent float.
  4. Contain, cap, anticipateThe measured face has to stay inside the window, forehead and chin included. Speed is capped at roughly one second to cross half the picture, and the constraint bites a third of a second early so the move spreads out instead of whipping.

Two cases are not moves and must not become one: a speaker change, and a subject jump on the same camera. Letting the solver join the two positions produces a 1 160-pixel displacement spread over four frames, 67 milliseconds: that is neither a clean cut nor a readable pan, it is a whip. Between two speakers an editor cuts, they do not sweep the room. Those instants are therefore treated as cuts, and the two halves are solved separately.

« When in doubt, a simple stable crop always beats a clever wrong one. »

What it recognises, shot by shot

The layout is decided on each shot, never once for the whole clip. The number of framable people decides the shape, and a measured overlay changes the answer.

One personFull-frame tracking. The nominal case.Two peopleSplit screen, one per half. Both are on screen, so there is no way to pick the wrong person.Three peopleTriptych: one full-width cell on top, two below. Nobody is sacrificed.Four peopleA 2x2 grid. A cell is the size of the bottom cells of a triptych; past four, none of them would show anything.Facecam over gameplayThe overlay is measured, not guessed: face on the top quarter, game on the other three, two static panels.Document and cameraThe document takes the top half, shown whole; the camera fills the bottom half edge to edge.Several webcamsFrames the column of overlays, without picking one of them.Vertical or square sourceNo reframing at all: it would be destructive.

There is no speaker switch inside a split screen: both people stay visible throughout. The mouth signal is measured and reported in the metadata, it does not decide the crop. Without the audio, nothing lets us verify that the winning face is the one talking.

Keeping the whole shot is a decision

When no layout is certain, the whole shot is kept. That is not an admission: on a landscape, a football pitch, a concert stage, a centred 9:16 throws away two thirds of the width to frame whatever sat in the middle. The director's frame beats a crop that pretends to be a decision.

That leaves the two free bands, a third of the output. They were black, then tinted with a halo. Today they are the scene itself, cropped to fill the vertical, blurred and darkened, with the sharp panel laid on top: the background moves with the picture instead of holding a colour, and the border stops showing.

The blur is never computed at full resolution. The source is reduced to a 120-pixel-tall buffer, blurred there, then stretched back: the cost drops from 430 milliseconds per frame to 0.44.

black bands
today's background
Both pictures are composed side by side by the same function, on the same frame of the same source. The only difference is the boolean the export carries.

What it will not decide

Most of the limits are not failures: they are refusals on purpose. The result is almost always the same, the whole shot, but the reason changes.

More than four peoplePicking one would pick the wrong person, usually whoever is closest to the camera and listening. Framing all of them would show none.An uncertain countSomeone just under the size threshold, or seen intermittently: we cannot tell four from five, and the layout follows the count.What is shown is not a faceHands, an object, a whiteboard, a minimap: nothing here detects an object or a region of interest. A face deliberately breaking the frame is read as composition, and left alone.A face that is too smallUnder 8 % of the height, following it would fill the vertical with a stretched thumbnail and hide what we came to see.A dissolveIt is not detected, and we do not pretend: measured, a dissolve peaks at 0.082 where real cuts sit at 0.35 and 0.54, against a median noise of 0.038. No threshold separates them.The action inside a gameThe gameplay band keeps the centre of the screen. Whatever falls outside, minimap or item bar, is lost by construction, not by choice.

One decision is also overturned after the fact: the analysis runs on a reduced picture, and a full-resolution counter-check throws away the detections it invented. The first case met in production was a wooden wall split in two for twelve seconds.

Taking it back, span by span

Reframing is not a switch you flip before generation, and that is not an oversight: deciding before seeing the result means guessing. What gets fixed is a clip that already exists, in front of you.

The Framing tab of the editor does not show the reframed clip: it shows the whole source, with the window in use at the current instant drawn on top, and the vertical preview beside it. So you see the gap between what the camera filmed and what the clip keeps. The block below is that gesture, on a still frame.

drag the frame, the wheel tightens it
The geometry is the editor's, line for line: same rectangle computation, same split into two halves, same floor at a quarter of the source height.
  1. Pick the mode on the rowThree choices per span: left to the AI, one frame, or two stacked frames. The manual vocabulary is deliberately smaller than the automatic one.
  2. Place the windowDrag the frame over the source, the wheel tightens it, down to a quarter of the source height. It cannot leave the picture, and its ratio is imposed by its destination.
  3. Re-cut if a boundary falls badlySplitting cuts at the playhead, including in the middle of a detected shot. Merging gives the span back to its left neighbour, keeping that one's crop.
  4. SaveOne render, for this clip. Spans left to the AI keep the computed crop; cancelling restores the original plan.

Only the spans you took over go to the worker. In practice one or two scenes are enough, and they are almost always the same ones: the on-screen demo, the object held up to the lens, the wide shot where you would rather see a face.

See what it does with your last videoPaste the link. Snip transcribes, picks the passages, frames every shot and shows you the scenes before the export.Clip it

Frequently asked questions

Can I fix a crop I do not like?

Yes, span by span, in the editor. The clip is pre-cut on the detected shots; each span is either left to the AI or taken over by hand with one frame or two, and can be split anywhere. Only the spans you took over go back into production.

60 minutes of video a month, no card needed

Reframing, captions and the editor are all in the free tier. The only way to know what it does with your shots is to give it one.

Try Snip

Read next

Snip against the other toolsAn alternative to Opus ClipHow Snip works