Skip to main content
Video you produce is predictable. You control the camera, the encoder and the settings, so the files are consistent and your pipeline can assume things about them. Video your users produce is not. It arrives from phones you have never heard of, from screen recorders, from downloads of downloads, from a decade of iOS versions. Every assumption your pipeline makes about it is a bet. This is what tends to be in that footage, where it surfaces, and how to order the work so problems appear early and cheaply.

What actually arrives

Not an exhaustive list — the four that account for most real failures: Variable frame rate. Phones drop frames in low light, when hot, or when the CPU is busy. The file still declares a rate; the gaps between frames are uneven. Anything computing time from frame counts drifts. See Variable frame rate. Rotation metadata. Vertical video is stored as a landscape frame plus a flag saying which way up. Players honour the flag; a lot of frame-extraction code does not. See Why vertical video arrives sideways. Codecs your stack tolerates unevenly. HEVC is normal from an iPhone and awkward everywhere else — decodable, but slower, and unevenly supported in browsers. It rarely fails outright; it makes everything cost more. Wildly inconsistent bitrate and resolution. A 4K 120Mbps clip and a 480p screen recording arrive through the same endpoint. Anything sized for the average falls over on the tail. Two more worth expecting: no audio track at all (screen recordings, some drone footage), and missing provenance — no creation time, no device — which matters if you need to order footage by when it was shot.

Where each one surfaces

The expensive property of these defects is when you find out. Almost all of them are invisible at upload and visible only after work has been done. Read that top to bottom and the shape of the problem is obvious: the two stages where you would notice are the two stages where nothing goes wrong. Everything else fails after you have spent money.

Order the work so failure is cheap

The instinct is to accept the upload, start processing, and handle errors as they come. That works, and it means every failure costs a full run. A cheaper order: 1. Inspect the header before anything else. Almost everything above is visible in the container’s metadata, which costs a fraction of a second and a few megabytes to read — you do not need the whole file. Frame rate, codec, dimensions, rotation, duration and audio presence are all there. 2. Decide against the job, not in general. There is no such thing as a good file, only a file that suits what you are about to do. 24fps is normal for editing and too slow for motion tracking. A rotation flag is invisible to an editor and fatal to a pose model. Deciding in general means picking one caller’s answer and being wrong for the rest. 3. Fix what is fixable, reject what is not, and know the difference. Variable frame rate can be conformed away. Frames that were never captured cannot be added — no amount of processing turns 12fps into 30. Sending unrecoverable footage to a transcode wastes the transcode and delays the bad news. 4. Only then do the expensive part. Transcode, transcribe, analyse — against footage you already know can support it. 5. Check the result of any fix. A transcode that ran is not a transcode that worked. If you conform a frame rate, verify the output is actually constant, the duration did not change, and the audio survived. Encoders accept instructions they do not honour, and the failure is silent.

What this looks like in practice

Roughly:
The two branches that matter are the ones people leave out. Telling a user their clip will not work while they still have the camera in their hand is worth more than telling them twenty minutes later, and it costs you nothing. And verifying a fix is the difference between “we ran a transcode” and “the problem is gone” — which are not the same claim, and only one of them is worth making to a customer.

If you build this yourself

It is a reasonable thing to build, and the pieces are not exotic — ffprobe reads the header, ffmpeg performs the fixes. The parts that take longer than expected:
  • Reading rotation from both places it hides, or you miss half of them.
  • Sorting frame timestamps by presentation order before measuring gaps, or B-frames make constant footage look variable.
  • Sampling more than the first few seconds, because a phone that stutters when hot records the opening perfectly.
  • Deciding what a defect means per use case, which is where most of the actual thinking lives and where a health score stops being enough.
  • Verifying fixes, which means measuring the output as carefully as you measured the input.
None of it is hard. All of it is fiddly, and most of it is only learned by getting it wrong once.
FastDrop is those steps as two API calls: one that reads the header and answers whether the footage suits a named job, and one that executes the prescribed fix and re-checks the result before handing it back. Why check footage before you process it covers the reasoning; the quickstart runs the whole loop on sample files for free.