Skip to main content
FastDrop’s /transcribe endpoint handles speech-to-text, language detection, English translation, and subtitle file generation in a single API call. This guide explains the pipeline, how to get the best results, and common integration patterns.

How it works

When you submit a video to /transcribe, FastDrop runs a multi-step pipeline:
  1. Audio extraction — The audio track is extracted from your video file
  2. Language detection — The spoken language is automatically identified (or uses your language hint)
  3. Transcription — Speech is converted to text with word-level and segment-level timestamps
  4. Translation — If the source language isn’t English and translate is true, a full English translation is generated
  5. File generation — If output_formats is specified, SRT and/or TXT subtitle files are generated and uploaded
The entire pipeline runs asynchronously. You submit the job and either receive a webhook or fetch the result with ?wait=N. See Getting your results.

Response modes

JSON only (default)

When you don’t specify output_formats, you get a structured JSON response with the full transcript, segments, word-level timestamps, and (if applicable) an English translation. This is ideal for:
  • Indexing transcripts for search
  • Extracting quotes or key moments
  • Building custom UI around transcript data
  • Feeding transcripts into downstream processing

With subtitle files

Pass output_formats to generate downloadable subtitle files alongside the JSON response. This is ideal for:
  • Adding captions to video players
  • Importing subtitles into editing software (Premiere, DaVinci, Final Cut)
  • Delivering accessible video content
  • Archiving transcripts as standalone files
Available formats: When translation is performed, you get both source-language and English files:
  • source_srt / source_txt — Transcript in the original language
  • english_srt / english_txt — English translation (only for non-English sources)
File download URLs are presigned and expire after 1 hour. Download or store them immediately after receiving results.

Language detection vs. language hints

Auto-detection

By default, FastDrop detects the spoken language from the audio. This works well for high-accuracy languages but can struggle with lower-resource languages that sound similar to a higher-resource one.

Using the language parameter

Passing a language hint bypasses detection entirely. This is recommended when:
  • You already know the language (e.g., from user input or metadata)
  • You’re working with a medium or low-accuracy language
  • You’re processing a batch of videos in the same language
  • Detection is returning the wrong language
See Supported Languages for all valid ISO 639-1 codes.

Translation

When the source language is not English and translate is true (the default), FastDrop generates a full English translation. The translation:
  • Preserves segment-level timestamps from the original transcription
  • Is generated from the transcribed text, not directly from audio
  • Works across all 99 supported languages

Disabling translation

If you only need the source-language transcript, set translate to false to skip the translation step. This doesn’t reduce the credit cost but slightly speeds up processing.

English source videos

When the detected (or specified) language is English, no translation is performed regardless of the translate setting. The translation field will be absent from the response.

Timestamps

FastDrop returns two levels of timing data:

Segments

Sentence-level chunks with start and end times. These map naturally to subtitle lines:

Words

Individual word timestamps for precise alignment. Useful for karaoke-style captions or precise text-to-video sync:
Word-level timestamps are available for the source transcription only. Translated text uses segment-level timestamps carried over from the source.

Accuracy considerations

Transcription accuracy varies by language. FastDrop classifies languages into three tiers based on training data availability: See Supported Languages for the complete breakdown.

Factors that affect accuracy

Beyond language tier, these factors influence transcription quality:
  • Audio quality — Clear audio with minimal background noise produces the best results
  • Speaker clarity — Single speakers with clear enunciation transcribe better than overlapping speakers
  • Domain vocabulary — Technical, medical, or domain-specific terms may be misrecognized
  • Accents and dialects — Strong regional accents or dialectal variation can reduce accuracy, especially for languages marked as “high accuracy” based on their standard form

Common workflows

Subtitles for multilingual content

Returns both source_srt (Arabic) and english_srt (English translation) ready to import into any video editor.

Search indexing

Use the text field to index the full transcript for search. The segments field lets you link search results to specific timestamps.

Batch transcription

Use the batch endpoint to transcribe up to 50 videos at once. Each video in the batch can include transcribe as a capability alongside other capabilities like classify or thumbnails.

Combine with classification

Request transcription alongside classification in a single job:
This costs 2 (classify) + 8 (transcribe) + 1 (thumbnails) = 11 credits and runs all capabilities in parallel.