/transcribe endpoint handles speech-to-text, language detection, English translation, and subtitle file generation in a single API call. This guide explains the pipeline, how to get the best results, and common integration patterns.
How it works
When you submit a video to/transcribe, FastDrop runs a multi-step pipeline:
- Audio extraction — The audio track is extracted from your video file
- Language detection — The spoken language is automatically identified (or uses your
languagehint) - Transcription — Speech is converted to text with word-level and segment-level timestamps
- Translation — If the source language isn’t English and
translateistrue, a full English translation is generated - File generation — If
output_formatsis specified, SRT and/or TXT subtitle files are generated and uploaded
?wait=N. See Getting your results.
Response modes
JSON only (default)
When you don’t specifyoutput_formats, you get a structured JSON response with the full transcript, segments, word-level timestamps, and (if applicable) an English translation. This is ideal for:
- Indexing transcripts for search
- Extracting quotes or key moments
- Building custom UI around transcript data
- Feeding transcripts into downstream processing
With subtitle files
Passoutput_formats to generate downloadable subtitle files alongside the JSON response. This is ideal for:
- Adding captions to video players
- Importing subtitles into editing software (Premiere, DaVinci, Final Cut)
- Delivering accessible video content
- Archiving transcripts as standalone files
When translation is performed, you get both source-language and English files:
source_srt/source_txt— Transcript in the original languageenglish_srt/english_txt— English translation (only for non-English sources)
Language detection vs. language hints
Auto-detection
By default, FastDrop detects the spoken language from the audio. This works well for high-accuracy languages but can struggle with lower-resource languages that sound similar to a higher-resource one.Using the language parameter
Passing a language hint bypasses detection entirely. This is recommended when:
- You already know the language (e.g., from user input or metadata)
- You’re working with a medium or low-accuracy language
- You’re processing a batch of videos in the same language
- Detection is returning the wrong language
Translation
When the source language is not English andtranslate is true (the default), FastDrop generates a full English translation. The translation:
- Preserves segment-level timestamps from the original transcription
- Is generated from the transcribed text, not directly from audio
- Works across all 99 supported languages
Disabling translation
If you only need the source-language transcript, settranslate to false to skip the translation step. This doesn’t reduce the credit cost but slightly speeds up processing.
English source videos
When the detected (or specified) language is English, no translation is performed regardless of thetranslate setting. The translation field will be absent from the response.
Timestamps
FastDrop returns two levels of timing data:Segments
Sentence-level chunks with start and end times. These map naturally to subtitle lines:Words
Individual word timestamps for precise alignment. Useful for karaoke-style captions or precise text-to-video sync:Word-level timestamps are available for the source transcription only. Translated text uses segment-level timestamps carried over from the source.
Accuracy considerations
Transcription accuracy varies by language. FastDrop classifies languages into three tiers based on training data availability:
See Supported Languages for the complete breakdown.
Factors that affect accuracy
Beyond language tier, these factors influence transcription quality:- Audio quality — Clear audio with minimal background noise produces the best results
- Speaker clarity — Single speakers with clear enunciation transcribe better than overlapping speakers
- Domain vocabulary — Technical, medical, or domain-specific terms may be misrecognized
- Accents and dialects — Strong regional accents or dialectal variation can reduce accuracy, especially for languages marked as “high accuracy” based on their standard form
Common workflows
Subtitles for multilingual content
source_srt (Arabic) and english_srt (English translation) ready to import into any video editor.
Search indexing
text field to index the full transcript for search. The segments field lets you link search results to specific timestamps.
Batch transcription
Use the batch endpoint to transcribe up to 50 videos at once. Each video in the batch can includetranscribe as a capability alongside other capabilities like classify or thumbnails.