> ## Documentation Index
> Fetch the complete documentation index at: https://developers.fastdrop.io/llms.txt
> Use this file to discover all available pages before exploring further.

# Transcription Guide

> How FastDrop transcription works, best practices, and common workflows

FastDrop's `/transcribe` endpoint handles speech-to-text, language detection, English translation, and subtitle file generation in a single API call. This guide explains the pipeline, how to get the best results, and common integration patterns.

## How it works

When you submit a video to `/transcribe`, FastDrop runs a multi-step pipeline:

1. **Audio extraction** — The audio track is extracted from your video file
2. **Language detection** — The spoken language is automatically identified (or uses your `language` hint)
3. **Transcription** — Speech is converted to text with word-level and segment-level timestamps
4. **Translation** — If the source language isn't English and `translate` is `true`, a full English translation is generated
5. **File generation** — If `output_formats` is specified, SRT and/or TXT subtitle files are generated and uploaded

The entire pipeline runs asynchronously. You submit the job and either receive a [webhook](/guides/webhooks) or fetch the result with `?wait=N`. See [Getting your results](/guides/delivery-modes).

## Response modes

### JSON only (default)

When you don't specify `output_formats`, you get a structured JSON response with the full transcript, segments, word-level timestamps, and (if applicable) an English translation. This is ideal for:

* Indexing transcripts for search
* Extracting quotes or key moments
* Building custom UI around transcript data
* Feeding transcripts into downstream processing

```bash theme={null}
curl -X POST https://api.fastdrop.io/api/v1/transcribe \
  -H "Content-Type: application/json" \
  -H "X-API-Key: fd_live_your_key_here" \
  -d '{"video_url": "https://example.com/video.mp4"}'
```

### With subtitle files

Pass `output_formats` to generate downloadable subtitle files alongside the JSON response. This is ideal for:

* Adding captions to video players
* Importing subtitles into editing software (Premiere, DaVinci, Final Cut)
* Delivering accessible video content
* Archiving transcripts as standalone files

```bash theme={null}
curl -X POST https://api.fastdrop.io/api/v1/transcribe \
  -H "Content-Type: application/json" \
  -H "X-API-Key: fd_live_your_key_here" \
  -d '{
    "video_url": "https://example.com/video.mp4",
    "output_formats": ["srt", "txt"]
  }'
```

Available formats:

| Format | Description |
| - | - |
| `srt` | SubRip subtitle format with timestamps. Compatible with most video players and editors. |
| `txt` | Plain text transcript without timestamps. Useful for reading, archiving, or text processing. |

When translation is performed, you get both source-language and English files:

* `source_srt` / `source_txt` — Transcript in the original language
* `english_srt` / `english_txt` — English translation (only for non-English sources)

<Warning>
  File download URLs are presigned and **expire after 1 hour**. Download or store them immediately after receiving results.
</Warning>

## Language detection vs. language hints

### Auto-detection

By default, FastDrop detects the spoken language from the audio. This works well for [high-accuracy languages](/guides/supported-languages#high-accuracy--8-wer) but can struggle with lower-resource languages that sound similar to a higher-resource one.

### Using the `language` parameter

Passing a `language` hint bypasses detection entirely. This is recommended when:

* You already know the language (e.g., from user input or metadata)
* You're working with a [medium or low-accuracy language](/guides/supported-languages)
* You're processing a batch of videos in the same language
* Detection is returning the wrong language

```json theme={null}
{
  "video_url": "https://example.com/swahili-interview.mp4",
  "language": "sw"
}
```

See [Supported Languages](/guides/supported-languages) for all valid ISO 639-1 codes.

## Translation

When the source language is not English and `translate` is `true` (the default), FastDrop generates a full English translation. The translation:

* Preserves segment-level timestamps from the original transcription
* Is generated from the transcribed text, not directly from audio
* Works across all 99 supported languages

### Disabling translation

If you only need the source-language transcript, set `translate` to `false` to skip the translation step. This doesn't reduce the credit cost but slightly speeds up processing.

```json theme={null}
{
  "video_url": "https://example.com/video.mp4",
  "translate": false
}
```

### English source videos

When the detected (or specified) language is English, no translation is performed regardless of the `translate` setting. The `translation` field will be absent from the response.

## Timestamps

FastDrop returns two levels of timing data:

### Segments

Sentence-level chunks with start and end times. These map naturally to subtitle lines:

```json theme={null}
{
  "segments": [
    { "start": 0.0, "end": 2.8, "text": "Welcome to the show." },
    { "start": 3.1, "end": 8.5, "text": "Today we're going to talk about video editing." }
  ]
}
```

### Words

Individual word timestamps for precise alignment. Useful for karaoke-style captions or precise text-to-video sync:

```json theme={null}
{
  "words": [
    { "start": 0.0, "end": 0.5, "word": "Welcome" },
    { "start": 0.5, "end": 0.7, "word": "to" },
    { "start": 0.7, "end": 0.9, "word": "the" },
    { "start": 0.9, "end": 1.3, "word": "show." }
  ]
}
```

<Note>
  Word-level timestamps are available for the source transcription only. Translated text uses segment-level timestamps carried over from the source.
</Note>

## Accuracy considerations

Transcription accuracy varies by language. FastDrop classifies languages into three tiers based on training data availability:

| Tier | WER range | Languages |
| - | :-: | - |
| High | Under 8% | English, Spanish, French, German, Japanese, and 25 more |
| Medium | 8–15% | Bengali, Tamil, Swahili, Persian, Welsh, and 30 more |
| Low | 15–40%+ | Amharic, Yoruba, Hausa, Hawaiian, Tibetan, and 28 more |

See [Supported Languages](/guides/supported-languages) for the complete breakdown.

### Factors that affect accuracy

Beyond language tier, these factors influence transcription quality:

* **Audio quality** — Clear audio with minimal background noise produces the best results
* **Speaker clarity** — Single speakers with clear enunciation transcribe better than overlapping speakers
* **Domain vocabulary** — Technical, medical, or domain-specific terms may be misrecognized
* **Accents and dialects** — Strong regional accents or dialectal variation can reduce accuracy, especially for languages marked as "high accuracy" based on their standard form

## Common workflows

### Subtitles for multilingual content

```json theme={null}
{
  "video_url": "https://example.com/arabic-interview.mp4",
  "output_formats": ["srt"]
}
```

Returns both `source_srt` (Arabic) and `english_srt` (English translation) ready to import into any video editor.

### Search indexing

```json theme={null}
{
  "video_url": "https://example.com/podcast.mp4"
}
```

Use the `text` field to index the full transcript for search. The `segments` field lets you link search results to specific timestamps.

### Batch transcription

Use the [batch endpoint](/api-reference/create-batch) to transcribe up to 50 videos at once. Each video in the batch can include `transcribe` as a capability alongside other capabilities like `classify` or `thumbnails`.

### Combine with classification

Request transcription alongside classification in a single job:

```json theme={null}
{
  "video_url": "https://example.com/footage.mp4",
  "capabilities": ["classify", "transcribe", "thumbnails"],
  "processing_tier": "basic"
}
```

This costs 2 (classify) + 8 (transcribe) + 1 (thumbnails) = 11 credits and runs all capabilities in parallel.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.