Transcribing podcasts locally with Parakeet TDT and ffmpeg
Why Local?
Podcast episodes can be 60+ minutes long. Uploading a 100 MB MP3 to a cloud API for transcription is slow, costs money, and raises privacy concerns. NVIDIA Parakeet TDT 0.6B v2 is an open-source ASR model that transcribes 60 minutes of audio in roughly 1 second on a compatible GPU (or Apple Silicon via parakeet-mlx). Everything stays on your machine.
This post covers a simple pipeline (split, compress, transcribe, save) using ffmpeg, curl, and a small bash script.
1. Split Long Audio into Chunks
Most local ASR servers have a max input duration. For a 1-hour podcast, splitting into 10-minute chunks keeps each request well under any limit.
#!/usr/bin/env bash
set -euo pipefail
usage() {
echo "Usage: $0 <input_file> [chunk_seconds]"
echo " chunk_seconds: max chunk duration (default: 600 = 10 min)"
exit 1
}
[[ $# -lt 1 || $# -gt 2 ]] && usage
input="$1"
chunk="${2:-600}"
[[ -f "$input" ]] || { echo "Error: '$input' not found." >&2; exit 1; }
command -v ffmpeg &>/dev/null || { echo "Error: ffmpeg not found." >&2; exit 1; }
duration=$(ffprobe -v error -show_entries format=duration \
-of default=noprint_wrappers=1:nokey=1 "$input")
duration=${duration%%.*}
count=$(( (duration + chunk - 1) / chunk ))
echo "Splitting $(basename "$input") (${duration}s) into ${count} chunks of ${chunk}s..."
ffmpeg -y -i "$input" -c copy \
-f segment -segment_time "$chunk" \
"${input%.*}_chunk_%03d.${input##*.}"
echo "Done: ${count} files created."
chmod +x split_audio.sh
./split_audio.sh podcast.mp3 # 10-min chunks
./split_audio.sh podcast.mp3 300 # 5-min chunks
-c copy performs zero re-encoding, so the split is instant and lossless.
2. Compress (Optional)
If your source is high-bitrate stereo and you send it over a network, shrinking the file first saves bandwidth:
# Voice/podcast: mono + 64 kbps (no audible loss for speech)
ffmpeg -i podcast.mp3 -b:a 64k -ac 1 podcast_compressed.mp3
# Music: VBR quality 4
ffmpeg -i podcast.mp3 -q:a 4 podcast_compressed.mp3
| Goal | Flags |
|---|---|
| Voice / podcast | -b:a 64k -ac 1 |
| Music, good quality | -b:a 192k or -q:a 2 |
| Smallest possible | -b:a 32k -ac 1 |
3. Transcribe with curl
NVIDIA NIM (Docker, local)
curl -s http://localhost:9000/v1/audio/transcriptions \
-F language=en-US \
-F file="@podcast_chunk_000.mp3" \
-o transcript_000.json
OpenAI-compatible local server
curl -s -X POST "http://localhost:5092/v1/audio/transcriptions" \
-F "file=@podcast_chunk_000.mp3" \
-F "model=parakeet-tdt-0.6b-v3" \
-F "response_format=srt" \
-F "language=en-US"
-o transcript_000.srt
To save the response, use -o filename for an explicit output path, or shell redirection (> file). Add -w "%{http_code}\n" to print the status code for quick verification:
curl -s -o transcript.json -w "%{http_code}\n" \
http://localhost:9000/v1/audio/transcriptions \
-F language=en-US \
-F file="@podcast.mp3"
4. Put It All Together
#!/usr/bin/env bash
set -euo pipefail
input="$1"
base="${input%.*}"
ext="${input##*.}"
api="http://localhost:9000/v1/audio/transcriptions"
# 1. Split
./split_audio.sh "$input" 600
# 2. Transcribe each chunk
for f in "${base}"_chunk_*.${ext}; do
out="${f%.*}.json"
echo "Transcribing $f → $out"
curl -s -o "$out" \
"$api" \
-F language=en-US \
-F "file=@$f"
done
# 3. (Optional) Clean up chunks
rm "${base}"_chunk_*.${ext}
echo "All done."
Summary
| Step | Tool | What it does |
|---|---|---|
| Split | ffmpeg -f segment | Breaks long audio into fixed-size chunks |
| Compress | ffmpeg -b:a / -q:a | Reduces file size for network transfer |
| Transcribe | curl → Parakeet API | Converts audio to text (JSON or SRT) |
| Save | curl -o / > | Writes the response to a file |
The entire pipeline is local, free, and fast, with no cloud round-trips, per-minute billing, or external data transfers.