Alex Leko
All content on this blog was fully or partially created using local AI (Apple MLX).

Transcribing podcasts locally with Parakeet TDT and ffmpeg

  • ffmpeg
  • transcription
  • parakeet
  • bash
  • audio
  • self-hosted

Why Local?

Podcast episodes can be 60+ minutes long. Uploading a 100 MB MP3 to a cloud API for transcription is slow, costs money, and raises privacy concerns. NVIDIA Parakeet TDT 0.6B v2 is an open-source ASR model that transcribes 60 minutes of audio in roughly 1 second on a compatible GPU (or Apple Silicon via parakeet-mlx). Everything stays on your machine.

This post covers a simple pipeline (split, compress, transcribe, save) using ffmpeg, curl, and a small bash script.

1. Split Long Audio into Chunks

Most local ASR servers have a max input duration. For a 1-hour podcast, splitting into 10-minute chunks keeps each request well under any limit.

#!/usr/bin/env bash
set -euo pipefail

usage() {
  echo "Usage: $0 <input_file> [chunk_seconds]"
  echo "  chunk_seconds: max chunk duration (default: 600 = 10 min)"
  exit 1
}

[[ $# -lt 1 || $# -gt 2 ]] && usage

input="$1"
chunk="${2:-600}"

[[ -f "$input" ]] || { echo "Error: '$input' not found." >&2; exit 1; }
command -v ffmpeg &>/dev/null || { echo "Error: ffmpeg not found." >&2; exit 1; }

duration=$(ffprobe -v error -show_entries format=duration \
  -of default=noprint_wrappers=1:nokey=1 "$input")
duration=${duration%%.*}

count=$(( (duration + chunk - 1) / chunk ))

echo "Splitting $(basename "$input") (${duration}s) into ${count} chunks of ${chunk}s..."

ffmpeg -y -i "$input" -c copy \
  -f segment -segment_time "$chunk" \
  "${input%.*}_chunk_%03d.${input##*.}"

echo "Done: ${count} files created."
chmod +x split_audio.sh
./split_audio.sh podcast.mp3        # 10-min chunks
./split_audio.sh podcast.mp3 300    # 5-min chunks

-c copy performs zero re-encoding, so the split is instant and lossless.

2. Compress (Optional)

If your source is high-bitrate stereo and you send it over a network, shrinking the file first saves bandwidth:

# Voice/podcast: mono + 64 kbps (no audible loss for speech)
ffmpeg -i podcast.mp3 -b:a 64k -ac 1 podcast_compressed.mp3
# Music: VBR quality 4
ffmpeg -i podcast.mp3 -q:a 4 podcast_compressed.mp3
Goal Flags
Voice / podcast -b:a 64k -ac 1
Music, good quality -b:a 192k or -q:a 2
Smallest possible -b:a 32k -ac 1

3. Transcribe with curl

NVIDIA NIM (Docker, local)

curl -s http://localhost:9000/v1/audio/transcriptions \
  -F language=en-US \
  -F file="@podcast_chunk_000.mp3" \
  -o transcript_000.json

OpenAI-compatible local server

curl -s -X POST "http://localhost:5092/v1/audio/transcriptions" \
  -F "file=@podcast_chunk_000.mp3" \
  -F "model=parakeet-tdt-0.6b-v3" \
  -F "response_format=srt" \
  -F "language=en-US"
  -o transcript_000.srt

To save the response, use -o filename for an explicit output path, or shell redirection (> file). Add -w "%{http_code}\n" to print the status code for quick verification:

curl -s -o transcript.json -w "%{http_code}\n" \
  http://localhost:9000/v1/audio/transcriptions \
  -F language=en-US \
  -F file="@podcast.mp3"

4. Put It All Together

#!/usr/bin/env bash
set -euo pipefail

input="$1"
base="${input%.*}"
ext="${input##*.}"
api="http://localhost:9000/v1/audio/transcriptions"

# 1. Split
./split_audio.sh "$input" 600

# 2. Transcribe each chunk
for f in "${base}"_chunk_*.${ext}; do
  out="${f%.*}.json"
  echo "Transcribing $f → $out"
  curl -s -o "$out" \
    "$api" \
    -F language=en-US \
    -F "file=@$f"
done

# 3. (Optional) Clean up chunks
rm "${base}"_chunk_*.${ext}

echo "All done."

Summary

Step Tool What it does
Split ffmpeg -f segment Breaks long audio into fixed-size chunks
Compress ffmpeg -b:a / -q:a Reduces file size for network transfer
Transcribe curl → Parakeet API Converts audio to text (JSON or SRT)
Save curl -o / > Writes the response to a file

The entire pipeline is local, free, and fast, with no cloud round-trips, per-minute billing, or external data transfers.