Align .srt / .lrc / .txt subtitles (or generate them) to audio/video with
WhisperX forced alignment — so each cue can
move independently instead of only applying one global timeline shift.
Tools like ffsubsync typically find a constant offset (or stretch) between speech activity and subtitle “on” times. That works well for whole-track drift, but leading/trailing silence or local timing errors can still leave lines early or late.
sub-align picks a strategy from the input type, then runs WhisperX
phoneme / word-level forced alignment so each line is refined against the
audio:
| Input | Strategy |
|---|---|
| Media only | Whisper ASR → word-align → split into timed cues |
.txt script |
ASR only for search windows → forced-align original lines |
.srt / .lrc |
Optional global offset → expand windows by --margin → forced-align |
Limitations: subtitle text must roughly match spoken content (no translation). Refine can still nudge already-good cues; very short lines may be merged with close neighbors (large pauses stay alone); VAD only blocks refine starts pulled earlier into silence. There is no speaker diarization. See docs/pipeline.md § Limitations.
More detail: docs/pipeline.md · scenarios & flags: docs/usage.md
Requires Python 3.10+ and ffmpeg on PATH. First run
downloads WhisperX alignment models (disk/RAM).
pip install 'sub-align[align]'
# or
uv pip install 'sub-align[align]'Extras [align], [cpu], and [gpu] all install WhisperX. Install a matching
PyTorch build first when you need a specific CPU/CUDA wheel:
# CPU
uv pip install torch --index-url https://download.pytorch.org/whl/cpu
uv pip install 'sub-align[cpu]'
# CUDA (example: cu124)
uv pip install torch --index-url https://download.pytorch.org/whl/cu124
uv pip install 'sub-align[gpu]'uv venv
uv sync --group dev # unit tests / lint (no WhisperX)
uv sync --group dev --extra align # full local alignment
uv run pytest
uv run ruff check src testsCI runs the same unit tests without the align extra. Path A/B/C coverage
also uses checked-in fixtures under tests/fixtures/ (timed WAVs + recorded
WhisperX ASR/align JSON), replayed via mocks in tests/test_fixture_replay.py.
To regenerate those fixtures locally (needs ffmpeg; recording needs WhisperX):
# 1) Timed TTS clips (clip 1 = main coverage, clip 2 = dash dialogue)
uv run --with edge-tts --with numpy --with soundfile \
scripts/generate_clip1_tts.py --clip 1
uv run --with edge-tts --with numpy --with soundfile \
scripts/generate_clip1_tts.py --clip 2
# 2) Record ASR + word-align JSON for replay tests
uv run --extra align scripts/record_whisperx_fixtures.py \
tests/fixtures/clip1_timed.wav tests/fixtures/clip2_timed.wavCompanion subtitle inputs (clip1_script.txt, clip1_drift.srt,
clip2_dialogue.srt) live beside the WAVs; edit those if the cue sheet changes.
# Timed subtitles: auto global offset + per-cue refine
sub-align media.mp4 subs.srt --language zh -o out.srt
# Lyrics (LRC): same strategy as SRT
sub-align audio.wav lyrics.lrc --language en --margin 1.0
# Untimed script: ASR windows, then force-align original lines
sub-align media.mkv script.txt --language en --model small
# Skip podcast intro/outro before aligning a script
sub-align media.mp3 script.txt --language en --trim-start 13 --trim-end 5
# Known whole-track shift (skips auto-offset ASR)
sub-align media.mkv subs.srt --language en --offset 12.5
# Audio only: transcribe + word-align into an SRT
sub-align lecture.mp4 --language en -o lecture.asr.srtAlways pass --language (e.g. en, zh) or --detect-language.
See docs/usage.md for when to use --model, --margin,
--offset, --fill-gaps, --trim-*, audio-only line limits, and a Whisper
model size / VRAM cheat sheet.
from sub_align import align_file
align_file(
media="a.mp4",
subtitle="a.srt", # omit for audio-only transcription
output="a.aligned.srt",
language="zh",
device="auto",
)- Load media as 16 kHz mono audio (via WhisperX / ffmpeg); optional
--trim-start/--trim-end. - Resolve language (
--languageor tiny-model detection). - Build search windows by input type (ASR token match for
.txt; offset + margin refine for.srt/.lrc; full ASR for media-only). - Run WhisperX forced alignment; remap word times onto original cues; trim
overlaps; optional
--fill-gaps; write.srtor.lrc.
Full diagram and tech notes: docs/pipeline.md.
uv build
uv publish # requires PyPI credentials / UV_PUBLISH_TOKENGitHub Actions publishes on tags matching v* (see .github/workflows/publish.yml).
Configure Trusted Publishing on PyPI or repository secret UV_PUBLISH_TOKEN.
MIT