Transcribe YouTube videos and other URLs supported by yt-dlp locally on an Apple Silicon Mac. The command downloads the audio, splits it near silence, and runs speech recognition through MLX. Completed chunks are saved so interrupted jobs can resume. Optional speaker diarization adds anonymous speaker labels and timestamps.
- macOS on Apple Silicon (M1 or newer). Intel Macs, Windows, and Linux are not supported for model inference by this project.
- Python 3.12 or 3.13; the repository defaults to 3.12.
uv, FFmpeg (includingffprobe), and yt-dlp.- Internet access to download source media and uncached model weights.
- Free disk space for models, downloaded audio, and temporary WAV chunks. Memory requirements depend on the model and input length; no minimum RAM configuration has been validated by this project.
The default transcription model is mlx-community/Qwen3-ASR-1.7B-bf16.
Speaker diarization uses OpenMOSS-Team/MOSS-Transcribe-Diarize and requires a
separate model download. Models are cached locally after their first use.
Speech recognition runs locally; source downloads still contact video services,
and uncached models are downloaded from Hugging Face.
With Homebrew installed, open a terminal and run:
brew install uv ffmpeg
git clone https://github.com/npjonath/yt-transcript.git
cd yt-transcript
uv sync --locked
uv run yt-transcript --helpuv installs the selected Python version if needed and creates .venv/.
The project dependencies include yt-dlp; uv run makes its command available.
A separate brew install yt-dlp is optional.
Run commands from the cloned repository. This guide installs from source; it does not require a package published on PyPI.
Replace VIDEO_ID with the ID of a video you can access:
uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID"The command prints the output paths when it finishes. In ./transcripts/:
Title [video-id].txtcontains the transcript.Title [video-id].jsoncontains source metadata, the model, settings, and segments.
For example, a plain text output might contain:
Welcome to this introduction to Python.
Today we will write our first function.
These are illustrative examples, not measured model results. Ordinary transcription provides timings for outer chunks, rather than word or sentence alignment. Only one video is processed per command; playlists are disabled.
uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID" \
--language French \
--context "MLX, Qwen3-ASR, Hugging Face, Jean Dupont" \
--output-dir ./transcripts/conferencesLanguage and vocabulary hints may help recognition, but output should still be reviewed.
uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID" --diarizeThis produces .diarized.txt and .diarized.json files. Example text:
[00:00:00–00:00:05] S01: Welcome.
[00:00:06–00:00:10] S02: Thank you.
Labels are anonymous and do not identify people. Diarization uses approximately
60-minute parts with silence-aware boundaries. For multiple parts, speaker labels
are scoped to each part, such as P01-S01 and P02-S01; these labels do not
establish whether the same person appears in both parts. If the model does not
return usable speaker segments, the CLI falls back to plain text for that chunk.
uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID" \
--cookies-from-browser chromeUse your own browser session for media you are authorized to access. Browser permissions and cookie extraction support depend on yt-dlp and the browser. Never commit cookies or share unsanitized diagnostic output.
uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID" \
--chunk-minutes 5 --keep-workTemporary audio and state.json live in OUTPUT_DIR/.work/. Press Ctrl+C to
interrupt, then rerun the same command with the same URL, output directory, model,
and transcription settings to resume completed chunks. The interrupted chunk
is processed again. Changing recognition settings starts a new transcription state.
By default, working files are deleted after success; --keep-work retains them.
uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID" \
--model mlx-community/Qwen3-ASR-1.7B-bf16--model also accepts a local model directory. Use a checkpoint compatible with
the MLX Audio Qwen transcription interface. For diarization, select the checkpoint
with --diarization-model instead.
| Option | Behavior / default |
|---|---|
url |
Required HTTP(S) video URL supported by yt-dlp |
-o, --output-dir PATH |
Output destination; ./transcripts |
--language NAME |
Spoken language, e.g. French; automatic detection otherwise |
--context TEXT |
Vocabulary hints |
--model MODEL |
Qwen checkpoint or local path; default listed above |
--chunk-minutes MINUTES |
Transcription target chunk size; 10, allowed range (0, 20] |
--cookies-from-browser BROWSER |
Read browser cookies through yt-dlp |
--keep-work |
Retain audio and resume state after success |
--diarize |
Enable speaker diarization with separate output files |
--diarization-model MODEL |
Diarization checkpoint; default listed above |
-h, --help |
Show command help |
Chunk sizes are targets: silence selection and merging a short final chunk can
make chunks longer. --chunk-minutes controls ordinary transcription;
diarization uses its own fixed target.
- Missing FFmpeg / ffprobe: run
brew install ffmpegand check that Homebrew is on yourPATH. Useuv runso the project's yt-dlp command is available. - Video download fails: check access in your browser, try the cookies option
when appropriate, and inspect yt-dlp's error. Website changes can require a
dependency update; contributors should update and commit
uv.locktogether. - Model loading fails or memory is exhausted: verify that the terminal uses native Apple Silicon Python, ensure sufficient disk space, close other memory intensive apps, or use a compatible smaller checkpoint. Smaller outer chunks reduce audio size but do not reduce the model's weight memory requirement.
- Interrupted job: rerun the same command; keep the output's
.work/directory. - Inaccurate text or speaker labels: provide language/context hints and review against the audio. Automated transcription can omit or invent content.
The default output directory is ignored by Git. Add custom output directories to
.gitignore before committing. Only download and process media you have permission
to use, respecting the source service's terms and applicable rights.
uv sync --locked
uv run pytest
uv run ruff check .
uv run ruff format --check .
uv buildThe CI runs unit tests, lint, formatting, CLI help, and package construction on Python 3.12 and 3.13. Unit tests mock inference and media commands; they do not validate real model accuracy or an end-to-end download/transcription.
See Contributing, the Code of Conduct, and the Security policy. Use GitHub issues for bugs and feature requests, and pull requests for focused improvements.
The project code is available under the MIT License. Dependencies, model weights, and source media retain their own licenses and terms.