Skip to content

About

Generate a transcript from a YouTube video with speaker diarization.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

yt-transcript

Transcribe YouTube videos and other URLs supported by yt-dlp locally on an Apple Silicon Mac. The command downloads the audio, splits it near silence, and runs speech recognition through MLX. Completed chunks are saved so interrupted jobs can resume. Optional speaker diarization adds anonymous speaker labels and timestamps.

Requirements

  • macOS on Apple Silicon (M1 or newer). Intel Macs, Windows, and Linux are not supported for model inference by this project.
  • Python 3.12 or 3.13; the repository defaults to 3.12.
  • uv, FFmpeg (including ffprobe), and yt-dlp.
  • Internet access to download source media and uncached model weights.
  • Free disk space for models, downloaded audio, and temporary WAV chunks. Memory requirements depend on the model and input length; no minimum RAM configuration has been validated by this project.

The default transcription model is mlx-community/Qwen3-ASR-1.7B-bf16. Speaker diarization uses OpenMOSS-Team/MOSS-Transcribe-Diarize and requires a separate model download. Models are cached locally after their first use. Speech recognition runs locally; source downloads still contact video services, and uncached models are downloaded from Hugging Face.

Installation

With Homebrew installed, open a terminal and run:

brew install uv ffmpeg
git clone https://github.com/npjonath/yt-transcript.git
cd yt-transcript
uv sync --locked
uv run yt-transcript --help

uv installs the selected Python version if needed and creates .venv/. The project dependencies include yt-dlp; uv run makes its command available. A separate brew install yt-dlp is optional.

Run commands from the cloned repository. This guide installs from source; it does not require a package published on PyPI.

Quick start

Replace VIDEO_ID with the ID of a video you can access:

uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID"

The command prints the output paths when it finishes. In ./transcripts/:

  • Title [video-id].txt contains the transcript.
  • Title [video-id].json contains source metadata, the model, settings, and segments.

For example, a plain text output might contain:

Welcome to this introduction to Python.

Today we will write our first function.

These are illustrative examples, not measured model results. Ordinary transcription provides timings for outer chunks, rather than word or sentence alignment. Only one video is processed per command; playlists are disabled.

Examples

French transcription with vocabulary hints

uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID" \
  --language French \
  --context "MLX, Qwen3-ASR, Hugging Face, Jean Dupont" \
  --output-dir ./transcripts/conferences

Language and vocabulary hints may help recognition, but output should still be reviewed.

Speaker labels and timestamps

uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID" --diarize

This produces .diarized.txt and .diarized.json files. Example text:

[00:00:00–00:00:05] S01: Welcome.
[00:00:06–00:00:10] S02: Thank you.

Labels are anonymous and do not identify people. Diarization uses approximately 60-minute parts with silence-aware boundaries. For multiple parts, speaker labels are scoped to each part, such as P01-S01 and P02-S01; these labels do not establish whether the same person appears in both parts. If the model does not return usable speaker segments, the CLI falls back to plain text for that chunk.

Authenticated media

uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID" \
  --cookies-from-browser chrome

Use your own browser session for media you are authorized to access. Browser permissions and cookie extraction support depend on yt-dlp and the browser. Never commit cookies or share unsanitized diagnostic output.

Smaller chunks and retained audio

uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID" \
  --chunk-minutes 5 --keep-work

Temporary audio and state.json live in OUTPUT_DIR/.work/. Press Ctrl+C to interrupt, then rerun the same command with the same URL, output directory, model, and transcription settings to resume completed chunks. The interrupted chunk is processed again. Changing recognition settings starts a new transcription state. By default, working files are deleted after success; --keep-work retains them.

Alternative checkpoint

uv run yt-transcript "https://www.youtube.com/watch?v=VIDEO_ID" \
  --model mlx-community/Qwen3-ASR-1.7B-bf16

--model also accepts a local model directory. Use a checkpoint compatible with the MLX Audio Qwen transcription interface. For diarization, select the checkpoint with --diarization-model instead.

Command options

Option Behavior / default
url Required HTTP(S) video URL supported by yt-dlp
-o, --output-dir PATH Output destination; ./transcripts
--language NAME Spoken language, e.g. French; automatic detection otherwise
--context TEXT Vocabulary hints
--model MODEL Qwen checkpoint or local path; default listed above
--chunk-minutes MINUTES Transcription target chunk size; 10, allowed range (0, 20]
--cookies-from-browser BROWSER Read browser cookies through yt-dlp
--keep-work Retain audio and resume state after success
--diarize Enable speaker diarization with separate output files
--diarization-model MODEL Diarization checkpoint; default listed above
-h, --help Show command help

Chunk sizes are targets: silence selection and merging a short final chunk can make chunks longer. --chunk-minutes controls ordinary transcription; diarization uses its own fixed target.

Troubleshooting

  • Missing FFmpeg / ffprobe: run brew install ffmpeg and check that Homebrew is on your PATH. Use uv run so the project's yt-dlp command is available.
  • Video download fails: check access in your browser, try the cookies option when appropriate, and inspect yt-dlp's error. Website changes can require a dependency update; contributors should update and commit uv.lock together.
  • Model loading fails or memory is exhausted: verify that the terminal uses native Apple Silicon Python, ensure sufficient disk space, close other memory intensive apps, or use a compatible smaller checkpoint. Smaller outer chunks reduce audio size but do not reduce the model's weight memory requirement.
  • Interrupted job: rerun the same command; keep the output's .work/ directory.
  • Inaccurate text or speaker labels: provide language/context hints and review against the audio. Automated transcription can omit or invent content.

The default output directory is ignored by Git. Add custom output directories to .gitignore before committing. Only download and process media you have permission to use, respecting the source service's terms and applicable rights.

Development and community

uv sync --locked
uv run pytest
uv run ruff check .
uv run ruff format --check .
uv build

The CI runs unit tests, lint, formatting, CLI help, and package construction on Python 3.12 and 3.13. Unit tests mock inference and media commands; they do not validate real model accuracy or an end-to-end download/transcription.

See Contributing, the Code of Conduct, and the Security policy. Use GitHub issues for bugs and feature requests, and pull requests for focused improvements.

License

The project code is available under the MIT License. Dependencies, model weights, and source media retain their own licenses and terms.

About

Generate a transcript from a YouTube video with speaker diarization.

Resources

Code of conduct

Contributing

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages