diff --git a/README.md b/README.md index 00be66e..193330c 100644 --- a/README.md +++ b/README.md @@ -1,20 +1,28 @@ # spectrogram-renderer -A Lambda function that renders one spectrogram PNG from one hydrophone audio segment. [Orcasite](https://github.com/orcasound/orcasite) invokes it by name from [`Orcasite.Radio.AwsClient`](https://github.com/orcasound/orcasite/blob/main/server/lib/orcasite/radio/aws_client.ex), once per `AudioImage`, passing the S3 location of the segment and where to put the image. It has no other entry point. +A Lambda function that renders one spectrogram PNG from one hydrophone audio segment. [Orcasite](https://github.com/orcasound/orcasite) invokes it by name, once per `AudioImage`, and stores what it answers as the image's parameters. It has no other entry point. -- [`core/app.py`](core/app.py): the handler. Downloads the segment, decodes it with ffmpeg, renders with [`core/spectrogram_generator.py`](core/spectrogram_generator.py), uploads the PNG. -- [`core/Dockerfile`](core/Dockerfile): the image, built on the AWS Lambda Python base with a static ffmpeg. -- [`template.yaml`](template.yaml): the function's memory, timeout and IAM policy, as a SAM template. Stack name (`audio-viz`, kept from when this lived in orcasite) and region are in [`samconfig.toml`](samconfig.toml). +The whole render is one ffmpeg command ([`core/app.py`](core/app.py), `render`): `showspectrumpic` draws the spectrogram as intensity and `pseudocolor` applies matplotlib's colour ramp. The handler around it fetches the audio, probes its sample rate and duration to size the image, and stores the PNG. There is no Python beyond the standard library, so the image is the Lambda base plus a static ffmpeg, and a cold start costs about a second. -## Cold starts +## Contract -A new Lambda container starts with an empty `/tmp`, where numba, librosa and matplotlib keep their caches, so the first render in a container recompiles librosa's numba functions and rebuilds the font list. The Dockerfile runs [`warm.py`](core/warm.py) once at build time and keeps the caches it leaves in the image; [`app.py`](core/app.py) copies them into `/tmp` before importing those libraries. If you add a library that caches on first use, give it a directory under `/tmp` in the Dockerfile and add it to the same `mv`. +```json +{ + "audio_url": "https://… (presigned GET)", + "image_url": "https://… (presigned PUT), or null to render and discard", + "parameters": {"n_fft": 1024, "hop_length": 256, "fmin": 1, "fmax": 15000, "db_min": 10, "db_max": 80, "cmap": "viridis", "frequency_scale": "linear"} +} +``` + +The function reads and writes only the URLs it is given and knows nothing about what is behind them; the caller decides where audio and images live and presigns accordingly. Every parameter is optional and defaults to what orcasite has always used, so images stay consistent with the existing archive. The response gives `image_size`, `sample_rate`, `width`, `height` and the `parameters` applied, plus the field names orcasite has stored since the first renderer. + +The image is `n_fft/2` pixels high and one pixel wide per `hop_length` samples, as the previous renderer's STFT was. `db_min`/`db_max` are the dB window relative to an amplitude of 0.01, the convention the archive was drawn with; `DB_OFFSET` in `app.py` converts it to ffmpeg's full-scale dB and was fitted against a production image, whose luminance distribution the new render reproduces to within 0.01. -To check a deployed function, find `platform.report` records in its CloudWatch log (the template sets `LogFormat: JSON`). Cold invocations carry `initDurationMs`; their `durationMs + initDurationMs` should be within a couple of seconds of a warm invocation's `durationMs`. Lambda's `Duration` metric excludes init time, so it alone understates a cold start. +Until orcasite has switched, the previous shape (`audio_bucket`/`audio_key`, `image_bucket`/`image_key`, read and written with boto3) is also accepted; [`template.yaml`](template.yaml) keeps the S3 policy for it and both go once no caller uses them. ## Testing -[`tests/smoke.py`](tests/smoke.py) renders a synthetic MPEG-TS clip inside the built image, run as an unprivileged user with a root-owned empty `/tmp` and the CPU and memory of the deployed tier, which is how Lambda runs it. It fails if the first render is slow enough to suggest the caches were not restored. [CI](.github/workflows/ci.yml) runs it on every pull request; locally: +[`tests/smoke.py`](tests/smoke.py) renders a synthetic AAC/MPEG-TS clip through the handler inside the built image, run as an unprivileged user with a root-owned empty `/tmp` and the CPU and memory of the deployed tier, which is how Lambda runs it. [CI](.github/workflows/ci.yml) runs it on every pull request; locally: ```bash docker build -t spectrogram-renderer core @@ -22,20 +30,15 @@ docker run --rm --user 1000:1000 --tmpfs /tmp:uid=0,gid=0,mode=1777 --cpus 0.58 --entrypoint python -v "$PWD/tests/smoke.py:/var/task/smoke.py" spectrogram-renderer smoke.py ``` -That rehearses the application, not the sandbox: Lambda's Runtime Interface Emulator in the base image reproduces the invoke API, not the filesystem rules. Anything that touches `/tmp`'s own metadata or relies on image contents under `/tmp` will pass locally on a plain tmpfs and fail deployed, hence the `--user` and `--tmpfs` flags above. +That rehearses the application, not the sandbox: Lambda's Runtime Interface Emulator in the base image reproduces the invoke API, not the filesystem rules, hence the `--user` and `--tmpfs` flags. ## Deploying -Merging to `main` deploys: the [CI workflow](.github/workflows/ci.yml) runs the smoke test, then `sam build && sam deploy` under the `production` environment, assuming the `spectrogram-renderer-deploy` IAM role through GitHub's OIDC provider. That role can change only this stack, its ECR repository and its function role. +Merging to `main` deploys: the [CI workflow](.github/workflows/ci.yml) runs the smoke test, then `sam build && sam deploy` under the `production` environment, assuming the `spectrogram-renderer-deploy` IAM role through GitHub's OIDC provider. That role can change only this stack (`audio-viz`, the name kept from when this lived in orcasite), its ECR repository and its function role. -To deploy from a machine instead, with an AWS profile for the Orcasound account and the [SAM CLI](https://docs.aws.amazon.com/serverless-application-model/latest/developerguide/serverless-sam-cli-install.html): - -```bash -sam build -AWS_PROFILE=orcasound sam deploy -``` +To deploy from a machine instead, with an AWS profile for the Orcasound account and the [SAM CLI](https://docs.aws.amazon.com/serverless-application-model/latest/developerguide/serverless-sam-cli-install.html): `sam build && AWS_PROFILE=orcasound sam deploy`. -To invoke the deployed function on a real segment without writing anything (`image_key` null skips the upload): +To invoke the deployed function on a real segment without writing anything ([`events/spectrogram_job.json`](events/spectrogram_job.json) reads from the public audio bucket and has `image_url` null): ```bash aws lambda invoke --function-name "$(aws lambda list-functions --query "Functions[?contains(FunctionName,'AudioVizFunction')].FunctionName | [0]" --output text)" \ diff --git a/core/Dockerfile b/core/Dockerfile index 08f2817..34d5b8b 100644 --- a/core/Dockerfile +++ b/core/Dockerfile @@ -1,35 +1,16 @@ FROM public.ecr.aws/lambda/python:3.14 +# ffmpeg does all the work; the handler is standard library plus the boto3 +# the base image ships, so there is nothing to pip install and no cache to +# warm. A static build keeps the image at the base plus one binary. ARG FFMPEG_VERSION=ffmpeg-7.0.2-amd64-static +RUN dnf install -y tar xz \ + && curl -sO https://johnvansickle.com/ffmpeg/releases/${FFMPEG_VERSION}.tar.xz \ + && tar -xf ${FFMPEG_VERSION}.tar.xz \ + && mv ${FFMPEG_VERSION}/ffmpeg ${FFMPEG_VERSION}/ffprobe /usr/local/bin/ \ + && rm -rf ${FFMPEG_VERSION} ${FFMPEG_VERSION}.tar.xz \ + && dnf clean all -# /tmp is the only writeable directory at runtime, but Lambda mounts a fresh, -# empty one for every new container, so anything put there at build time is -# gone by the first invocation. The caches these libraries build on first use -# are therefore built into the image (see warm.py below) and copied into /tmp -# by app.py when a container starts. -ENV CACHE_SEED=/opt/cache-seed -ENV MPLCONFIGDIR=/tmp/matplotlib -ENV LIBROSA_CACHE_DIR=/tmp/librosa_cache -ENV NUMBA_CACHE_DIR=/tmp/numba_cache +COPY app.py ./ -RUN dnf install -y tar gzip xz - -RUN curl -O https://johnvansickle.com/ffmpeg/releases/${FFMPEG_VERSION}.tar.xz -RUN tar -xvf "${FFMPEG_VERSION}.tar.xz" -RUN mv ${FFMPEG_VERSION}/ffmpeg /usr/local/bin -RUN mv ${FFMPEG_VERSION}/ffprobe /usr/local/bin - -COPY app.py requirements.txt spectrogram_generator.py warm.py ./ -RUN mkdir -p data - -RUN python3 -m pip install -r requirements.txt -t . - -# Render one clip so that every cache a render populates exists before the -# first real invocation, then keep the result where app.py can find it. -RUN mkdir -p ${MPLCONFIGDIR} ${LIBROSA_CACHE_DIR} ${NUMBA_CACHE_DIR} \ - && python3 warm.py \ - && mkdir -p ${CACHE_SEED} \ - && mv ${MPLCONFIGDIR} ${LIBROSA_CACHE_DIR} ${NUMBA_CACHE_DIR} ${CACHE_SEED}/ - -# Command can be overwritten by providing a different command in the template directly. CMD ["app.lambda_handler"] diff --git a/core/__init__.py b/core/__init__.py deleted file mode 100644 index e69de29..0000000 diff --git a/core/app.py b/core/app.py index 4b0373e..fbbcecf 100644 --- a/core/app.py +++ b/core/app.py @@ -1,159 +1,145 @@ -import os -import shutil - -# Every new Lambda container starts with an empty /tmp, where numba, librosa -# and matplotlib keep their caches. Without them the first render recompiles -# librosa's numba functions and rebuilds the font list, which at 512 MB took -# longer than the function's timeout. The Dockerfile builds the caches into -# the image; restore them before those libraries load and look for them. -# -# Copy each cache directory separately: copying onto /tmp itself makes -# copytree set /tmp's permissions, which Lambda refuses. A failure here only -# costs a slow first render, so it must never fail the invocation. -_cache_seed = os.environ.get("CACHE_SEED") -if _cache_seed and os.path.isdir(_cache_seed): - for _name in os.listdir(_cache_seed): - try: - shutil.copytree( - os.path.join(_cache_seed, _name), - os.path.join("/tmp", _name), - dirs_exist_ok=True, - ) - except OSError as e: - print(f"Could not restore cache {_name}: {e}") - -from matplotlib.pyplot import imshow -from spectrogram_generator import SpectrogramGenerator -from subprocess import check_output -from typing import TypedDict, Optional -import boto3 -import io -import json -import librosa -import matplotlib -import sys - -SpectrogramJob = TypedDict( - "SpectrogramJob", - { - "id": str, - "audio_bucket": str, - "audio_key": str, - "sample_rate": int, - "image_key": Optional[str], - "image_bucket": Optional[str], - }, -) - - -def lambda_handler(event, context): - """Sample pure Lambda function - - Parameters - ---------- - event: dict, required - API Gateway Lambda Proxy Input Format - - Event doc: https://docs.aws.amazon.com/apigateway/latest/developerguide/set-up-lambda-proxy-integrations.html#api-gateway-simple-proxy-for-lambda-input-format - - context: object, required - Lambda Context runtime methods and attributes +"""Renders one spectrogram PNG from one audio segment, with ffmpeg. - Context doc: https://docs.aws.amazon.com/lambda/latest/dg/python-context-object.html +The request says where to read the audio and where to write the image, as +URLs; the function knows nothing about what is behind them. Render +parameters are explicit, with the defaults orcasite has always used, and +the response echoes the ones that were applied. - Returns - ------ - API Gateway Lambda Proxy Output Format: dict + { + "audio_url": "https://... (presigned GET)", + "image_url": "https://... (presigned PUT), or null to render and discard", + "parameters": {"fmax": 15000, ...} # optional, see DEFAULTS + } - Return doc: https://docs.aws.amazon.com/apigateway/latest/developerguide/set-up-lambda-proxy-integrations.html - """ - result = make_spectrogram(event) +For one release the previous shape is also accepted: `audio_bucket` and +`audio_key`, `image_bucket` and `image_key`, read and written with boto3. +""" - return {"status": 200, **result} +import os +import subprocess +import tempfile +import urllib.request +import wave + +DEFAULTS = { + "n_fft": 1024, + "hop_length": 256, + "fmin": 1, + "fmax": 15000, + "db_min": 10, + "db_max": 80, + "cmap": "viridis", + "frequency_scale": "linear", +} + +# ffmpeg's dB scale is relative to full scale, while the dB window orcasite's +# images were drawn with is relative to an amplitude of 0.01 in librosa's +# STFT. This offset lines the two up: it was fitted against a production +# image, whose luminance distribution it reproduces to within 0.01. +DB_OFFSET = -122 + +FSCALE = {"linear": "lin", "log": "log"} -def make_spectrogram( - job: SpectrogramJob, store_audio=False, show_image=False, store_image=False -): - print(f"Received job: {json.dumps(job)}") - local_path = f"/tmp/{job["id"]}" - s3 = boto3.client("s3") +def lambda_handler(event, context): + return {"status": 200, **make_spectrogram(event)} - if not os.path.exists(local_path): - response = s3.get_object(Bucket=job["audio_bucket"], Key=job["audio_key"]) - audio_bytes = response["Body"].read() - with open(local_path, "wb") as file: - file.write(audio_bytes) - sample_rate = job.get("sample_rate") - if sample_rate is None: - metadata = get_audio_metadata(local_path) - sample_rate = int(metadata["streams"][0]["sample_rate"]) +def make_spectrogram(job): + params = {**DEFAULTS, **(job.get("parameters") or {})} - audio, sr = load_audio(local_path, sample_rate) + with tempfile.TemporaryDirectory(dir="/tmp") as tmp: + audio_path = os.path.join(tmp, "audio") + wav_path = os.path.join(tmp, "audio.wav") + fetch_audio(job, audio_path) + sample_rate, samples = decode(audio_path, wav_path) + width = 1 + samples // params["hop_length"] + height = params["n_fft"] // 2 + png = render(wav_path, params, width, height) - params = {"linear": True, "fmin": 1, "fmax": 15000, "cmap": "viridis"} - - generator = SpectrogramGenerator( - sr, n_fft=1024, hop_length=256, db_range=(10, 80), **params - ) - spectrogram = generator(audio) - if show_image: - matplotlib.pyplot.axis("off") - imshow(spectrogram) - - # Store spectrogram image either as a file or in-memory - image_size = None - image_store = f"/tmp/{job["id"]}.png" if store_image else io.BytesIO() - matplotlib.image.imsave(image_store, spectrogram) - if store_image: - image_size = os.path.getsize(image_store) - else: - image_store.seek(0) - image_size = sys.getsizeof(image_store) - image_store.seek(0) - - if job["image_key"] is not None and job["image_bucket"] is not None: - image_content = open(image_store, "rb") if store_image else image_store - s3.put_object( - Bucket=job["image_bucket"], Key=job["image_key"], Body=image_content - ) + store_image(job, png) return { + "image_size": len(png), "sample_rate": sample_rate, - "image_size": image_size, - "frequency_spacing": "linear" if params["linear"] else "log", + "width": width, + "height": height, + "parameters": params, + # The names orcasite has stored since the first renderer. + "frequency_spacing": params["frequency_scale"], "freq_min": params["fmin"], "freq_max": params["fmax"], "color_map": params["cmap"], } -def load_audio(local_path, sample_rate): - """Decodes the segment with ffmpeg and hands librosa a WAV. +def render(audio_path, params, width, height): + """One ffmpeg invocation: spectrogram as intensity, then the colour map. - The hydrophones stream AAC in MPEG-TS, which libsndfile does not read, and - librosa 1.0 dropped the ffmpeg fallback that used to cover it. + showspectrumpic's own colour maps start at black, so the image is drawn + in grey and matplotlib's ramp applied with pseudocolor, which carries the + same tables. Its window size follows from the height (n_fft/2 bins). """ - wav = check_output( - ["ffmpeg", "-hide_banner", "-loglevel", "error", "-i", local_path, "-f", "wav", "-"] + spectrum = ":".join( + [ + f"s={width}x{height}", + "mode=combined", + "color=intensity", + "scale=log", + f"fscale={FSCALE[params['frequency_scale']]}", + "win_func=hann", + f"start={params['fmin']}", + f"stop={params['fmax']}", + f"limit={params['db_max'] + DB_OFFSET}", + f"drange={params['db_max'] - params['db_min']}", + "legend=0", + ] ) - return librosa.load(io.BytesIO(wav), sr=sample_rate) + graph = f"showspectrumpic={spectrum},format=gbrp,pseudocolor=preset={params['cmap']}" + return subprocess.run( + ["ffmpeg", "-hide_banner", "-loglevel", "error", "-i", audio_path, + "-lavfi", graph, "-f", "image2pipe", "-c:v", "png", "-"], + check=True, + capture_output=True, + ).stdout -def get_audio_metadata(local_path): - result = check_output( - [ - "ffprobe", - "-hide_banner", - "-loglevel", - "panic", - "-show_format", - "-show_streams", - "-of", - "json", - local_path, - ] +def decode(audio_path, wav_path): + """Decodes the segment to WAV and returns its sample rate and length. + + The image is sized from the samples actually decoded: the container's + duration undercounts an AAC stream by a frame or two, which would make + each tile a few pixels narrower than the archive's. + """ + subprocess.run( + ["ffmpeg", "-hide_banner", "-loglevel", "error", "-i", audio_path, "-f", "wav", wav_path], + check=True, + capture_output=True, ) + with wave.open(wav_path, "rb") as wav: + return wav.getframerate(), wav.getnframes() + - return json.loads(result) +def fetch_audio(job, path): + if job.get("audio_url"): + with urllib.request.urlopen(job["audio_url"]) as response, open(path, "wb") as file: + file.write(response.read()) + else: + import boto3 + + boto3.client("s3").download_file(job["audio_bucket"], job["audio_key"].lstrip("/"), path) + + +def store_image(job, png): + if job.get("image_url"): + request = urllib.request.Request( + job["image_url"], data=png, method="PUT", headers={"Content-Type": "image/png"} + ) + with urllib.request.urlopen(request): + pass + elif job.get("image_bucket") and job.get("image_key"): + import boto3 + + boto3.client("s3").put_object( + Bucket=job["image_bucket"], Key=job["image_key"].lstrip("/"), Body=png, ContentType="image/png" + ) diff --git a/core/requirements.txt b/core/requirements.txt deleted file mode 100644 index cfbc5eb..0000000 --- a/core/requirements.txt +++ /dev/null @@ -1,5 +0,0 @@ -boto3 -librosa -matplotlib -numpy -opencv-python-headless diff --git a/core/sandbox.ipynb b/core/sandbox.ipynb deleted file mode 100644 index f47aabe..0000000 --- a/core/sandbox.ipynb +++ /dev/null @@ -1,317 +0,0 @@ -{ - "cells": [ - { - "cell_type": "markdown", - "metadata": {}, - "source": [ - "# Spectrogram sandbox\n", - "\n", - "This notebook is for experimenting with spectrogram parameters.\n", - "\n", - "Ensure you have `ffprobe` available to your computer via your computer's package manager (`brew install ffprobe` on MacOS or `sudo apt-get install ffprobe` on Ubuntu, etc.)" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "orca_call_audio_url = \"https://s3-us-west-2.amazonaws.com/streaming-orcasound-net/rpi_orcasound_lab/hls/1719430219/live322.ts\"\n", - "fish_grunt_audio_url = \"https://s3-us-west-2.amazonaws.com/streaming-orcasound-net/rpi_orcasound_lab/hls/1723491018/live516.ts\"\n", - "ferry_vessel_audio_url = \"https://s3-us-west-2.amazonaws.com/streaming-orcasound-net/rpi_north_sjc/hls/1723100418/live6160.ts\"\n", - "hydrophone_hum_water_splashes_audio_url = \"https://s3-us-west-2.amazonaws.com/streaming-orcasound-net/rpi_north_sjc/hls/1718175617/live7424.ts\"" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "%pip install boto3\n", - "%pip install imutils \n", - "%pip install librosa\n", - "%pip install matplotlib\n", - "%pip install numpy<2.0\n", - "%pip install opencv-python \n", - "%pip install requests" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "# Spectrogram generator via: https://github.com/kylemcdonald/AudioNotebooks/blob/master/Generating%20Spectrograms.ipynb\n", - "\n", - "# Copyright (c) 2016- Kyle McDonald\n", - "# Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the \"Software\"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions:\n", - "# The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.\n", - "# THE SOFTWARE IS PROVIDED \"AS IS\", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE.\n", - "\n", - "import librosa\n", - "import matplotlib.pyplot as plt\n", - "import numpy as np\n", - "import cv2\n", - "\n", - "class SpectrogramGenerator:\n", - " \"\"\"Create a spectrogram generator.\"\"\"\n", - " \n", - " def __init__(self, sr,\n", - " n_fft=1024,\n", - " linear=False,\n", - " hop_length=None,\n", - " db_ref=0.01,\n", - " db_range=None,\n", - " fmin=None,\n", - " fmax=None,\n", - " cmap=None,\n", - " window='hann',\n", - " drop_nyquist=True,\n", - " flipud=True):\n", - " \"\"\"\n", - " Args:\n", - " sr (int): sample rate\n", - " n_fft (int): size of STFT window in samples\n", - " linear (bool): use linear instead of logarithmic frequency spacing\n", - " hop_length (int): samples between STFT frames\n", - " db_ref (float): `ref` parameter when converting amplitude to dB\n", - " db_range (tuple): min and max dB for output\n", - " fmin (float): minimum frequency to output (hertz)\n", - " fmax (float): maximum frequency to output (hertz)\n", - " cmap (str): matplotlib color map to use for output\n", - " window (str): windowing function for STFT\n", - " drop_nyquist (bool): if True, output has exactly `n_fft//2` bins\n", - " flipud (bool): whether to flip the STFT output\n", - " \"\"\"\n", - "\n", - " self.n_fft = n_fft\n", - " self.hop_length = hop_length\n", - " self.db_ref = db_ref\n", - " self.db_range = db_range\n", - " self.cmap = cmap\n", - " self.window = window\n", - " self.drop_nyquist = drop_nyquist\n", - " \n", - " if hop_length is None:\n", - " self.hop_length = n_fft // 4\n", - " \n", - " if cmap is not None:\n", - " self.cmap = plt.get_cmap(cmap)\n", - " \n", - " bin_count = (n_fft // 2) + 1 \n", - " if drop_nyquist:\n", - " bin_count -= 1\n", - " nyquist = sr // 2\n", - " if fmin == None:\n", - " min_bin = 1\n", - " else:\n", - " min_bin = (fmin / nyquist) * (bin_count - 1)\n", - " if fmax == None:\n", - " max_bin = bin_count - 1\n", - " else:\n", - " max_bin = (fmax / nyquist) * (bin_count - 1)\n", - " bins = np.arange(bin_count)\n", - " if linear:\n", - " y_remap = (bins + min_bin) / ((bin_count - 1) / (max_bin - min_bin))\n", - " else:\n", - " scale_factor = bin_count / np.log10(max_bin / min_bin)\n", - " y_remap = min_bin * 10 ** (bins / scale_factor)\n", - " if flipud:\n", - " y_remap = y_remap[::-1]\n", - " self.mapy_base = y_remap.reshape(-1,1).astype(np.float32)\n", - " \n", - " def __call__(self, audio):\n", - " \"\"\"Generate a spectrogram for visualization.\n", - " \n", - " Returns:\n", - " np.array of type np.uint8 if `db_range` or `cmap` is set, otherwise np.float32\n", - " \"\"\"\n", - " stft = librosa.stft(audio,\n", - " n_fft=self.n_fft,\n", - " hop_length=self.hop_length,\n", - " window=self.window)\n", - " if self.drop_nyquist:\n", - " stft = stft[:-1]\n", - " db = librosa.amplitude_to_db(np.abs(stft), ref=self.db_ref)\n", - " mapx = np.arange(db.shape[1], dtype=np.float32).reshape(1,-1).repeat(db.shape[0], axis=0)\n", - " mapy = self.mapy_base.repeat(db.shape[1], axis=1)\n", - " out = cv2.remap(db, mapx, mapy, cv2.INTER_LANCZOS4)\n", - " \n", - " if self.db_range is not None:\n", - " out -= self.db_range[0]\n", - " out /= self.db_range[1] - self.db_range[0]\n", - " \n", - " if self.cmap is not None:\n", - " if self.db_range is None:\n", - " out -= out.min()\n", - " out /= out.max()\n", - " out = self.cmap(out)[...,:3]\n", - " \n", - " if self.cmap is not None or self.db_range is not None:\n", - " out = np.clip(out * 256, 0, 255).astype(np.uint8)\n", - " \n", - " return out\n", - " " - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "from matplotlib.pyplot import imshow\n", - "from spectrogram_generator import SpectrogramGenerator\n", - "from subprocess import check_output\n", - "from typing import TypedDict, Optional\n", - "import boto3\n", - "import io\n", - "import json\n", - "import librosa\n", - "import matplotlib\n", - "import os.path\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "def get_audio_metadata(local_path):\n", - " result = check_output(['ffprobe',\n", - " '-hide_banner', '-loglevel', 'panic',\n", - " '-show_format',\n", - " '-show_streams',\n", - " '-of',\n", - " 'json', local_path])\n", - "\n", - " return json.loads(result)" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "def spectrogram(job, show_image=True, store_image=True):\n", - " local_path = f\"{job[\"id\"]}.ts\"\n", - " s3 = boto3.client(\"s3\")\n", - "\n", - " if not os.path.exists(local_path):\n", - " response = s3.get_object(Bucket=job[\"audio_bucket\"], Key=job[\"audio_key\"])\n", - " audio_bytes = response[\"Body\"].read()\n", - " with open(local_path, \"wb\") as file:\n", - " file.write(audio_bytes)\n", - "\n", - " sample_rate = job.get(\"sample_rate\")\n", - " if sample_rate is None:\n", - " metadata = get_audio_metadata(local_path)\n", - " sample_rate = int(metadata[\"streams\"][0][\"sample_rate\"])\n", - "\n", - " audio, sr = librosa.load(local_path, sr=sample_rate)\n", - "\n", - " generator = SpectrogramGenerator(sr,\n", - " linear=False,\n", - " n_fft=1024,\n", - " # fmin=1,\n", - " # fmax=20000,\n", - " hop_length=256,\n", - " db_range=(10, 80),\n", - " cmap='inferno')\n", - " spectrogram = generator(audio)\n", - " if show_image:\n", - " matplotlib.pyplot.axis('off')\n", - " imshow(spectrogram)\n", - "\n", - " # Store spectrogram image either as a file or in-memory\n", - " image_store = f\"{job[\"id\"]}.png\" if store_image else io.BytesIO()\n", - " matplotlib.image.imsave(image_store, spectrogram)\n", - " if not store_image:\n", - " image_store.seek(0)" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "from urllib.parse import urlparse\n", - "\n", - "def s3_params(url):\n", - " path_parts = urlparse(url).path.split(\"/\")\n", - " return {\"bucket\": path_parts.pop(1), \"key\": \"/\".join(path_parts)[1:]}\n" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "jobs = [\n", - " {\n", - " \"id\": \"orca_call\",\n", - " \"audio_bucket\": s3_params(orca_call_audio_url)[\"bucket\"],\n", - " \"audio_key\": s3_params(orca_call_audio_url)[\"key\"],\n", - " },\n", - " {\n", - " \"id\": \"fish_grunt\",\n", - " \"audio_bucket\": s3_params(fish_grunt_audio_url)[\"bucket\"],\n", - " \"audio_key\": s3_params(fish_grunt_audio_url)[\"key\"],\n", - " },\n", - " {\n", - " \"id\": \"ferry_vessel\",\n", - " \"audio_bucket\": s3_params(ferry_vessel_audio_url)[\"bucket\"],\n", - " \"audio_key\": s3_params(ferry_vessel_audio_url)[\"key\"],\n", - " },\n", - " {\n", - " \"id\": \"water_splash\",\n", - " \"audio_bucket\": s3_params(hydrophone_hum_water_splashes_audio_url)[\"bucket\"],\n", - " \"audio_key\": s3_params(hydrophone_hum_water_splashes_audio_url)[\"key\"],\n", - " },\n", - "]\n", - "orca = jobs[0]\n", - "fish = jobs[1]\n", - "ferry_vessel = jobs[2]\n", - "water_splash = jobs[3]\n", - "jobs" - ] - }, - { - "cell_type": "code", - "execution_count": null, - "metadata": {}, - "outputs": [], - "source": [ - "spectrogram(orca)" - ] - } - ], - "metadata": { - "kernelspec": { - "display_name": "Python 3", - "language": "python", - "name": "python3" - }, - "language_info": { - "codemirror_mode": { - "name": "ipython", - "version": 3 - }, - "file_extension": ".py", - "mimetype": "text/x-python", - "name": "python", - "nbconvert_exporter": "python", - "pygments_lexer": "ipython3", - "version": "3.12.4" - } - }, - "nbformat": 4, - "nbformat_minor": 2 -} diff --git a/core/spectrogram_generator.py b/core/spectrogram_generator.py deleted file mode 100644 index 4f97252..0000000 --- a/core/spectrogram_generator.py +++ /dev/null @@ -1,110 +0,0 @@ -# Spectrogram generator via: https://github.com/kylemcdonald/AudioNotebooks/blob/master/Generating%20Spectrograms.ipynb - -# Copyright (c) 2016- Kyle McDonald -# Permission is hereby granted, free of charge, to any person obtaining a copy of this software and associated documentation files (the "Software"), to deal in the Software without restriction, including without limitation the rights to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the Software, and to permit persons to whom the Software is furnished to do so, subject to the following conditions: -# The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software. -# THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY, FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM, OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE SOFTWARE. - -import librosa -import matplotlib.pyplot as plt -import numpy as np -import cv2 - -class SpectrogramGenerator: - """Create a spectrogram generator.""" - - def __init__(self, sr, - n_fft=1024, - linear=False, - hop_length=None, - db_ref=0.01, - db_range=None, - fmin=None, - fmax=None, - cmap=None, - window='hann', - drop_nyquist=True, - flipud=True): - """ - Args: - sr (int): sample rate - n_fft (int): size of STFT window in samples - linear (bool): use linear instead of logarithmic frequency spacing - hop_length (int): samples between STFT frames - db_ref (float): `ref` parameter when converting amplitude to dB - db_range (tuple): min and max dB for output - fmin (float): minimum frequency to output (hertz) - fmax (float): maximum frequency to output (hertz) - cmap (str): matplotlib color map to use for output - window (str): windowing function for STFT - drop_nyquist (bool): if True, output has exactly `n_fft//2` bins - flipud (bool): whether to flip the STFT output - """ - - self.n_fft = n_fft - self.hop_length = hop_length - self.db_ref = db_ref - self.db_range = db_range - self.cmap = cmap - self.window = window - self.drop_nyquist = drop_nyquist - - if hop_length is None: - self.hop_length = n_fft // 4 - - if cmap is not None: - self.cmap = plt.get_cmap(cmap) - - bin_count = (n_fft // 2) + 1 - if drop_nyquist: - bin_count -= 1 - nyquist = sr // 2 - if fmin == None: - min_bin = 1 - else: - min_bin = (fmin / nyquist) * (bin_count - 1) - if fmax == None: - max_bin = bin_count - 1 - else: - max_bin = (fmax / nyquist) * (bin_count - 1) - bins = np.arange(bin_count) - if linear: - y_remap = (bins + min_bin) / ((bin_count - 1) / (max_bin - min_bin)) - else: - scale_factor = bin_count / np.log10(max_bin / min_bin) - y_remap = min_bin * 10 ** (bins / scale_factor) - if flipud: - y_remap = y_remap[::-1] - self.mapy_base = y_remap.reshape(-1,1).astype(np.float32) - - def __call__(self, audio): - """Generate a spectrogram for visualization. - - Returns: - np.array of type np.uint8 if `db_range` or `cmap` is set, otherwise np.float32 - """ - stft = librosa.stft(audio, - n_fft=self.n_fft, - hop_length=self.hop_length, - window=self.window) - if self.drop_nyquist: - stft = stft[:-1] - db = librosa.amplitude_to_db(np.abs(stft), ref=self.db_ref) - mapx = np.arange(db.shape[1], dtype=np.float32).reshape(1,-1).repeat(db.shape[0], axis=0) - mapy = self.mapy_base.repeat(db.shape[1], axis=1) - out = cv2.remap(db, mapx, mapy, cv2.INTER_LANCZOS4) - - if self.db_range is not None: - out -= self.db_range[0] - out /= self.db_range[1] - self.db_range[0] - - if self.cmap is not None: - if self.db_range is None: - out -= out.min() - out /= out.max() - out = self.cmap(out)[...,:3] - - if self.cmap is not None or self.db_range is not None: - out = np.clip(out * 256, 0, 255).astype(np.uint8) - - return out \ No newline at end of file diff --git a/core/warm.py b/core/warm.py deleted file mode 100644 index fff8cf1..0000000 --- a/core/warm.py +++ /dev/null @@ -1,51 +0,0 @@ -"""Renders one synthetic clip at image build time. - -The first render in a fresh container is slow: numba compiles the librosa -functions it meets, and matplotlib scans the system fonts to build its font -list. Running one render here leaves those caches on disk, and the Dockerfile -moves them into the image for app.py to restore on cold start. - -The clip is AAC in MPEG-TS, the format the hydrophones stream, so the render -goes through the same decoder as a real segment. make_spectrogram skips the S3 -download when the clip already exists at /tmp/, and skips the upload when -image_key is None, so this touches no AWS. -""" - -import subprocess - -import app - -CLIP = "warm.ts" - -subprocess.run( - [ - "ffmpeg", - "-hide_banner", - "-loglevel", - "error", - "-y", - "-f", - "lavfi", - "-i", - "anoisesrc=d=10:c=pink:r=48000", - "-ac", - "1", - "-c:a", - "aac", - "-f", - "mpegts", - f"/tmp/{CLIP}", - ], - check=True, -) - -app.make_spectrogram( - { - "id": CLIP, - "audio_bucket": "none", - "audio_key": "none", - "sample_rate": None, - "image_key": None, - "image_bucket": None, - } -) diff --git a/docs/comparison.png b/docs/comparison.png new file mode 100644 index 0000000..7d97129 Binary files /dev/null and b/docs/comparison.png differ diff --git a/events/spectrogram_job.json b/events/spectrogram_job.json index 4363cf7..9abc5fb 100644 --- a/events/spectrogram_job.json +++ b/events/spectrogram_job.json @@ -1,7 +1,11 @@ { - "id": "spectrogram_job_example", - "audio_bucket": "audio-orcasound-net", - "audio_key": "rpi_orcasound_lab/hls/1790665213/live5238.ts", - "image_bucket": null, - "image_key": null + "audio_url": "https://audio-orcasound-net.s3.us-west-2.amazonaws.com/rpi_orcasound_lab/hls/1790665213/live5238.ts", + "image_url": null, + "parameters": { + "fmin": 1, + "fmax": 15000, + "db_min": 10, + "db_max": 80, + "cmap": "viridis" + } } diff --git a/template.yaml b/template.yaml index 0c8b932..900ccdf 100644 --- a/template.yaml +++ b/template.yaml @@ -20,11 +20,12 @@ Resources: PackageType: Image Architectures: - x86_64 - # CPU scales with memory (a full vCPU at 1769 MB). A cold render at 512 MB - # peaked at 500 MB and did not finish inside the 60 s timeout. + # CPU scales with memory (a full vCPU at 1769 MB); ffmpeg renders a 10 s + # segment in about 2 s at this size. MemorySize: 1024 - # The function reads one segment and writes one image. The buckets are - # those orcasite's environments use: production and staging read from + # Only for the previous request shape, which names buckets instead of + # presigned URLs; remove once orcasite passes URLs. The buckets are those + # orcasite's environments use: production and staging read from # audio-orcasound-net and write to audio-deriv-orcasound-net; dev also # reads dev-streaming-orcasound-net and writes dev-audio-viz. Policies: diff --git a/tests/smoke.py b/tests/smoke.py index d6cc678..2c52283 100644 --- a/tests/smoke.py +++ b/tests/smoke.py @@ -1,48 +1,47 @@ -"""Renders one clip in a fresh container and reports how long each step took. +"""Renders one synthetic clip through the handler and checks the image. -Run inside the built image, the way CI does, to rehearse a Lambda cold start: -an unprivileged user, a root-owned empty /tmp, and the CPU and memory of the -deployed tier. The clip is AAC in MPEG-TS like a real segment. The render -skips S3 both ways, as in warm.py. - -Fails if the first render is slow enough to suggest the caches were not -restored, or if the two renders disagree about the image. +Run inside the built image, as CI does, as an unprivileged user with a +root-owned empty /tmp and the CPU and memory of the deployed tier, which is +how Lambda runs it. The clip is AAC in MPEG-TS like a real segment, read +through a file:// URL so nothing outside the container is touched. """ +import struct import subprocess import sys import time -CLIP = "smoke.ts" -JOB = { - "id": CLIP, - "audio_bucket": "none", - "audio_key": "none", - "sample_rate": None, - "image_key": None, - "image_bucket": None, -} -# A cold render without the caches took over 30 s at this CPU share. -SLOW_FIRST_RENDER = 15.0 +import app + +CLIP = "/tmp/smoke.ts" +SECONDS = 10 +SAMPLE_RATE = 48000 subprocess.run( ["ffmpeg", "-hide_banner", "-loglevel", "error", "-y", "-f", "lavfi", "-i", - "anoisesrc=d=10:c=pink:r=48000", "-ac", "1", "-c:a", "aac", "-f", "mpegts", f"/tmp/{CLIP}"], + f"anoisesrc=d={SECONDS}:c=pink:r={SAMPLE_RATE}", "-ac", "1", "-c:a", "aac", "-f", "mpegts", CLIP], check=True, ) -t0 = time.time() -import app # noqa: E402 (timed on purpose: this is where the caches are restored) - -t1 = time.time() -first = app.make_spectrogram(JOB) -t2 = time.time() -second = app.make_spectrogram(JOB) -t3 = time.time() - -print(f"import {t1 - t0:.1f}s | first render {t2 - t1:.1f}s | second render {t3 - t2:.1f}s | image {first['image_size']} bytes") +captured = {} +app.store_image = lambda job, png: captured.update(png=png) -if first["image_size"] != second["image_size"]: - sys.exit(f"renders differ: {first['image_size']} vs {second['image_size']} bytes") -if t2 - t1 > SLOW_FIRST_RENDER: - sys.exit(f"first render took {t2 - t1:.1f}s; were the caches restored?") +t0 = time.time() +result = app.lambda_handler({"audio_url": f"file://{CLIP}", "image_url": None}, None) +elapsed = time.time() - t0 + +png = captured["png"] +if png[:8] != b"\x89PNG\r\n\x1a\n": + sys.exit("output is not a PNG") +width, height = struct.unpack(">II", png[16:24]) +expected_width = 1 + int(result["sample_rate"] * SECONDS) // app.DEFAULTS["hop_length"] + +print(f"rendered {width}x{height} in {elapsed:.1f}s, {len(png)} bytes, sample rate {result['sample_rate']}") + +if height != app.DEFAULTS["n_fft"] // 2: + sys.exit(f"height {height}, expected {app.DEFAULTS['n_fft'] // 2}") +# AAC pads the clip by up to a couple of frames, so allow a few columns. +if abs(width - expected_width) > 8: + sys.exit(f"width {width}, expected about {expected_width}") +if result["image_size"] != len(png): + sys.exit(f"image_size {result['image_size']} does not match the {len(png)} bytes rendered")