Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
37 changes: 20 additions & 17 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,41 +1,44 @@
# spectrogram-renderer

A Lambda function that renders one spectrogram PNG from one hydrophone audio segment. [Orcasite](https://github.com/orcasound/orcasite) invokes it by name from [`Orcasite.Radio.AwsClient`](https://github.com/orcasound/orcasite/blob/main/server/lib/orcasite/radio/aws_client.ex), once per `AudioImage`, passing the S3 location of the segment and where to put the image. It has no other entry point.
A Lambda function that renders one spectrogram PNG from one hydrophone audio segment. [Orcasite](https://github.com/orcasound/orcasite) invokes it by name, once per `AudioImage`, and stores what it answers as the image's parameters. It has no other entry point.

- [`core/app.py`](core/app.py): the handler. Downloads the segment, decodes it with ffmpeg, renders with [`core/spectrogram_generator.py`](core/spectrogram_generator.py), uploads the PNG.
- [`core/Dockerfile`](core/Dockerfile): the image, built on the AWS Lambda Python base with a static ffmpeg.
- [`template.yaml`](template.yaml): the function's memory, timeout and IAM policy, as a SAM template. Stack name (`audio-viz`, kept from when this lived in orcasite) and region are in [`samconfig.toml`](samconfig.toml).
The whole render is one ffmpeg command ([`core/app.py`](core/app.py), `render`): `showspectrumpic` draws the spectrogram as intensity and `pseudocolor` applies matplotlib's colour ramp. The handler around it fetches the audio, probes its sample rate and duration to size the image, and stores the PNG. There is no Python beyond the standard library, so the image is the Lambda base plus a static ffmpeg, and a cold start costs about a second.

## Cold starts
## Contract

A new Lambda container starts with an empty `/tmp`, where numba, librosa and matplotlib keep their caches, so the first render in a container recompiles librosa's numba functions and rebuilds the font list. The Dockerfile runs [`warm.py`](core/warm.py) once at build time and keeps the caches it leaves in the image; [`app.py`](core/app.py) copies them into `/tmp` before importing those libraries. If you add a library that caches on first use, give it a directory under `/tmp` in the Dockerfile and add it to the same `mv`.
```json
{
"audio_url": "https://… (presigned GET)",
"image_url": "https://… (presigned PUT), or null to render and discard",
"parameters": {"n_fft": 1024, "hop_length": 256, "fmin": 1, "fmax": 15000, "db_min": 10, "db_max": 80, "cmap": "viridis", "frequency_scale": "linear"}
}
```

The function reads and writes only the URLs it is given and knows nothing about what is behind them; the caller decides where audio and images live and presigns accordingly. Every parameter is optional and defaults to what orcasite has always used, so images stay consistent with the existing archive. The response gives `image_size`, `sample_rate`, `width`, `height` and the `parameters` applied, plus the field names orcasite has stored since the first renderer.

The image is `n_fft/2` pixels high and one pixel wide per `hop_length` samples, as the previous renderer's STFT was. `db_min`/`db_max` are the dB window relative to an amplitude of 0.01, the convention the archive was drawn with; `DB_OFFSET` in `app.py` converts it to ffmpeg's full-scale dB and was fitted against a production image, whose luminance distribution the new render reproduces to within 0.01.

To check a deployed function, find `platform.report` records in its CloudWatch log (the template sets `LogFormat: JSON`). Cold invocations carry `initDurationMs`; their `durationMs + initDurationMs` should be within a couple of seconds of a warm invocation's `durationMs`. Lambda's `Duration` metric excludes init time, so it alone understates a cold start.
Until orcasite has switched, the previous shape (`audio_bucket`/`audio_key`, `image_bucket`/`image_key`, read and written with boto3) is also accepted; [`template.yaml`](template.yaml) keeps the S3 policy for it and both go once no caller uses them.

## Testing

[`tests/smoke.py`](tests/smoke.py) renders a synthetic MPEG-TS clip inside the built image, run as an unprivileged user with a root-owned empty `/tmp` and the CPU and memory of the deployed tier, which is how Lambda runs it. It fails if the first render is slow enough to suggest the caches were not restored. [CI](.github/workflows/ci.yml) runs it on every pull request; locally:
[`tests/smoke.py`](tests/smoke.py) renders a synthetic AAC/MPEG-TS clip through the handler inside the built image, run as an unprivileged user with a root-owned empty `/tmp` and the CPU and memory of the deployed tier, which is how Lambda runs it. [CI](.github/workflows/ci.yml) runs it on every pull request; locally:

```bash
docker build -t spectrogram-renderer core
docker run --rm --user 1000:1000 --tmpfs /tmp:uid=0,gid=0,mode=1777 --cpus 0.58 --memory 1024m \
--entrypoint python -v "$PWD/tests/smoke.py:/var/task/smoke.py" spectrogram-renderer smoke.py
```

That rehearses the application, not the sandbox: Lambda's Runtime Interface Emulator in the base image reproduces the invoke API, not the filesystem rules. Anything that touches `/tmp`'s own metadata or relies on image contents under `/tmp` will pass locally on a plain tmpfs and fail deployed, hence the `--user` and `--tmpfs` flags above.
That rehearses the application, not the sandbox: Lambda's Runtime Interface Emulator in the base image reproduces the invoke API, not the filesystem rules, hence the `--user` and `--tmpfs` flags.

## Deploying

Merging to `main` deploys: the [CI workflow](.github/workflows/ci.yml) runs the smoke test, then `sam build && sam deploy` under the `production` environment, assuming the `spectrogram-renderer-deploy` IAM role through GitHub's OIDC provider. That role can change only this stack, its ECR repository and its function role.
Merging to `main` deploys: the [CI workflow](.github/workflows/ci.yml) runs the smoke test, then `sam build && sam deploy` under the `production` environment, assuming the `spectrogram-renderer-deploy` IAM role through GitHub's OIDC provider. That role can change only this stack (`audio-viz`, the name kept from when this lived in orcasite), its ECR repository and its function role.

To deploy from a machine instead, with an AWS profile for the Orcasound account and the [SAM CLI](https://docs.aws.amazon.com/serverless-application-model/latest/developerguide/serverless-sam-cli-install.html):

```bash
sam build
AWS_PROFILE=orcasound sam deploy
```
To deploy from a machine instead, with an AWS profile for the Orcasound account and the [SAM CLI](https://docs.aws.amazon.com/serverless-application-model/latest/developerguide/serverless-sam-cli-install.html): `sam build && AWS_PROFILE=orcasound sam deploy`.

To invoke the deployed function on a real segment without writing anything (`image_key` null skips the upload):
To invoke the deployed function on a real segment without writing anything ([`events/spectrogram_job.json`](events/spectrogram_job.json) reads from the public audio bucket and has `image_url` null):

```bash
aws lambda invoke --function-name "$(aws lambda list-functions --query "Functions[?contains(FunctionName,'AudioVizFunction')].FunctionName | [0]" --output text)" \
Expand Down
39 changes: 10 additions & 29 deletions core/Dockerfile
Original file line number Diff line number Diff line change
@@ -1,35 +1,16 @@
FROM public.ecr.aws/lambda/python:3.14

# ffmpeg does all the work; the handler is standard library plus the boto3
# the base image ships, so there is nothing to pip install and no cache to
# warm. A static build keeps the image at the base plus one binary.
ARG FFMPEG_VERSION=ffmpeg-7.0.2-amd64-static
RUN dnf install -y tar xz \
&& curl -sO https://johnvansickle.com/ffmpeg/releases/${FFMPEG_VERSION}.tar.xz \
&& tar -xf ${FFMPEG_VERSION}.tar.xz \
&& mv ${FFMPEG_VERSION}/ffmpeg ${FFMPEG_VERSION}/ffprobe /usr/local/bin/ \
&& rm -rf ${FFMPEG_VERSION} ${FFMPEG_VERSION}.tar.xz \
&& dnf clean all

# /tmp is the only writeable directory at runtime, but Lambda mounts a fresh,
# empty one for every new container, so anything put there at build time is
# gone by the first invocation. The caches these libraries build on first use
# are therefore built into the image (see warm.py below) and copied into /tmp
# by app.py when a container starts.
ENV CACHE_SEED=/opt/cache-seed
ENV MPLCONFIGDIR=/tmp/matplotlib
ENV LIBROSA_CACHE_DIR=/tmp/librosa_cache
ENV NUMBA_CACHE_DIR=/tmp/numba_cache
COPY app.py ./

RUN dnf install -y tar gzip xz

RUN curl -O https://johnvansickle.com/ffmpeg/releases/${FFMPEG_VERSION}.tar.xz
RUN tar -xvf "${FFMPEG_VERSION}.tar.xz"
RUN mv ${FFMPEG_VERSION}/ffmpeg /usr/local/bin
RUN mv ${FFMPEG_VERSION}/ffprobe /usr/local/bin

COPY app.py requirements.txt spectrogram_generator.py warm.py ./
RUN mkdir -p data

RUN python3 -m pip install -r requirements.txt -t .

# Render one clip so that every cache a render populates exists before the
# first real invocation, then keep the result where app.py can find it.
RUN mkdir -p ${MPLCONFIGDIR} ${LIBROSA_CACHE_DIR} ${NUMBA_CACHE_DIR} \
&& python3 warm.py \
&& mkdir -p ${CACHE_SEED} \
&& mv ${MPLCONFIGDIR} ${LIBROSA_CACHE_DIR} ${NUMBA_CACHE_DIR} ${CACHE_SEED}/

# Command can be overwritten by providing a different command in the template directly.
CMD ["app.lambda_handler"]
Empty file removed core/__init__.py
Empty file.
252 changes: 119 additions & 133 deletions core/app.py
Original file line number Diff line number Diff line change
@@ -1,159 +1,145 @@
import os
import shutil

# Every new Lambda container starts with an empty /tmp, where numba, librosa
# and matplotlib keep their caches. Without them the first render recompiles
# librosa's numba functions and rebuilds the font list, which at 512 MB took
# longer than the function's timeout. The Dockerfile builds the caches into
# the image; restore them before those libraries load and look for them.
#
# Copy each cache directory separately: copying onto /tmp itself makes
# copytree set /tmp's permissions, which Lambda refuses. A failure here only
# costs a slow first render, so it must never fail the invocation.
_cache_seed = os.environ.get("CACHE_SEED")
if _cache_seed and os.path.isdir(_cache_seed):
for _name in os.listdir(_cache_seed):
try:
shutil.copytree(
os.path.join(_cache_seed, _name),
os.path.join("/tmp", _name),
dirs_exist_ok=True,
)
except OSError as e:
print(f"Could not restore cache {_name}: {e}")

from matplotlib.pyplot import imshow
from spectrogram_generator import SpectrogramGenerator
from subprocess import check_output
from typing import TypedDict, Optional
import boto3
import io
import json
import librosa
import matplotlib
import sys

SpectrogramJob = TypedDict(
"SpectrogramJob",
{
"id": str,
"audio_bucket": str,
"audio_key": str,
"sample_rate": int,
"image_key": Optional[str],
"image_bucket": Optional[str],
},
)


def lambda_handler(event, context):
"""Sample pure Lambda function

Parameters
----------
event: dict, required
API Gateway Lambda Proxy Input Format

Event doc: https://docs.aws.amazon.com/apigateway/latest/developerguide/set-up-lambda-proxy-integrations.html#api-gateway-simple-proxy-for-lambda-input-format

context: object, required
Lambda Context runtime methods and attributes
"""Renders one spectrogram PNG from one audio segment, with ffmpeg.

Context doc: https://docs.aws.amazon.com/lambda/latest/dg/python-context-object.html
The request says where to read the audio and where to write the image, as
URLs; the function knows nothing about what is behind them. Render
parameters are explicit, with the defaults orcasite has always used, and
the response echoes the ones that were applied.

Returns
------
API Gateway Lambda Proxy Output Format: dict
{
"audio_url": "https://... (presigned GET)",
"image_url": "https://... (presigned PUT), or null to render and discard",
"parameters": {"fmax": 15000, ...} # optional, see DEFAULTS
}

Return doc: https://docs.aws.amazon.com/apigateway/latest/developerguide/set-up-lambda-proxy-integrations.html
"""
result = make_spectrogram(event)
For one release the previous shape is also accepted: `audio_bucket` and
`audio_key`, `image_bucket` and `image_key`, read and written with boto3.
"""

return {"status": 200, **result}
import os
import subprocess
import tempfile
import urllib.request
import wave

DEFAULTS = {
"n_fft": 1024,
"hop_length": 256,
"fmin": 1,
"fmax": 15000,
"db_min": 10,
"db_max": 80,
"cmap": "viridis",
"frequency_scale": "linear",
}

# ffmpeg's dB scale is relative to full scale, while the dB window orcasite's
# images were drawn with is relative to an amplitude of 0.01 in librosa's
# STFT. This offset lines the two up: it was fitted against a production
# image, whose luminance distribution it reproduces to within 0.01.
DB_OFFSET = -122

FSCALE = {"linear": "lin", "log": "log"}


def make_spectrogram(
job: SpectrogramJob, store_audio=False, show_image=False, store_image=False
):
print(f"Received job: {json.dumps(job)}")
local_path = f"/tmp/{job["id"]}"
s3 = boto3.client("s3")
def lambda_handler(event, context):
return {"status": 200, **make_spectrogram(event)}

if not os.path.exists(local_path):
response = s3.get_object(Bucket=job["audio_bucket"], Key=job["audio_key"])
audio_bytes = response["Body"].read()
with open(local_path, "wb") as file:
file.write(audio_bytes)

sample_rate = job.get("sample_rate")
if sample_rate is None:
metadata = get_audio_metadata(local_path)
sample_rate = int(metadata["streams"][0]["sample_rate"])
def make_spectrogram(job):
params = {**DEFAULTS, **(job.get("parameters") or {})}

audio, sr = load_audio(local_path, sample_rate)
with tempfile.TemporaryDirectory(dir="/tmp") as tmp:
audio_path = os.path.join(tmp, "audio")
wav_path = os.path.join(tmp, "audio.wav")
fetch_audio(job, audio_path)
sample_rate, samples = decode(audio_path, wav_path)
width = 1 + samples // params["hop_length"]
height = params["n_fft"] // 2
png = render(wav_path, params, width, height)

params = {"linear": True, "fmin": 1, "fmax": 15000, "cmap": "viridis"}

generator = SpectrogramGenerator(
sr, n_fft=1024, hop_length=256, db_range=(10, 80), **params
)
spectrogram = generator(audio)
if show_image:
matplotlib.pyplot.axis("off")
imshow(spectrogram)

# Store spectrogram image either as a file or in-memory
image_size = None
image_store = f"/tmp/{job["id"]}.png" if store_image else io.BytesIO()
matplotlib.image.imsave(image_store, spectrogram)
if store_image:
image_size = os.path.getsize(image_store)
else:
image_store.seek(0)
image_size = sys.getsizeof(image_store)
image_store.seek(0)

if job["image_key"] is not None and job["image_bucket"] is not None:
image_content = open(image_store, "rb") if store_image else image_store
s3.put_object(
Bucket=job["image_bucket"], Key=job["image_key"], Body=image_content
)
store_image(job, png)

return {
"image_size": len(png),
"sample_rate": sample_rate,
"image_size": image_size,
"frequency_spacing": "linear" if params["linear"] else "log",
"width": width,
"height": height,
"parameters": params,
# The names orcasite has stored since the first renderer.
"frequency_spacing": params["frequency_scale"],
"freq_min": params["fmin"],
"freq_max": params["fmax"],
"color_map": params["cmap"],
}


def load_audio(local_path, sample_rate):
"""Decodes the segment with ffmpeg and hands librosa a WAV.
def render(audio_path, params, width, height):
"""One ffmpeg invocation: spectrogram as intensity, then the colour map.

The hydrophones stream AAC in MPEG-TS, which libsndfile does not read, and
librosa 1.0 dropped the ffmpeg fallback that used to cover it.
showspectrumpic's own colour maps start at black, so the image is drawn
in grey and matplotlib's ramp applied with pseudocolor, which carries the
same tables. Its window size follows from the height (n_fft/2 bins).
"""
wav = check_output(
["ffmpeg", "-hide_banner", "-loglevel", "error", "-i", local_path, "-f", "wav", "-"]
spectrum = ":".join(
[
f"s={width}x{height}",
"mode=combined",
"color=intensity",
"scale=log",
f"fscale={FSCALE[params['frequency_scale']]}",
"win_func=hann",
f"start={params['fmin']}",
f"stop={params['fmax']}",
f"limit={params['db_max'] + DB_OFFSET}",
f"drange={params['db_max'] - params['db_min']}",
"legend=0",
]
)
return librosa.load(io.BytesIO(wav), sr=sample_rate)
graph = f"showspectrumpic={spectrum},format=gbrp,pseudocolor=preset={params['cmap']}"
return subprocess.run(
["ffmpeg", "-hide_banner", "-loglevel", "error", "-i", audio_path,
"-lavfi", graph, "-f", "image2pipe", "-c:v", "png", "-"],
check=True,
capture_output=True,
).stdout


def get_audio_metadata(local_path):
result = check_output(
[
"ffprobe",
"-hide_banner",
"-loglevel",
"panic",
"-show_format",
"-show_streams",
"-of",
"json",
local_path,
]
def decode(audio_path, wav_path):
"""Decodes the segment to WAV and returns its sample rate and length.

The image is sized from the samples actually decoded: the container's
duration undercounts an AAC stream by a frame or two, which would make
each tile a few pixels narrower than the archive's.
"""
subprocess.run(
["ffmpeg", "-hide_banner", "-loglevel", "error", "-i", audio_path, "-f", "wav", wav_path],
check=True,
capture_output=True,
)
with wave.open(wav_path, "rb") as wav:
return wav.getframerate(), wav.getnframes()


return json.loads(result)
def fetch_audio(job, path):
if job.get("audio_url"):
with urllib.request.urlopen(job["audio_url"]) as response, open(path, "wb") as file:
file.write(response.read())
else:
import boto3

boto3.client("s3").download_file(job["audio_bucket"], job["audio_key"].lstrip("/"), path)


def store_image(job, png):
if job.get("image_url"):
request = urllib.request.Request(
job["image_url"], data=png, method="PUT", headers={"Content-Type": "image/png"}
)
with urllib.request.urlopen(request):
pass
elif job.get("image_bucket") and job.get("image_key"):
import boto3

boto3.client("s3").put_object(
Bucket=job["image_bucket"], Key=job["image_key"].lstrip("/"), Body=png, ContentType="image/png"
)
5 changes: 0 additions & 5 deletions core/requirements.txt

This file was deleted.

Loading
Loading