Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
33 commits
Select commit Hold shift + click to select a range
83aefda
Feat/vendored opencode binary (#299)
guillaume-byte Aug 26, 2026
454d039
Fix/v2.1 UI fixes (#300)
guillaume-byte Aug 26, 2026
40ad072
fix CI core dumped
guillaume-byte Aug 26, 2026
0ed0500
fix stucked CI
guillaume-byte Aug 26, 2026
9e9820a
Deselect scale tests from the fast CI job; cap test job at 15 minutes
guillaume-byte Aug 26, 2026
4728db2
Update CHANGELOG - v2.0.1.dev0
guillaume-byte Aug 26, 2026
87f19ae
Merge branch 'main' into dev
guillaume-byte Aug 26, 2026
328e46e
refactor(examples): one import surface — `import weightslab as wl`
guillaume-byte Sep 4, 2026
3c8c1b4
Resolve MConflicts And merge main to dev
guillaume-byte Sep 6, 2026
c74dc0e
Merge main into dev: docs sidebar tree, UL example rework, label fixes
guillaume-byte Sep 8, 2026
c25bd45
fix(notebook): keep the trainer's stdout out of notebook cell outputs…
guillaume-byte Sep 10, 2026
b0347d3
Interactivity on 100GB+ datasets: O(change) view sync and ledger writ…
AlexGrayBox Sep 10, 2026
0ab50a9
test: fix the four failures in the unit suite
guillaume-byte Sep 10, 2026
5cd4920
Fix Ultralytics issue with v8.*
guillaume-byte Sep 11, 2026
ac40533
Merge branch 'main' into dev
guillaume-byte Sep 22, 2026
a2b8902
test: fix the gRPC serve flake that failed CI with identical code
guillaume-byte Sep 22, 2026
caa1d74
fix(notebook): shut the embedded kernel down instead of aborting at exit
guillaume-byte Sep 23, 2026
e9f7f23
fix utest error
guillaume-byte Sep 25, 2026
9927abc
update cls example with basic CNN 1M parameters
guillaume-byte Sep 25, 2026
b2c23fa
fix chkpt restoring from several root dirs
guillaume-byte Sep 25, 2026
0b13524
fix cls example
guillaume-byte Sep 25, 2026
21fcb4d
se/start: OS-aware cert generation, certs dir fallback, CLI help + do…
guillaume-byte Sep 29, 2026
9078eb3
logs: terminal/file split, survive dictConfig, tqdm mirror; rewind sa…
guillaume-byte Oct 1, 2026
91495d4
Fix certs generation for encrypted communication
guillaume-byte Oct 1, 2026
7ee87b8
Remove useless logs, redundant, and fix dataloader indexing related t…
guillaume-byte Oct 1, 2026
be170bf
Update documentation
guillaume-byte Oct 1, 2026
bef1702
fix related utests with recent bugs fixed and changes
guillaume-byte Oct 1, 2026
1fd1e9c
fix dataloader pinmemory failing test
guillaume-byte Oct 1, 2026
804d235
fix UI plots points budgets
guillaume-byte Oct 1, 2026
476b927
fix test
guillaume-byte Oct 1, 2026
2097c02
Fix pytest timeout, too short
guillaume-byte Oct 1, 2026
df74096
[skip ci] Fix Claude artefacts from doc and add quickstart migration
guillaume-byte Oct 1, 2026
cffe256
[skip ci] Comment auto start training for the examples
guillaume-byte Oct 1, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -242,7 +242,7 @@ jobs:
# test hitting the per-test timeout instead of being deselected burns
# 10 minutes AND leaves the job stuck at the same progress % across
# unrelated pushes, which read as a hang.
python -m pytest ./tests -v --timeout=600 -m "not scale"
python -m pytest ./tests -v --timeout=1200 -m "not scale"

# ── Agent smoke test on a pip-installed package ───────────────────────────
# Proves the Option-2 promise end-to-end: install weightslab into a CLEAN
Expand Down
4 changes: 2 additions & 2 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -19,8 +19,8 @@ venv
runs
data
outputs
!./tests/data/
!./weightslab/data/
!tests/data/
!weightslab/data/
MagicMock
drop
htmlcov
Expand Down
40 changes: 27 additions & 13 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -100,8 +100,11 @@ Working starting points live in
(each is a `main.py` + `config.yaml`) — find the closest example and mirror it.

UI deployment details (port, TLS, certs) are documented in
`weightslab/docs/weights_studio.rst`. TLS is opt-in: run `weightslab se` once,
then `weightslab start --certs`.
`weightslab/docs/weights_studio.rst`. TLS turns on once certs exist: run
`weightslab se` once, and `weightslab start` and the backend then use them
automatically (`weightslab start --no-certs` forces HTTP). On Windows `se` uses the PowerShell script and
the Windows `openssl`; `weightslab se --force-ubuntu` uses the bash script
through WSL instead.

---

Expand All @@ -120,10 +123,10 @@ the global ledger (`weightslab/weightslab/backend/ledgers.py`,

Conventions that matter for correctness:

- Wrap the train step in `with guard_training_context:` and eval in
`with guard_testing_context:` (from
`weightslab.components.global_monitoring`). This is how pause/resume and
train/test separation work — **skip it and pause/resume or stats will misbehave.**
- Wrap the train step in `with wl.guard_training_context:` and eval in
`with wl.guard_testing_context:` (both re-exported at package level; no deep
import). This is how pause/resume and train/test separation work — **skip it
and pause/resume or stats will misbehave.**
- Use `model.get_age()` (steps actually trained; survives checkpoint reloads),
not the raw loop counter.
- `task_type` on the dataset/model selects rendering: `classification`,
Expand Down Expand Up @@ -175,11 +178,13 @@ ones when debugging:

| Variable | Default | Why you touch it |
|---|---|---|
| `WEIGHTSLAB_LOG_LEVEL` | `INFO` | Set `DEBUG` to see what's happening. (`WATCHDOG` level sits between WARNING/ERROR.) |
| `WEIGHTSLAB_LOG_LEVEL` | `INFO` | **Terminal only.** Set `DEBUG` to see what's happening. (`WATCHDOG` level sits between WARNING/ERROR.) |
| `WEIGHTSLAB_LOG_FILE_LEVEL` | *(unset = all)* | The session log file keeps every record whatever the terminal shows; set this to cap the file too. Log lives in `<root_log_dir>/weightslab_logs/`. |
| `WEIGHTSLAB_TQDM_LOG_INTERVAL` | `30` | Seconds between snapshots of live `tqdm` bars into the log (`0` disables). tqdm never goes through `logging`, so without this the file has no record of training progress. |
| `GRPC_BACKEND_HOST` / `GRPC_BACKEND_PORT` | `0.0.0.0` / `50051` | Backend gRPC bind address. |
| `GRPC_TLS_ENABLED` | `0` | TLS on the gRPC socket. Set `1` with `weightslab start --certs`. |
| `GRPC_TLS_REQUIRE_CLIENT_AUTH` | `0` | mTLS. Must match what `weightslab start --certs` presents. |
| `WEIGHTSLAB_CERTS_DIR` | `~/.weightslab-certs` | Where cert files are looked up (single source of truth). |
| `GRPC_TLS_ENABLED` | `0` | TLS on the gRPC socket. Set to `1` automatically when certs are found; `0`/`false` forces plaintext (also for `weightslab start`). |
| `GRPC_TLS_REQUIRE_CLIENT_AUTH` | `0` | mTLS. Must match what `weightslab start` presents when it uses certs. |
| `WEIGHTSLAB_CERTS_DIR` | `~/.weightslab-certs` | Where cert files are looked up (single source of truth). Falls back to `~/.weightslab-certs` when unset, not an absolute path, or holding no certs. |
| `GRPC_AUTH_TOKEN` | *(unset)* | Optional metadata-token auth on top of mTLS. |
| `GRPC_MAX_MESSAGE_BYTES` | `268435456` (256 MB) | Raise it if large tensors/image batches fail. |
| `WEIGHTSLAB_DISABLE_WATCHDOGS` | `0` | Set `1` when debugging with breakpoints (see §5). |
Expand Down Expand Up @@ -216,9 +221,18 @@ distilled from issues hit in development).
**UI loads but the sample grid is empty / "failed to fetch" / gRPC errors.**
The wire path (§1) is broken somewhere. Check in order: (1) backend actually
serving on `0.0.0.0:50051`; (2) `weightslab start` is running and the browser
can reach it on `:8080`; (3) **TLS mismatch** if using `--certs` — run
`weightslab se` first and export `WEIGHTSLAB_CERTS_DIR`. For local debugging
drop TLS entirely (omit `--certs`; `GRPC_TLS_ENABLED=0`).
can reach it on `:8080`; (3) **TLS mismatch** — the UI and the backend each
turn TLS on when they find certs, so both must see the same
`WEIGHTSLAB_CERTS_DIR` (the browser console prints `TLS: ENABLED/DISABLED`).
For local debugging drop TLS on both sides (`weightslab start --no-certs`;
`GRPC_TLS_ENABLED=0` for the backend).

**`weightslab se` hangs with no output (Windows).**
Only the WSL path can do this: `--force-ubuntu`, or the fallback after the
PowerShell script fails. There `bash` is the WSL launcher, the script's output
is captured, and there is no timeout, so a stuck WSL distro blocks forever.
Confirm with `wsl -e echo ok` (it hangs too). Fix with `wsl --shutdown`, or drop
`--force-ubuntu` so the PowerShell script runs.

**Changed an env var, restarted, but the UI still uses the old value.**
- `VITE_*` is build-time → you must **rebuild** the frontend, not just restart.
Expand Down
2 changes: 1 addition & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1 +1 @@
# Changelog - 2026-08-26 v2.0.1 (2)
# Changelog - 2026-08-26 v2.0.1
6 changes: 2 additions & 4 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -244,8 +244,6 @@ the exact samples causing them — so you can fix your data, not just log it.

-import wandb
+import weightslab as wl
+from weightslab.components.global_monitoring import (
+ guard_training_context, guard_testing_context)
+
+@wl.signal(name="byte_adjusted_loss", subscribe_to="loss/CE")
+def byte_adjusted_loss(ctx): return ctx.subscribed_value / ctx.image_bytes
Expand Down Expand Up @@ -285,7 +283,7 @@ the exact samples causing them — so you can fix your data, not just log it.
for epoch in range(1, args.epochs + 1):
model.train()
for x, y in train_loader:
+ with guard_training_context:
+ with wl.guard_training_context:
logits = model(x.to(device))
loss = criterion(logits, y.to(device))
optimizer.zero_grad(); loss.backward(); optimizer.step()
Expand All @@ -298,7 +296,7 @@ the exact samples causing them — so you can fix your data, not just log it.
model.eval()
with torch.no_grad():
for x, y in test_loader:
+ with guard_testing_context:
+ with wl.guard_testing_context:
accuracy.update(model(x.to(device)), y)
- wandb.log({"test/acc": accuracy.compute().item(), "epoch": epoch})
+ wl.save_signals(preds_raw=logits, targets=y,
Expand Down
17 changes: 13 additions & 4 deletions agent_config.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,16 @@ agent:
opencode_url: http://127.0.0.1:4096

# OpenCode model, "providerID/modelID" (can also be set as env variable
# OPENCODE_MODEL). Empty string self-heals to whatever OpenCode's own
# config was last set to, or a configured provider default, falling back
# to the free-tier "opencode/deepseek-v4-flash-free" if neither resolves.
opencode_model: "opencode/deepseek-v4-flash-free"
# OPENCODE_MODEL).
#
# Left EMPTY on purpose. Empty means "follow OpenCode's own config" -- which
# is what the studio's model picker (.wl-ag-model) writes on every pick, via
# PUT /config -- so choosing a model in the UI also changes the model
# weightslab's own queries use, and it is re-checked before every turn.
# When OpenCode names no model either, the fallback is "opencode/big-pickle".
#
# A value here SEEDS the choice: it is used when nothing has been chosen in
# OpenCode's config yet (and is published there, so the studio shows it), but
# a model picked in the UI afterwards wins. Set OPENCODE_MODEL instead to
# PIN a model that nothing can override.
opencode_model: ""
2 changes: 1 addition & 1 deletion docs/_scripts/update_whats_new.py
Original file line number Diff line number Diff line change
Expand Up @@ -275,7 +275,7 @@ def build_rst(releases: list) -> str:
def main() -> None:
releases = stable_releases(fetch_releases())
OUTPUT.write_text(build_rst(releases), encoding="utf-8")
print(f"wrote {OUTPUT} — {len(releases)} releases, "
print(f"wrote {OUTPUT}, {len(releases)} releases, "
f"{releases[-1]['tag_name']} … {releases[0]['tag_name']}")


Expand Down
6 changes: 3 additions & 3 deletions docs/_static/custom.css
Original file line number Diff line number Diff line change
Expand Up @@ -160,7 +160,7 @@ body[data-theme="dark"] .wl-only-dark {
color: var(--color-foreground-primary);
}

/* Hide Furo's native theme-toggle buttons everywhere — only the topnav toggle is used */
/* Hide Furo's native theme-toggle buttons everywhere, only the topnav toggle is used */
.theme-toggle {
display: none !important;
}
Expand Down Expand Up @@ -438,14 +438,14 @@ body[data-theme="dark"] .wl-only-dark {
color: var(--color-foreground-primary);
}

/* Active state — "All" */
/* Active state, "All" */
.wl-eg-filter-btn.wl-eg-filter--active {
background: var(--color-foreground-primary);
color: var(--color-background-primary);
border-color: var(--color-foreground-primary);
}

/* Active state — per framework color */
/* Active state, per framework color */
.wl-eg-filter-btn[data-color="pytorch"].wl-eg-filter--active { background:#de4e20; border-color:#de4e20; color:#fff; }
.wl-eg-filter-btn[data-color="lightning"].wl-eg-filter--active { background:#7b2bdb; border-color:#7b2bdb; color:#fff; }
.wl-eg-filter-btn[data-color="ultralytics"].wl-eg-filter--active { background:#0a9e4e; border-color:#0a9e4e; color:#fff; }
Expand Down
18 changes: 9 additions & 9 deletions docs/_static/examples-gallery.js
Original file line number Diff line number Diff line change
Expand Up @@ -16,31 +16,31 @@
var EXAMPLES = [
{
badge: 'PyTorch', color: 'pytorch',
title: 'Classification — MNIST',
title: 'Classification, MNIST',
desc: 'CNN digit classifier on MNIST. Register hyperparameters, monitor per-sample loss, and use the deny-aware sampler to focus on hard examples.',
tags: ['classification', 'supervised', 'mnist', 'cnn'],
url: 'examples/pytorch/classification.html',
colab: COLAB + 'PyTorch/wl-classification.ipynb'
},
{
badge: 'PyTorch', color: 'pytorch',
title: 'Segmentation — BDD100k',
title: 'Segmentation, BDD100k',
desc: 'Per-pixel semantic segmentation with a UNet. Track per-sample IoU and visualise mask overlays directly in the studio.',
tags: ['segmentation', 'semantic', 'bdd100k', 'masks', 'dense prediction'],
url: 'examples/pytorch/segmentation.html',
colab: COLAB + 'PyTorch/wl-segmentation.ipynb'
},
{
badge: 'PyTorch', color: 'pytorch',
title: 'Detection — Penn-Fudan',
title: 'Detection, Penn-Fudan',
desc: 'Bounding-box detection on Penn-Fudan pedestrians. Per-instance multi-index dataframe with (sample_id, annotation_id) keys.',
tags: ['detection', 'object detection', 'bounding boxes', 'penn-fudan'],
url: 'examples/pytorch/detection.html',
colab: COLAB + 'PyTorch/wl-detection.ipynb'
},
{
badge: 'PyTorch', color: 'pytorch',
title: 'Clustering — Face Recognition',
title: 'Clustering, Face Recognition',
desc: 'Metric learning with triplet loss on face datasets. Store and explore high-dimensional embeddings per sample in the studio.',
tags: ['clustering', 'unsupervised', 'embeddings', 'face recognition', 'metric learning'],
url: 'examples/pytorch/clustering.html',
Expand All @@ -56,22 +56,22 @@
},
{
badge: 'Lightning', color: 'lightning',
title: 'Classification — MNIST (Lightning)',
desc: 'Same MNIST classification wrapped in a LightningModule. WeightsLab hooks replace only the guard functions — the rest is unchanged.',
title: 'Classification, MNIST (Lightning)',
desc: 'Same MNIST classification wrapped in a LightningModule. WeightsLab hooks replace only the guard functions, the rest is unchanged.',
tags: ['classification', 'supervised', 'mnist', 'pytorch lightning'],
url: 'examples/lightning/classification.html'
},
{
badge: 'Ultralytics', color: 'ultralytics',
title: 'Detection — YOLO',
title: 'Detection, YOLO',
desc: 'Drop-in WLAwareTrainer for YOLO training. Track mAP, per-image loss, and discard low-quality samples without touching the model.',
tags: ['detection', 'yolo', 'object detection', 'mAP'],
url: 'examples/ultralytics/detection.html',
colab: COLAB + 'Ultralytics/wl-how-to-train-ultralytics-yolo-on-kitti-detection-dataset.ipynb'
},
{
badge: 'Usecase', color: 'usecase',
title: 'LiDAR Detection — 2D and 3D',
title: 'LiDAR Detection, 2D and 3D',
desc: 'Point-cloud BEV previews, dual 2D/3D bounding box signals, streaming GetPointCloud RPC, and an interactive three.js 3D viewer.',
tags: ['lidar', 'point cloud', '3d detection', 'bev', 'streaming'],
url: 'examples/usecases/lidar_detection.html'
Expand All @@ -86,7 +86,7 @@
},
{
badge: 'Usecase', color: 'usecase',
title: 'Model Signals — Fashion-MNIST',
title: 'Model Signals, Fashion-MNIST',
desc: 'Per-step training dynamics: global and per-layer gradient norms, weight norms and activation statistics, from one argument on the model wrap.',
tags: ['model signals', 'gradient norm', 'activations', 'per-layer', 'training dynamics'],
url: 'examples/usecases/model_signals.html'
Expand Down
2 changes: 1 addition & 1 deletion docs/_static/screenshots/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,7 @@

One image per feature section in `docs/weights_studio.rst`.

Every file here is currently a **placeholder** — a grey frame naming the
Every file here is currently a **placeholder**, a grey frame naming the
feature it stands in for. To add a real screenshot, overwrite the file
**keeping its exact filename**; the docs reference these paths directly, so
nothing else needs editing.
Expand Down
20 changes: 10 additions & 10 deletions docs/_static/wl-ribbon.js
Original file line number Diff line number Diff line change
Expand Up @@ -2,27 +2,27 @@
'use strict';

var TIPS = [
'Edit <code>hyperparameters.yaml</code> while training — changes apply within 1 second, no restart needed.',
'Edit <code>hyperparameters.yaml</code> while training, changes apply within 1 second, no restart needed.',
'Click any sample in the studio to deny it from future batches. The deny-aware sampler persists tags across runs.',
'Call <code>wl.keep_serving()</code> after your training loop to keep the studio live for post-training analysis.',
'Add <code>per_sample=True</code> to a <code>@wl.signal</code> decorator to store one value per sample per step.',
'Set <code>is_training=True</code> on your DataLoader kwargs to activate the deny-aware sampler.',
'The studio streams signals in real-time — no need to wait for an epoch to end to see results.',
'The studio streams signals in real-time, no need to wait for an epoch to end to see results.',
'<code>weightslab start example --cls</code> launches a full MNIST classification demo in one command.',
'Use <code>subscribe_to=</code> on a signal to build reactive per-sample analytics derived from other signals.',
'Run <code>weightslab start --certs</code> to enable HTTPS + mTLS for secure remote studio access.',
'Run <code>weightslab se</code> once: <code>weightslab start</code> then serves HTTPS + mTLS automatically.',
'Set <code>preload_labels=False</code> for large datasets to speed up startup; labels are loaded lazily.',
'Use <code>array_return_proxies=True</code> (default) to avoid loading the full dataset array into RAM.',
'Set <code>WEIGHTSLAB_LOG_LEVEL=DEBUG</code> to see full gRPC logs when debugging connectivity issues.',
'Call <code>wl.ai_report_generation()</code> — or run <code>report</code> in <code>weightslab cli</code> — for a full HTML report with plots, dataset analysis, and agent-written insights.',
'Type <code>/init</code> in the experiment agent bar to bring the integrated OpenCode agent online for a running experiment — no separate server to start.',
'Call <code>wl.ai_report_generation()</code>, or run <code>report</code> in <code>weightslab cli</code>, for a full HTML report with plots, dataset analysis, and agent-written insights.',
'Type <code>/init</code> in the experiment agent bar to bring the integrated OpenCode agent online for a running experiment, no separate server to start.',
'Type <code>/loop 30m &lt;prompt&gt;</code> in the experiment agent bar to have the agent check in on your training on a recurring interval.',
'Export tagged samples straight to CVAT, Label Studio, or V7 with <code>wl.export_annotations("cvat", tags=["ToReview"])</code> — no custom relabeling script needed.',
'Right-click any curve to add a step note, hide it, load weights from that step, or change its color — no separate panel needed.',
'Export tagged samples straight to CVAT, Label Studio, or V7 with <code>wl.export_annotations("cvat", tags=["ToReview"])</code>, no custom relabeling script needed.',
'Right-click any curve to add a step note, hide it, load weights from that step, or change its color, no separate panel needed.',
'Curves now render error bands and flag outlier steps automatically, so anomalies stand out without manual smoothing.',
'Filter the plots panel with a regex in the search bar to isolate exactly the curves you want.',
'From a live Jupyter or Colab notebook, ask the agent to generate analysis code against your on-training experiment — no need to stop training first.',
'WeightsLab now tracks GPU, CPU, and RAM usage automatically during training and agent runs — check the resource panel, no separate monitoring setup needed.',
'From a live Jupyter or Colab notebook, ask the agent to generate analysis code against your on-training experiment, no need to stop training first.',
'WeightsLab now tracks GPU, CPU, and RAM usage automatically during training and agent runs, check the resource panel, no separate monitoring setup needed.',
];

var INTERVAL = 5000; // ms between rotations
Expand Down Expand Up @@ -60,7 +60,7 @@
}

showTip(idx);
window.addEventListener('resize', avoidCollision);
window.addEventListener('resize', hideIfTooNarrow);

setInterval(function () {
idx = (idx + 1) % TIPS.length;
Expand Down
Loading