Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@ predictions/
predictions.csv
*.shards/
clusters.csv
*.duckdb

# Local working notes, not part of the repo
CLAUDE.md
Expand Down
1 change: 1 addition & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -108,6 +108,7 @@ result = linker.cluster("predictions.parquet", truth="sample.truth.csv")
print(result.quality)
```

`cpplink-viewer` (`pip install "cpplink[viewer]"`) then serves the clusters as a page: each cluster's members side by side with every disagreeing cell highlighted, and every prediction behind it as the ledger the scorer produced.
Between `init` and `estimate` sit the diagnostics that cost seconds and decide the quality of the result: `profile` says what each column can be worth before any model exists, `levels` places the fuzzy thresholds from the data, `explain-blocking` prices every source without enumerating a pair, and `recall` measures what blocking reaches.
See [Getting started](https://4ment.github.io/cpplink/getting-started/) for the whole pipeline with its output explained, [Commands](https://4ment.github.io/cpplink/commands/) for every command at a glance, the [schema reference](https://4ment.github.io/cpplink/reference/schema/) for every field, and [From Python](https://4ment.github.io/cpplink/python/) for the package.

Expand Down
4 changes: 4 additions & 0 deletions docs/commands/cluster.md
Original file line number Diff line number Diff line change
Expand Up @@ -223,6 +223,10 @@ The `cluster_id` is the representative record's own `unique_id`, so the output s
record the others collapse onto. Only records in a cluster of at least `--min-size` are
written; singletons are excluded by default.

## Reading the clusters

[The cluster viewer](../viewer.md) serves this file beside the predictions and their waterfalls: the members of each cluster side by side with every disagreeing cell highlighted, and every prediction touching it, including the ones clustering at a higher threshold overruled.

## Cost

The union–find is `uint32` parent plus `uint8` rank — **5 bytes a record**, 100 MB at 20M rows —
Expand Down
4 changes: 2 additions & 2 deletions docs/commands/explain.md
Original file line number Diff line number Diff line change
Expand Up @@ -216,8 +216,8 @@ A level's label, `m` and `u` are the model's rather than the pair's, so the file
Ids are resolved through a sorted index over the id column, 4 bytes a record, so a file of millions of predictions costs one load and one pass.
A prediction naming an id no record holds is skipped and counted, and so is one whose stored `gamma` the comparisons no longer produce, which means the schema changed after `predict` ran and the ledgers explain today's schema rather than the file's weights.

This is what the cluster viewers in `tools/` draw from: `tools/cluster_view.py --waterfalls` embeds a ledger per prediction, and `tools/cluster_server.py --waterfalls` loads the file beside the predictions so a click is one lookup.
Nothing in either recomputes a bit of the weight, and neither runs this binary.
This is what [the cluster viewer](../viewer.md) draws from: `cpplink-viewer --waterfalls` loads the file beside the predictions so a click is one lookup.
Nothing in it recomputes a bit of the weight, and it does not run this binary.

## See also

Expand Down
6 changes: 6 additions & 0 deletions docs/python.md
Original file line number Diff line number Diff line change
Expand Up @@ -168,6 +168,12 @@ None of them calls back into Python, so another thread can run while they do.

Errors the core reports as `false` and a message become `cpplink.Error`, a `RuntimeError` carrying the core's own text.

## The cluster viewer

`cpplink-viewer`, the `cpplink_viewer` package in the same wheel, serves the clusters a run produced as a page; `pip install "cpplink[viewer]"` adds DuckDB, its one dependency beyond the binding's.
It reads the files the pipeline wrote and never the compiled module, so it also runs from a checkout with nothing built.
See [The cluster viewer](viewer.md).

## Not in this version

- **In-memory input.** The `Linker` reads parquet files, as the command does. A pyarrow, polars or pandas table through `__arrow_c_stream__` is the next step, and needs the loader split so a table can be appended without a file.
Expand Down
67 changes: 67 additions & 0 deletions docs/viewer.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,67 @@
# The cluster viewer

`cpplink-viewer` serves a page over the clusters a run produced: a list of clusters on the left, and for the one picked, its members side by side with every disagreeing cell highlighted, and every prediction touching it with the one picked drawn as the ledger the scorer produced.
It reads what the pipeline wrote and runs none of it: the clusters and the predictions come from the files, the waterfalls from `explain --predictions`, and nothing in the viewer recomputes a bit of a weight.

```sh
cpplink predict --schema schema.json --model model.json --out predictions.parquet sample.parquet
cpplink cluster --schema schema.json --predictions predictions.parquet --out clusters.csv sample.parquet
cpplink explain --schema schema.json --model model.json --predictions predictions.parquet --out waterfalls.parquet sample.parquet
cpplink-viewer --schema schema.json --clusters clusters.csv --predictions predictions.parquet \
--waterfalls waterfalls.parquet --model model.json --truth sample.truth.csv --open sample.parquet
```

The page is served at `http://127.0.0.1:8770/` (`--port` moves it, `--open` opens a browser on it).

## Install

The viewer is the `cpplink_viewer` package, shipped in the same wheel as the binding, with DuckDB as its one extra dependency:

```sh
pip install "cpplink[viewer]"
```

It never imports the compiled module, so it also runs from a checkout with nothing built, against a run the command line produced:

```sh
pip install duckdb pyarrow numpy
PYTHONPATH=python python -m cpplink_viewer --schema schema.json --clusters clusters.csv sample.parquet
```

## What it reads

| Option | File | Written by |
|---|---|---|
| `data` | the parquet file(s) the run read, in order | you |
| `--schema` | the schema, for the id column and the columns to show | `init` |
| `--clusters` | `unique_id, cluster_id, cluster_size`, csv or parquet | `cluster --out` |
| `--predictions` | one csv or parquet file, or the shard directory | `predict --out` |
| `--waterfalls` | one wide row per prediction, csv or parquet | `explain --predictions --out` |
| `--model` | the level labels and rates the waterfall shows | `estimate --out` |
| `--truth` | known pairs, to colour members by true entity | you, or `gen-sample --truth` |

Only `--schema` and the data are required, plus one of `--clusters` or `--threshold`.
Without `--predictions` the page shows the members and nothing about why they are together; without `--waterfalls` and `--model` it shows each prediction's weight but not the ledger behind it.
A shard directory names records by row, so it costs one pass over the id column that a merged file does not.

## The cache

The first run builds a DuckDB file (`--cache`, `cpplink_viewer.duckdb` by default) holding only what the page can ever show: the clustered records, never the singletons, their predictions, the waterfalls of those predictions, and one row of statistics per cluster.
At 20M records with 2.9M of them clustered the parquet is read exactly once and every later request touches a table seven times smaller than the file.
The cache is keyed on every input's path, size and modification time and on the options that shape it, so a changed input rebuilds it and an unchanged one is reused; `--rebuild` forces it.
`--memory` and `--threads` are DuckDB's limits while building.

## Thresholds and rejected predictions

`--threshold` keeps only the predictions at or above it.
With `--clusters`, the file is taken as what `cluster --threshold` wrote at that threshold and read as it stands.
Without one, the predictions are clustered here by the same union-find in the same order, so the cache holds exactly the partition that command would write, named by the same representatives; the test suite holds the two to the same file.

A prediction whose two records clustering put in different clusters, above the write threshold and below the clustering one, is the prediction most worth reading, so it is kept and listed under both of its clusters as `rejected`, each naming the other cluster its second record went to.
`--min-size` and `--max-size` bound the clusters listed; `--max-rows` bounds the members shown per cluster, the rest being counted only.

## The list

The list is sorted by disagreement (columns whose values differ within the cluster), size, weakest prediction or id, and filtered to every cluster, the split ones, the ones held together only transitively (fewer predictions than pairs), or the ones mixing entities against the truth file.
The search box reads any value, one column, the cluster id, or a comma-separated list of record ids matched whole, in which case it says which of them named a record.
The `network` checkbox draws the selected cluster's predictions as a graph, which is where a chain shows itself as a chain.
2 changes: 2 additions & 0 deletions environment.yml
Original file line number Diff line number Diff line change
Expand Up @@ -21,3 +21,5 @@ dependencies:
- numpy
- pytest
- ruff
# The viewer's cache; the `viewer` extra of the wheel.
- python-duckdb
1 change: 1 addition & 0 deletions mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,7 @@ nav:
- Home: index.md
- Getting started: getting-started.md
- From Python: python.md
- The cluster viewer: viewer.md
- Concepts:
- The model: model.md
- Comparisons: comparisons.md
Expand Down
9 changes: 7 additions & 2 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -14,6 +14,11 @@ dependencies = ["numpy", "pyarrow"]

[project.optional-dependencies]
test = ["pytest"]
# `cpplink-viewer` serves the clusters of a run over a DuckDB cache.
viewer = ["duckdb"]

[project.scripts]
cpplink-viewer = "cpplink_viewer.server:main"

[project.urls]
Homepage = "https://github.com/4ment/cpplink"
Expand All @@ -26,7 +31,7 @@ metadata.version.input = "CMakeLists.txt"
metadata.version.regex = 'project\(cpplink VERSION (?P<value>[0-9.]+)'
cmake.args = ["-DCPPLINK_BUILD_PYTHON=ON"]
cmake.build-type = "Release"
wheel.packages = ["python/cpplink"]
wheel.packages = ["python/cpplink", "python/cpplink_viewer"]
# Keep the extension's build out of build/, which is the CLI's.
build-dir = "build/python"

Expand Down Expand Up @@ -62,4 +67,4 @@ indent-style = "space"
line-ending = "auto"

[tool.ruff.lint.isort]
known-first-party = ["cpplink"]
known-first-party = ["cpplink", "cpplink_viewer"]
13 changes: 13 additions & 0 deletions python/cpplink_viewer/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,13 @@
# Copyright 2026 Mathieu Fourment
# SPDX-License-Identifier: MIT
"""``cpplink_viewer``: the cluster viewer, served over a DuckDB cache of a run.

It reads what a run wrote and never the compiled module, so it runs from a
checkout with no build (``PYTHONPATH=python python -m cpplink_viewer``) as well
as from the wheel (``cpplink-viewer``). DuckDB is the ``viewer`` extra.
"""

from .cache import Options, open_cache
from .server import Viewer, main, make_server

__all__ = ["Options", "Viewer", "main", "make_server", "open_cache"]
10 changes: 10 additions & 0 deletions python/cpplink_viewer/__main__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,10 @@
# Copyright 2026 Mathieu Fourment
# SPDX-License-Identifier: MIT
"""``python -m cpplink_viewer [options] data.parquet``: the viewer, served."""

import sys

from .server import main

if __name__ == "__main__":
sys.exit(main())
Loading
Loading