What VCF-RDFizer can produce from an RDF aggregate, how each artifact is built, what is verified about it, and where each one runs out of road.
Compression is deliberately not one option. It is three:
| Decision | Flag | Values | Default |
|---|---|---|---|
| How the aggregate is staged | --rdf-storage-mode |
plain, space-optimized |
space-optimized |
| Which raw RDF artifacts to keep | --rdf-compression |
gzip, brotli, none |
gzip,brotli |
| Which queryable representations to build | --representations |
hdt, cottas, none |
hdt |
| How to package those representations | --artifact-compression |
gzip, brotli, none |
none |
Each selector takes a comma-separated list. none must appear alone.
--artifact-compression requires at least one selected representation.
The gzip used by --rdf-storage-mode space-optimized is staging, not
automatically a final artifact. It is retained as the gzip artifact only when
gzip is also selected in --rdf-compression.
For the smallest output: --rdf-compression none, one representation, and
--remove-rdf-storage-output.
| Artifact | Queryable | Built by | Verified by |
|---|---|---|---|
.nt |
with any RDF store | the merge step | Raptor syntax check during validation |
.nt.gz, .nt.br |
after decompression | gzip / brotli |
— |
.hdt + .hdt.index.v1-1 |
yes, directly | hdtc create (+ hdtc index) |
streaming decode + triple-count equality |
.cottas |
yes, directly | pycottas + PyArrow merge |
streaming decode + triple-count equality |
.hdt.gz, .hdt.br, .cottas.gz, .cottas.br |
no | packaging | — |
The packaged forms are archives. Keep the unwrapped .hdt / .cottas if
queries must run without a decompression step.
When HDT or COTTAS is selected, the aggregate is read sequentially and split
into chunks on complete N-Triples or N-Quads line boundaries. Only one uncompressed
chunk exists at a time: it is consumed by both converters and removed before the
next is read. That property is what makes a space-optimized .nt.gz aggregate
usable without ever expanding a second full raw copy.
| Flag | Meaning | Default |
|---|---|---|
--chunk-target-bytes |
target uncompressed bytes per chunk | 512 MiB |
--chunk-min-bytes |
minimum before a chunk group is flushed | 128 MiB |
--chunk-max-bytes |
hard ceiling; boundaries stay on complete lines | 1 GiB |
--hdt-strategy chooses the policy: auto (build chunks and merge with native
hdtc), partitioned (always chunk), single (one rdf2hdt run).
single is a verification path, not a faster alternative. There is no size
regime in which it wins: below --chunk-min-bytes the partitioned path emits
one chunk and its merge is trivial, so the two converge, and above it
partitioned is the only one that survives cohort scale. What single is for is
producing a one-shot HDT of the same graph so the chunk-and-merge result can be
checked against it — identical triple counts and identical query answers are
what justify partitioned generation being the default, rather than it merely
being convenient.
It applies only to an HDT-only run over an uncompressed aggregate, and both other cases are refused rather than silently downgraded:
- a gzip aggregate cannot be read in one pass without expanding a second full
uncompressed copy, so
singleis incompatible withspace-optimizedand needs an explicit--rdf-storage-mode plain; cottasamong the--representationsalways needs bounded chunks, so the partitioned path would run for both representations and the strategy would have no effect. Earlier releases ignored the flag silently here, which meant a--representations hdt,cottas --hdt-strategy singlerun measured partitioned HDT under asinglelabel.
So the one configuration where single does what it says is:
vcf-rdfizer --mode full --input cohort.vcf.gz \
--rdf-storage-mode plain --representations hdt \
--rdf-compression none --hdt-strategy single --out resultsThe whole partitioned stage runs in an ephemeral Docker-managed volume. Chunks, scratch, and merge files never reach the output directory, and the volume is removed on success and on failure alike.
HDT generation, merging, and indexing use the pinned Rust hdtc 1.1.0, not
hdtCat, hdtSearch.sh, or any Java HDT process. This is not a stylistic
preference: the JVM path fails with java.lang.OutOfMemoryError on
cohort-scale graphs, while hdtc create merges chunk HDTs through disk-backed
external sorts and hdtc index streams BitmapTriples to build the
object/predicate orderings. The sidecar it produces is the canonical HDT v1-1
file, <file>.hdt.index.v1-1, readable by hdt-java and hdt-cpp.
Memory is bounded by two environment variables, both defaulting to 512M in
the image and forwarded by the wrapper when set on the host:
HDT_MERGE_MEMORY_LIMIT=2G HDT_INDEX_MEMORY_LIMIT=2G vcf-rdfizer ...Lower values reduce in-memory sort buffers and increase temporary I/O. Both
stages need scratch space in the container's /work, which lives on the
Docker data volume — free space in the output filesystem does not help.
An HDT whose data is readable but whose sidecar could not be built is a
degraded success in full mode: it is validated with the index check skipped,
marked index_status: "failed", reported in reports/index_warnings.json, and
the raw RDF is retained so it can be repaired later with --mode index.
COTTAS is a Parquet-based representation with its index inside the artifact;
there is no sidecar. Chunk conversion uses
pycottas.rdf2cottas(..., disk=True) with a fresh container-local DuckDB
workspace per operation. --cottas-indexes accepts spo (default), sop,
pso, pos, osp, ops, a comma-separated selection, or all; case is
ignored and repeated orders are built once. Dataset orders accept all 24
permutations of spog (such as spog,gspo,pgos), or all-quads for all 24.
The first order uses name.cottas;
additional orders use name.<order>.cottas. Each file contains the whole graph
in one order. Query one copy at a time; loading them together repeats the graph.
The RDF chunk is parsed and deduplicated once. Additional orders sort that
chunk's Parquet data using DuckDB, without repeating RDF parsing or DISTINCT.
Indexes merge sequentially, so adding orders increases disk use and work without
multiplying merge memory. Gzip/Brotli packaging and round-trip checks cover each
copy. Per-index paths, sizes and validation are in JSON details.indexes; the
existing CSV size columns describe the primary copy, with timings for all orders.
--mode compress --rdf dataset.nq (also .nq.gz) preserves named and default
graphs. Every selected order must include g; HDT and triple-only orders are
rejected for datasets. N-Quads uses a streaming RDFLib parser with bounded
Parquet batches because the pinned pycottas parser renames blank nodes per
chunk. This preserves shared blank nodes, literal lexical forms and graph
identity across chunks. Default graphs are stored as NULL, sorted last.
Graph-aware indexes on VCF/N-Triples input add the default graph. Reindexing
also accepts dataset orders and normalizes legacy pycottas DEFAULT values.
When using pycottas 1.1.0 directly, pass a four-term RDFLib tuple to search
for quad patterns; its string-pattern parser ignores the fourth term.
The final merge deliberately does not call pycottas.cat. In version 1.1.0
that runs a global DISTINCT plus ORDER BY through an unbounded in-memory
DuckDB connection, which gets OOM-killed on large condensed graphs. VCF-RDFizer
instead performs a k-way PyArrow merge of the already index-sorted Parquet
chunks: it holds at most COTTAS_MERGE_BATCH_ROWS (default 2048) rows per
input, drops adjacent duplicate triples/quads, and writes the result incrementally.
The same triple in different graphs remains a separate quad in each graph.
There is no graph-wide hash table and no external-sort spill directory; memory
is a function of the batch size and the chunk count, not of the graph size.
A COTTAS failure degrades differently from HDT: because the .cottas file
itself is unusable, dependent COTTAS artifacts are marked not generated rather
than published with a warning.
Every HDT and COTTAS base artifact is verified before packaging and before
raw RDF cleanup: src/validate_compression.py
reads the source triple count, streams the artifact back through its native
decoder, and requires equality. Compression fails closed if the artifact
cannot be decoded or the counts differ. Results and the
source_triples/decoded_triples pair land in the per-run compression JSON and
in the HDT/COTTAS columns of metrics.csv.
Be clear about what this proves and what it does not. A matching triple count
proves the artifact decodes and contains the right number of statements. It
does not prove the statements are the right ones. The stronger claim comes from
running the semantic suite against the decoded artifact — --validate-artifacts hdt,cottas — which re-derives every VCF summary from it. See
validation.md.
--mode index regenerates an existing artifact's index in place. It is the one
deliberate exception to "never overwrite a planned artifact".
| Input | Behaviour |
|---|---|
--hdt file.hdt |
Existing versioned sidecars are moved aside, regenerated, and restored if indexing fails; incomplete replacements are removed first |
--cottas file.cottas |
--cottas-indexes selects the replacement order and optional additional copies; default spo. Changed orders use bounded sorted batches and streaming merges (at most 128 runs per merge). The primary is replaced only after all indexes build successfully; existing additional outputs are refused |
No conversion, packaging, or decompression output is produced. Standalone index mode is strict — unlike the in-run degradation above, a failure is a failure. The same operation runs automatically after each partitioned HDT merge.
--mode decompress decodes raw RDF archives and HDT/COTTAS artifacts. For a
COTTAS dataset, specify --decompress-out <out>/dataset.nq to preserve named graphs;
the default .nt output is accepted only for triple/default-graph data.
The streaming quad decoder omits the graph term for default-graph statements.
A packaged COTTAS is unwrapped inside the container, so the intermediate
unwrapped file never appears on the host. Raw .nq.gz/.nq.br archives keep
their .nq extension when decompressed.
| Situation | Selection |
|---|---|
| Load into an existing triple store | --representations none --rdf-compression gzip |
| Queryable, single artifact, smallest footprint | --representations hdt --rdf-compression none --remove-rdf-storage-output |
| Comparing HDT against COTTAS | --representations hdt,cottas |
| Archival transfer | --artifact-compression brotli on top of the chosen representation |
| Memory-constrained host | space-optimized + partitioned (both default) + lower --chunk-*-bytes |
| Verifying the partitioned merge | plain + --representations hdt + single, compared against partitioned |
- Docker volume space, not output space, is the binding constraint for partitioned merges and HDT indexing. This is the single most common cause of a failed large run, and the error surfaces as an OOM/SIGKILL rather than as a disk message.
exit_code=-9(or137) means the kernel or Docker OOM-killer intervened, not that the RDF is invalid. Checkstderr_tail,max_rss_kb, and the workspace free-space samples instages/partitioned/<sample>.jsonbefore suspecting the data.- COTTAS is the more fragile path. If it cannot fit the available memory
even at a reduced
COTTAS_MERGE_BATCH_ROWS,--representations hdtis independent and remains queryable. - Packaged representations are not queryable, which is easy to forget when
--artifact-compressionis set and--remove-rdf-storage-outputhas removed the alternative. - HDT cannot carry named graphs. COTTAS preserves them with dataset indexes; VCF conversion itself still produces triples in the default graph.
- No incremental update. Adding variants means reconverting and rebuilding the representation from scratch.
- Architecture — where these stages run
- Output and metrics — what each stage records
- Validation — proving a representation still means what the VCF meant
- CLI reference — every flag, with constraints