VCF-RDFizer converts VCF files into RDF, optionally into compressed and queryable representations (HDT, COTTAS), and can prove that the result still reproduces the source VCF's semantics.
The top-level README is the task-oriented quick start: install, flags, worked commands. These documents are the explanation — how each part works, why it was built that way, and where it stops working.
Every document here states its own limits. If you only read one page before deciding whether the tool fits your problem, read Limitations.
| If you want to… | Read |
|---|---|
| Understand how the tool is put together | Architecture |
| Know exactly what happens to a VCF | Conversion |
| Decide what to convert into | Representations |
| Trust the output | Validation |
| Know what the tool cannot do | Limitations |
- Architecture — the host/container split, what runs where, the three places the split deliberately leaks, the failure policy, and the pinned toolchain.
- Conversion — VCF → TSV → RDF stage by stage: the
awkparser and its blind spots, the RML mapping, the wrapper's own emitters, the IRI templates, datatype and missing-value decisions. - Representations — the compression plan's three
independent decisions, record-safe chunking, HDT via native
hdtc, the bounded COTTAS merge, round-trip verification, and index maintenance. - Output and metrics — the output layout, the
run_metrics/tree, input size accounting, progress, interrupts, exit codes.
- Custom RML mappings — the
--rulescontract, thevcf-rdfizer-rulesCLI, and an honest account of what a custom mapping does not control and what it costs in validation. - Data linking — three runnable plug-in tiers, full and post-hoc usage, the authoring CLI, reference/assembly checks, network policy, provenance, and measured implementation limits. The broader design records what remains planned.
- Privacy policy design — proposal, not implemented. Granular, machine-readable disclosure control over parts of the graph: an ODRL profile with graph selectors, three enforcement tiers, and verification — plus a candid account of why access control is not anonymization when the genotypes are themselves identifiers.
- Policy attachment v0.1.0 — implemented;
walkthrough in
examples/policy/. The first slice of the privacy design: ODRL policies attached to any resource or declared graph selection, a partition rule that says what a withheld resource takes with it, per-request release views, and checks — including a VCF-text oracle — that each view withholds exactly what it should. Selectors and partitions are Turtle, not code.
- Sample representations — expanded versus condensed genotype shapes, worked examples, scaling arithmetic, the trade-off, and an assessment of query-time decoding options.
- VCF coverage matrix — element by element: is it represented, and would a corruption of it be detected. Includes the current mutation score and the vocabulary alignment status.
- Validation — how to run the semantic suite, which artifact and engine to choose, what the thirteen queries and the preflights check, and a candid "what is not tested".
- Validation methodology — how coverage is measured rather than asserted: the mutation-testing harness, why known gaps are assertions rather than comments, and how to reproduce the score.
- Validation migration notes — what the move to the VCF Core vocabulary leaves to do in the validation suite, and how the converter's SHACL conformance was verified in the meantime.
- Limitations — everything the tool cannot do, does badly, or does surprisingly, in one place.
- Roadmap — what is planned, what is known-broken, and what has been assessed and deliberately rejected.
"I have a VCF and I want RDF." README → Conversion → Sample representations → Representations
"I need to defend this output in a paper." Validation → Validation methodology → VCF coverage matrix → Limitations
"I want to change what RDF comes out." Architecture → Conversion → Custom RML mappings
"I want to add my own domain's links." Data linking → Data linking design → Custom RML mappings → Validation methodology
"I need to release only part of this cohort." Privacy policy design → Sample representations → Conversion §6 → Validation methodology
"A cohort-scale run just failed." Representations §10 → Output and metrics §2 → Limitations §1
Three conventions, kept deliberately:
- Claims are measured, not asserted. Where a document says a defect would
be caught, there is a named mutation in
test/validation_mutations.pythat proves it — and where one would not be caught, that gap is also an assertion, so closing it fails a test rather than passing silently. - Limitations sit next to capabilities, not in a footnote. A section that describes what something does also describes where it stops.
- Design rationale is recorded, including for options that were assessed and rejected, so the same ground is not re-covered later.
Related files outside docs/: rules/README.md (the
mapping directory), test/README.md (the test suite),
scripts/RELEASING.md, and
ACKNOWLEDGEMENTS.md. Per-version release notes are
published on the releases page.