Skip to content

Latest commit

 

History

History
145 lines (115 loc) · 6.79 KB

File metadata and controls

145 lines (115 loc) · 6.79 KB

VCF-RDFizer documentation

VCF-RDFizer converts VCF files into RDF, optionally into compressed and queryable representations (HDT, COTTAS), and can prove that the result still reproduces the source VCF's semantics.

The top-level README is the task-oriented quick start: install, flags, worked commands. These documents are the explanation — how each part works, why it was built that way, and where it stops working.

Every document here states its own limits. If you only read one page before deciding whether the tool fits your problem, read Limitations.


Start here

If you want to… Read
Understand how the tool is put together Architecture
Know exactly what happens to a VCF Conversion
Decide what to convert into Representations
Trust the output Validation
Know what the tool cannot do Limitations

The whole set

How it works

  • Architecture — the host/container split, what runs where, the three places the split deliberately leaks, the failure policy, and the pinned toolchain.
  • Conversion — VCF → TSV → RDF stage by stage: the awk parser and its blind spots, the RML mapping, the wrapper's own emitters, the IRI templates, datatype and missing-value decisions.
  • Representations — the compression plan's three independent decisions, record-safe chunking, HDT via native hdtc, the bounded COTTAS merge, round-trip verification, and index maintenance.
  • Output and metrics — the output layout, the run_metrics/ tree, input size accounting, progress, interrupts, exit codes.

How to extend it

  • Custom RML mappings — the --rules contract, the vcf-rdfizer-rules CLI, and an honest account of what a custom mapping does not control and what it costs in validation.
  • Data linking — three runnable plug-in tiers, full and post-hoc usage, the authoring CLI, reference/assembly checks, network policy, provenance, and measured implementation limits. The broader design records what remains planned.
  • Privacy policy design — proposal, not implemented. Granular, machine-readable disclosure control over parts of the graph: an ODRL profile with graph selectors, three enforcement tiers, and verification — plus a candid account of why access control is not anonymization when the genotypes are themselves identifiers.
  • Policy attachment v0.1.0 — implemented; walkthrough in examples/policy/. The first slice of the privacy design: ODRL policies attached to any resource or declared graph selection, a partition rule that says what a withheld resource takes with it, per-request release views, and checks — including a VCF-text oracle — that each view withholds exactly what it should. Selectors and partitions are Turtle, not code.

What the graph looks like

  • Sample representations — expanded versus condensed genotype shapes, worked examples, scaling arithmetic, the trade-off, and an assessment of query-time decoding options.
  • VCF coverage matrix — element by element: is it represented, and would a corruption of it be detected. Includes the current mutation score and the vocabulary alignment status.

Whether to trust it

  • Validation — how to run the semantic suite, which artifact and engine to choose, what the thirteen queries and the preflights check, and a candid "what is not tested".
  • Validation methodology — how coverage is measured rather than asserted: the mutation-testing harness, why known gaps are assertions rather than comments, and how to reproduce the score.
  • Validation migration notes — what the move to the VCF Core vocabulary leaves to do in the validation suite, and how the converter's SHACL conformance was verified in the meantime.

Where it is going

  • Limitations — everything the tool cannot do, does badly, or does surprisingly, in one place.
  • Roadmap — what is planned, what is known-broken, and what has been assessed and deliberately rejected.

Reading paths

"I have a VCF and I want RDF." README → Conversion → Sample representations → Representations

"I need to defend this output in a paper." Validation → Validation methodology → VCF coverage matrix → Limitations

"I want to change what RDF comes out." Architecture → Conversion → Custom RML mappings

"I want to add my own domain's links." Data linking → Data linking design → Custom RML mappings → Validation methodology

"I need to release only part of this cohort." Privacy policy design → Sample representations → Conversion §6 → Validation methodology

"A cohort-scale run just failed." Representations §10 → Output and metrics §2 → Limitations §1


A note on how these documents are written

Three conventions, kept deliberately:

  1. Claims are measured, not asserted. Where a document says a defect would be caught, there is a named mutation in test/validation_mutations.py that proves it — and where one would not be caught, that gap is also an assertion, so closing it fails a test rather than passing silently.
  2. Limitations sit next to capabilities, not in a footnote. A section that describes what something does also describes where it stops.
  3. Design rationale is recorded, including for options that were assessed and rejected, so the same ground is not re-covered later.

Related files outside docs/: rules/README.md (the mapping directory), test/README.md (the test suite), scripts/RELEASING.md, and ACKNOWLEDGEMENTS.md. Per-version release notes are published on the releases page.