What is planned, what is merely known-to-be-wrong, and what has been assessed and deliberately rejected. Every item links to the document where it is described in full.
This is a working document, not a commitment. Items are ordered by how much they block something else, not by effort.
Both items that previously blocked publication are closed. They are kept here briefly because the reasoning is referenced elsewhere.
The 17 condensed-representation terms the conversion emitted without an
ontology behind them — CohortCallMatrix, FormatValueVector, SampleSet,
VCFSample, VCFTextVector, representationProfile, sampleIndex and the
rest — are all defined in VCF Core, with SHACL shapes. Condensed graphs are
ontology-backed and citable as linked data.
vcfc:VCFRecordShape now accepts xsd:string or vcfc:Null for vcfc:alt, so
a record with ALT=. satisfies both the shape and vcfc:missingValuePolicy.
vcfc:VariantCallShape likewise accepts the full VCF Float lexical space for
qual.
The suite passes: 707 tests, mutation score 96/113 (85%) across 60 mutations.
The runner and the wrapper now share one vocabulary module instead of mirroring
each other, and the census derives every new resource family from the VCF. Ten
of the eighteen new mutations are recorded gaps, each with what would close it.
What changed and what remains:
validation-migration-notes.md.
Checking the output against the vocabulary's own SHACL profiles found four real
modelling bugs during the migration, all of the same kind: a resource typed into
a class whose shape requires a property the record did not supply. That is a
cheap, strong invariant and it should not be a manual step. It needs pyshacl
and rdflib as dev dependencies and access to the vocabulary's ontology/ and
shacl/ directories.
VCF Core now supplies vcfc:IntegerLiteralShape, which accepts xsd:integer
and every integer-derived XSD datatype with the bound applied numerically. The
ontology range and the shapes no longer disagree.
validation_runner.py supports --mapping-policy report-only, which is exactly
what a custom --rules mapping needs: it records the mapping-dependent queries
(q09–q13) without failing on them. The vcf-rdfizer wrapper never passes
it, so the runner always executes strict and a correct custom-mapping
conversion is reported as MISMATCH.
Fix: expose a --mapping-policy flag on the wrapper, and default it to
report-only automatically whenever --rules is not the shipped default.
Small change; it makes the custom-mapping extension point actually usable.
Detail.
From vcf-coverage.md:
vcfc:contigCountis counted, not read. A wrong derived contig total is undetected. Needs the value in a comparison, not just the predicate in the census.- Header line values are not compared.
q08compares how many lines carry each key andq10their types;vcfc:headerValueitself is only counted. The structured attributes that matter are already covered.
Both are query-set additions with matching mutation-catalogue entries, so
closing either one will fail the mutation harness until
vcf-coverage.md is updated — by design.
An initial implementation now connects the graph through declarative, distributable linker plug-ins, with examples of all three tiers: dbSNP token links, a synthetic GFF3 interval bundle, and an Ensembl API resolver.
Full design, including the three join strategies, the three plugin tiers, the
network safeguards, the provenance model and the build order, is in
datalinking-design.md. The implemented contract and
commands are in datalinking.md. Allele joins/normalization,
link merging, and plug-in validation/mutation auto-discovery remain planned;
the manifest vocabulary is provisional.
The hardest part is not the plugin system — it is allele normalization (§9 of that document), because an under-matching join looks exactly like a true negative.
Condensed mode stores each FORMAT key as one tab-separated vector, which is what makes cohort scale tractable. Reconstructing sample i's value at query time needs a split operation SPARQL does not portably provide.
The options have been assessed in detail in
sample-representation-guide.md:
| Option | Verdict |
|---|---|
| GeoSPARQL geometry functions | No. Genotype vectors are not geometries; the analogy does not survive contact |
GraphDB spif:split |
Useful as a prototype tokenizer and validation aid, not as a production join |
spif:split + spif:for |
Building blocks, but no ordinal pairing mechanism, so no trustworthy sample↔value join |
| A custom vector-extraction function, or an application-side decoder | The promising direction |
The recommended shape is a function that takes a vector and a sample index (or a
VCFSample IRI) and returns both the position and the value, with the six
safeguards listed in that section. Crucially, this belongs in a separate
decoding/projection layer — the condensed model itself should not change.
There is also a scale caveat that any implementation has to respect: expanding every vector in a cohort-wide query produces millions of bindings before filtering. A targeted single-position extractor is the right primitive; a whole-graph expansion is not.
Today the tool has exactly one disclosure setting: convert everything. There is no way to release a cohort graph with some participants withheld, some regions degraded, or some fields suppressed, and no machine-readable record of what a given artifact was permitted to contain.
The proposal is an ODRL profile whose assets are graph selectors (by class,
predicate, sample, genomic region, declared field or pattern), compiled to a
release plan and enforced at the cheapest available point in the existing
pipeline — TSV pre-filtering, emitter-time filtering, or a post-hoc pass.
Full design in privacy-policy-design.md. The first
slice, a v0.1.0 demonstrator over single-sample fixtures, is specified in
policy-demonstrator.md; its §10 maps later versions
onto the full design's build order.
Two findings from that design are worth surfacing here because they affect work outside it:
ParsedSampleRecorddoes not carryCHROMorPOS._parse_rowreads columns 0, 1, 7, 9, 10 and −1. Any region-scoped feature evaluated at emission time needs both, and adding them is two fields and two index reads.- Condensed mode is not per-sample enforceable by triple filtering. One
FormatValueVectorliteral holds every participant's value for a FORMAT key, so the unit of protection is finer than the unit of storage. Redaction has to rewrite the literal, and a single masked position is itself disclosive.
The honest framing, which the design keeps throughout: this is governed release, not anonymization. Genotypes identify people, and no access-control layer changes that.
Stated so the absence reads as a decision rather than an oversight.
| Not planned | Why |
|---|---|
| A Docker-free installation path | The pinned image is what makes results reproducible across machines |
| Variant normalization inside the conversion | The graph is a faithful transcription of the file; normalization is a linking-time concern |
| Distributed or multi-node execution | Out of scope; chunking already bounds memory on one machine |
| Incremental graph update | Would require identity and provenance machinery the current model does not have |
| Clinical interpretation or pathogenicity assertion | The tool transcribes; it is not a clinical authority |
| Differential privacy on aggregate queries | Formal guarantees need a persistent per-recipient budget ledger; without one it is decoration. See privacy-policy-design.md §9 |
| Deciding whether a release is legally compliant | A tool can implement and evidence a policy; it cannot make a data-protection determination |
.bcf input |
Would pull htslib into the parsing path; convert with bcftools first |
- Limitations — the current state, honestly
- Data linking design — the largest planned addition
- Privacy policy design — governed release over parts of the graph
- Validation methodology — how coverage is measured, so gaps stay falsifiable
- Releases — what has actually shipped