Skip to content

Latest commit

 

History

History
642 lines (496 loc) · 26.9 KB

File metadata and controls

642 lines (496 loc) · 26.9 KB

Expanded and Condensed Knowledge Representations

Part of the VCF-RDFizer documentation. How these triples are emitted: conversion.md. Where section 7's assessment leads: roadmap.md.

This guide explains the two ways VCF-RDFizer can represent sample genotype information in RDF:

  • expanded, the default representation;
  • condensed, selected with --sample-representation condensed.

The word representation here means the shape of the RDF graph. It does not mean that the input VCF is changed or that genotype values are discarded.

1. Why sample representation matters

A VCF is naturally a wide table. A small part of a multi-sample VCF might look like this:

#CHROM  POS  REF  ALT  FORMAT       SAMPLE_A          SAMPLE_B       SAMPLE_C
20      100  A    G    GT:DP:AD     0/1:42:30,12      0/0:18:18,0   1/1:9:9,0

There is one variant row, followed by one sample column for every individual. The FORMAT column says that each sample value has three fields:

  • GT: genotype;
  • DP: read depth;
  • AD: allele depths.

For a real cohort, the numbers can be much larger. For example, a file with 10,000 variants, 2,500 samples, and three FORMAT fields contains:

10,000 × 2,500 × 3 = 75,000,000 individual FORMAT values

RDF can describe each of those values separately, but doing so creates a very large number of RDF resources and relationships. The two modes let users choose between a graph that is easy to query one value at a time and a graph that is much more economical for large cohorts.

2. Expanded representation: one RDF object per sample value

Expanded mode preserves the original, explicit vocabulary model. For every variant/sample pair, VCF-RDFizer creates a vcfc:SampleCall. For every FORMAT field in that pair, it creates a vcfc:FormatFieldValue.

For the example above, the conceptual graph includes resources like these (the full output has one line per RDF statement):

<file://cohort.vcf#call/1>
    vcfc:hasSampleCall <file://cohort.vcf#sample/1/SAMPLE_A> .

<file://cohort.vcf#sample/1/SAMPLE_A>
    a vcfc:SampleCall ;
    vcfc:sampleId "SAMPLE_A" ;
    vcfc:forSample <file://cohort.vcf#samples/SAMPLE_A> ;
    vcfc:hasFormatValue <file://cohort.vcf#sample/1/SAMPLE_A/fmt/GT> ;
    vcfc:hasGenotype <file://cohort.vcf#sample/1/SAMPLE_A/genotype> .

<file://cohort.vcf#sample/1/SAMPLE_A/fmt/GT>
    a vcfc:FormatFieldValue ;
    vcfc:declaredBy <file://cohort.vcf#header/line/8> ;
    vcfc:fieldValue "0/1" .

<file://cohort.vcf#sample/1/SAMPLE_A/genotype>
    a vcfc:Genotype ;
    vcfc:genotypeString "0/1"^^vcfc:GenotypeString ;
    vcfc:ploidy 2 ;
    vcfc:phasingStatus vcfc:Unphased ;
    vcfc:hasAlleleCall <file://cohort.vcf#sample/1/SAMPLE_A/genotype/call/0> ,
                       <file://cohort.vcf#sample/1/SAMPLE_A/genotype/call/1> .

<file://cohort.vcf#sample/1/SAMPLE_A/genotype/call/1>
    a vcfc:GenotypeAlleleCall ;
    vcfc:callIndex 1 ;
    vcfc:isNoCall false ;
    vcfc:calledAllele <file://cohort.vcf#record/1/allele/1> .

The same structure is repeated for DP and AD, and then repeated again for SAMPLE_B and SAMPLE_C.

What expanded mode provides

Expanded mode makes each value easy to address directly:

  • a consumer can find one sample call without decoding a vector;
  • a consumer can ask for one FORMAT field as one RDF resource, and follow vcfc:declaredBy to the ##FORMAT declaration that defines it;
  • the sample identifier is attached to every sample-call resource, and vcfc:forSample additionally gives it a reusable, file-scoped identity, so a sample is a resource rather than a repeated literal;
  • genotypes are queryable directly. GT is parsed into an ordered vcfc:Genotype: a consumer can filter on vcfc:phasingStatus, count vcfc:ploidy, or join vcfc:calledAllele straight to the record's allele resources — without string-matching "0/1" or "0|1" in SPARQL. FT, PS/PSL, LAA, the copy-number and haplotype keys, and the M/DPM/ADM base modifications get their own resources on the same principle.

The VCF file is marked with vcfc:representationProfile vcfc:ExpandedRepresentation, so a consumer can identify the graph shape explicitly. The SPARQL SHACL profile enforces that a graph never carries both profiles at once.

Expanded growth

Let:

  • V = number of variant records;
  • S = number of samples;
  • F = number of FORMAT keys on a record;
  • p = ploidy of a genotype, and r the positions in it that resolve to an allele (r ≤ p).

Both profiles begin with the same file-level block, which is why it cancels out of every comparison below:

B = 3 + 5S + D

one statement for vcfc:representationProfile, two for the vcfc:SampleSet, five per sample column (hasSample, rdf:type, sampleName, sampleIndex, and the #CHROM line's hasGenotypeColumns), and D for the FORMAT definitions — one statement per distinct key declared by a ##FORMAT line, since the header emitter owns its attributes, or six for an undeclared key the sample emitter has to synthesize. With every key declared, D = F.

Expanded mode then emits, per record and sample:

T_expanded = B + V × S × (4 + Σ c_k)

Four statements for the vcfc:SampleCall — hasSampleCall, rdf:type, sampleId, forSample — plus c_k for each FORMAT key k:

Contribution to c_k Statements When
hasFormatValue, rdf:type, declaredBy 3 always
fieldValue 1 the cell is not empty
fieldValueInteger / fieldValueDecimal 1 single-valued, declared Integer/Float, and parses
vcfc:FieldValueItem 5 per item the declared Number makes positions meaningful (A, R, LA, LR, G, LG, P, and the version's tuple keys); a sixth per item for a tuple key's tupleArity
the genotype layer 5 + 5p + r k is GT

The genotype layer is five statements for the vcfc:Genotype (hasGenotype, rdf:type, genotypeString, ploidy, phasingStatus), five per position for the vcfc:GenotypeAlleleCall (hasAlleleCall, rdf:type, callIndex, isNoCall, phaseIndicator), and one calledAllele per position that names an allele the record actually has. A fully resolvable diploid gives 5 + 10 + 2 = 17; a ./. gives 15, since a no-call calls no allele. GT is a String, so it earns no typed companion and costs 4 + 17 = 21 in total.

In the simplest uniform case — F single-valued Integer/Float keys, no GT — every c_k is 5 and the whole expression collapses to:

T_expanded = 3 + 5S + F + V × S × (4 + 5F)

For the worked example (V=1, S=3, FORMAT=GT:DP:AD on a biallelic record):

Expanded item Count
SampleCall resources 3
FormatFieldValue resources 9
Genotype resources 3
GenotypeAlleleCall resources 6
FieldValueItem resources (AD is Number=R) 6
Per-sample statements (3 × 44) 132
File-level SampleSet + VCFSample statements 17
FORMAT definition statements 3
Representation-profile statement 1
Total 153

Expanded mode buys direct queryability at a cost that multiplies by both samples and FORMAT fields.

3. Condensed representation: shared samples plus ordered vectors

Condensed mode avoids creating a separate RDF resource for every scalar sample value. It describes the sample columns once, then stores the values for each variant and FORMAT key in one ordered vector.

For the same example, the graph first declares a reusable sample set:

<file://cohort.vcf#samples>
    a vcfc:SampleSet ;
    vcfc:hasSample <file://cohort.vcf#samples/SAMPLE_A> ;
    vcfc:hasSample <file://cohort.vcf#samples/SAMPLE_B> ;
    vcfc:hasSample <file://cohort.vcf#samples/SAMPLE_C> .

<file://cohort.vcf#samples/SAMPLE_A>
    a vcfc:VCFSample ;
    vcfc:sampleName "SAMPLE_A" ;
    vcfc:sampleIndex "1" .

The variant then has one cohort matrix. The matrix points back to the sample set and has one vector for each FORMAT key:

<file://cohort.vcf#call/1/matrix>
    a vcfc:CohortCallMatrix ;
    vcfc:appliesToSampleSet <file://cohort.vcf#samples> ;
    vcfc:hasFormatValueVector <file://cohort.vcf#call/1/matrix/fmt/GT> .

<file://cohort.vcf#call/1/matrix/fmt/GT>
    a vcfc:FormatValueVector ;
    vcfc:valueEncoding vcfc:VCFTextVector ;
    vcfc:encodedValues "0/1\t0/0\t1/1" .

The three tab-separated items in encodedValues are in sampleIndex order:

sampleIndex Sample GT value in the vector
1 SAMPLE_A 0/1
2 SAMPLE_B 0/0
3 SAMPLE_C 1/1

The DP vector is "42\t18\t9", and the AD vector is "30,12\t18,0\t9,0". A comma inside one VCF value remains inside that item; it is not treated as a sample separator. If a value is absent, the vector uses . to keep every sample position aligned.

The file is marked with vcfc:representationProfile vcfc:CondensedRepresentation.

What condensed mode provides

  • Sample resources are declared once per file, not once per variant.
  • Each variant has one CohortCallMatrix, rather than one sample-call resource for every sample.
  • Each variant/FORMAT-key pair has one FormatValueVector, rather than one FormatFieldValue resource per sample.
  • The original sample order and all lexical values remain available.
  • A vector-aware consumer can reconstruct sample i by looking up the sample with sampleIndex=i and selecting item i from the corresponding vector.

Condensed mode therefore moves repetition from RDF structure into the literal payload of a vector. It is a storage and graph-shape optimization, not a loss of genotype information.

What it does not carry. Condensed mode stops at the vector: it derives no vcfc:Genotype, vcfc:PhaseSet, vcfc:LocalAlleleSet or vcfc:BaseModification. Those are one resource per sample per record, which is precisely the growth this profile exists to prevent. The values are still all there — inside the vectors — but a SPARQL engine cannot filter on them without a decoder or an application-level vector function. If you need to query genotypes directly, that is the argument for expanded mode.

Everything not per-sample is identical in both profiles: the header layer, the allele layer, structured INFO with its value items, the SV carriers, and the record-level ID/FILTER/FORMAT-key decompositions are all emitted the same way.

Condensed growth

Condensed mode emits the same file-level block B, then three statements per record for the vcfc:CohortCallMatrix — hasCallMatrix, rdf:type, appliesToSampleSet — and five per vector: hasFormatValueVector, rdf:type, declaredBy, valueEncoding, encodedValues.

T_condensed = B + V × (3 + 5F)
            = 3 + 5S + F + V × (3 + 5F)

with every key declared. --sample-data-raw adds one more statement per record.

Two properties follow directly, and they are the point of the profile:

  • S appears only in the file-level term. The S values for a variant and key live inside one encodedValues literal, not as S separate statements.
  • The count is independent of ploidy, of GT, and of how many values a FORMAT key holds. The genotype layer and the vcfc:FieldValueItem decomposition are expanded-only, so a Number=R field costs a condensed graph exactly what a scalar does.

For the worked example (V=1, S=3, FORMAT=GT:DP:AD):

Condensed item Count
Reusable VCFSample resources 3
CohortCallMatrix resources 1
FormatValueVector resources 3
Per-record statements (3 + 5 × 3) 18
File-level SampleSet + VCFSample statements 17
FORMAT definition statements 3
Representation-profile statement 1
Total 39

Because a condensed graph has fixed metadata for the sample set, matrix, and FORMAT definitions, it can be similar in size to—or even slightly larger than— an expanded graph for a tiny file. Its advantage appears when many variants reuse the same sample columns.

4. The scaling difference in numbers

Assume a cohort has:

V = 1,000 variants
S = 2,500 samples
F = 3 FORMAT fields

The difference

Subtracting the two closed forms, the shared file-level block cancels and the whole difference is per-record:

ΔT = T_expanded − T_condensed = V × [ S × (4 + Σ c_k) − (3 + 5F) ]

and the ratio has a limit that depends only on the cohort width:

T_expanded / T_condensed  ──V→∞──▶  S × (4 + Σ c_k) / (3 + 5F)  ──F→∞──▶  S

Condensation removes exactly the per-sample dimension, so the saving in statements approaches the number of samples. Convergence in V is slow while S is large, because the 5S term dominates T_condensed until V ≫ S.

Counts for the example cohort

Expanded totals depend on the FORMAT shape, so three are given. Condensed is 30,506 in every one of them — it does not depend on Σ c_k at all:

FORMAT shape per sample call Expanded Condensed Ratio
3 single-valued typed scalars, no GT 19 47,512,506 30,506 1,557×
GT + 2 typed scalars, diploid 35 87,512,506 30,506 2,869×
GT:DP:AD, diploid, biallelic 44 110,012,506 30,506 3,606×

The limit for the first shape is 2,500 × 19/18 ≈ 2,639×, which the ratio at V = 1,000 has not yet reached.

This does not mean the compressed file shrinks by the same factor. The same scalar values still exist inside the vector literals, and their text length, escaping, dictionary compression, and the rest of the VCF graph affect the final .nt, .hdt, or .cottas size. The large saving is the removal of millions of repeated RDF nodes, predicates, and type statements.

Every constant in these formulas is asserted against the emitters' real output in test/test_representation_growth_unit.py, so a change to any per-resource triple count fails a test that names the resource it belongs to.

5. Why VCF-RDFizer keeps both strategies

The modes serve different downstream needs. Combining them into one universal graph would either make small files unnecessarily complicated or make large cohorts unnecessarily expensive.

Expanded is useful when direct RDF querying is the priority

Expanded mode is a good fit when:

  • the VCF has one sample or a small number of samples;
  • downstream software already expects SampleCall and FormatFieldValue resources;
  • queries need to match one sample value directly, without splitting a vector;
  • maximum vocabulary-level transparency is more important than graph size.

For example, a consumer can navigate directly from a call to the sample whose sampleId is SAMPLE_A, then to its GT FormatFieldValue. Every value is a normal RDF resource and literal.

Condensed is useful when cohort scale is the priority

Condensed mode is a good fit when:

  • there are hundreds or thousands of samples;
  • the same sample columns repeat across many variant rows;
  • materializing one RDF object per sample/variant/FORMAT combination causes excessive memory, disk, or indexing work;
  • downstream software can use sampleIndex and decode tab-separated vectors.

For a 2,500-sample cohort, expanded mode creates 2,500 sample-call resources for every variant. Condensed mode creates the 2,500 sample descriptions once and reuses them for all variants.

The semantic trade-off

Question Expanded Condensed
Where is each sample value? Its own RDF resource An item in an ordered vector literal
How often are samples declared? Repeated per variant/sample call Once per file/sample set
Simple one-value RDF query Easier Requires vector lookup/decoding
Large-cohort graph size Grows with V × S × F Structure grows roughly with S + V × F
All source values retained? Yes Yes
Vocabulary profile ExpandedRepresentation CondensedRepresentation

Neither mode is universally “more correct.” They expose the same underlying VCF information through different graph contracts.

6. How this is implemented in VCF-RDFizer

The command-line choice is explicit:

# Expanded is the default
vcf-rdfizer --mode full --input cohort.vcf.gz \
  --sample-representation expanded \
  --rdf-storage-mode plain --out results

# Condensed is intended for large multi-sample cohorts
vcf-rdfizer --mode full --input cohort.vcf.gz \
  --sample-representation condensed \
  --rdf-storage-mode space-optimized \
  --representations hdt --out results

There is no automatic “switch at N samples” threshold. This is deliberate: the same command should produce the same graph contract every time, and a downstream consumer should not unexpectedly receive a different RDF shape just because a file happened to contain more samples.

For the built-in rules, both modes read the wide sample payload in records.tsv and stream the selected RDF representation directly. The current implementation does not first materialize the enormous helper tables for the built-in maps. That implementation optimization is separate from the semantic choice:

  • expanded still emits the expanded SampleCall/FormatFieldValue graph;
  • condensed still emits the condensed SampleSet/CohortCallMatrix/ FormatValueVector graph.

Custom rules that consume sample_calls.tsv or sample_format_values.tsv retain the materialized helper-table behavior. Such custom helper-table mappings are rejected in condensed mode, because running them alongside the condensed emitter would produce both graph shapes and recreate the expansion that condensed mode is intended to avoid.

7. Assessment of GeoSPARQL and GraphDB SPARQL extensions

GraphDB's SPARQL extensions reference contains several different kinds of functionality. They should not all be treated as ways to read a condensed VCF vector.

7.1 GeoSPARQL geometry functions are not a direct match

The standard geof: functions in the GraphDB reference operate on a geomLiteral, such as a literal with the datatype geo:wktLiteral or geo:gmlLiteral. They perform geometry operations such as distance, buffer, intersection, union, envelope, and spatial relationship tests. GraphDB's additional geoext: functions provide operations such as area, geometry validity, simplification, and Hausdorff distance. These functions are intended for points, lines, polygons, and other spatial objects, not arbitrary ordered text values. See the GeoSPARQL function table and GraphDB GeoSPARQL extensions.

Our condensed value is instead a normal RDF string associated with vcfc:VCFTextVector, for example:

"0/1\t0/0\t1/1"

It is not a WKT geometry.

The analogy with WKT remains useful at the design level: both pack a sequence into one literal. The GeoSPARQL operations themselves, however, are not reusable for genotype-vector extraction.

7.2 The useful GraphDB feature is string splitting

The reference also documents the GraphDB/SPIN magic predicate spif:split. It takes a string and a regular expression and produces one result row for each split item. The implementation is described as using Java's String.split() method. A prototype query for all tokens in condensed vectors could look like this:

PREFIX vcfc: <https://w3id.org/vcf-core/vocab#>
PREFIX spif: <http://spinrdf.org/spif#>

SELECT ?vector ?token WHERE {
  ?vector a vcfc:FormatValueVector ;
          vcfc:valueEncoding vcfc:VCFTextVector ;
          vcfc:encodedValues ?encoded .
  ?token spif:split (STR(?encoded) "\t") .
}

For a GT vector, this would expose 0/1, 0/0, and 1/1 as separate query bindings. It is useful for experimentation, validation, or exporting values to an application.

There is an important limitation: splitting gives the values, but the query also needs the ordinal position of each value in order to identify the sample. The condensed model says that the first token belongs to sampleIndex=1, the second to sampleIndex=2, and so on. The documented spif:split pattern does not itself return that ordinal. GraphDB also documents spif:for, which can generate a sequence of integers, but the reference does not provide a direct “zip these generated indexes with the split results” operation. This means that spif:split alone is not a complete, reliable sample lookup mechanism.

The page also lists RDF-list functions such as list:index and list:length. Those operate on RDF Collections. vcfc:encodedValues is a single string literal, not an RDF Collection, so these functions do not apply unless the data model is changed to materialize every vector item as list structure. That would reintroduce much of the per-item RDF overhead that condensed mode was designed to avoid. GraphDB's helper:tuple and helper:iterate functions similarly operate on internal query-time lists; they do not automatically parse a persisted VCF vector literal.

7.3 What is useful now and what is not

GraphDB feature Usefulness for VCFTextVector Assessment
geof:* GeoSPARQL functions Geometry calculations Not applicable to genotype text vectors
geoext:* geometry extensions Geometry validity and transformations Not applicable unless a future dataset contains real sample geometries
spif:split Split tab-separated vector text Useful prototype, but loses the token ordinal
spif:for Generate expected sample positions Helpful support function, but does not pair positions with split tokens
list:index / list:length Index RDF Collections Not applicable to the current string encoding
helper:tuple / helper:iterate Work with internal query lists Not a persisted-vector parser
A custom vcfc: vector function Return value at a sample index Most direct future SPARQL integration
Application-side parsing Decode one selected vector after retrieval Best portable near-term approach

GraphDB explicitly labels these as extensions beyond the W3C SPARQL specification, so a query using spif:* or a custom function would be tied to GraphDB (or to another engine that implements the same extension). That can be acceptable for a GraphDB deployment, but it should not be presented as a portable SPARQL solution. These are query-time GraphDB operations; using them would not change VCF-RDFizer's native HDT creation or indexing path, but their memory and latency behavior would still need to be benchmarked inside the GraphDB server.

7.4 Recommended future extraction design

If query-time sample lookup becomes an important use case, the most useful addition would be an extractor that understands the VCF-RDFizer metadata and returns both the position and the value. Conceptually, it could have a function or magic-predicate contract like:

vcfc:vectorValue(?vector, ?sampleIndex) → ?value

or:

?vector vcfc:valueAt (?sampleIndex ?value)

For a sample-aware form, the function could accept the vcfc:VCFSample IRI instead of an integer, resolve that resource's sampleIndex, and then extract the matching token. Returning the index as well as the value would make it possible to validate that a vector has the expected number of positions.

A production implementation should also consider these safeguards:

  1. Verify that the vector belongs to the expected SampleSet through its CohortCallMatrix.
  2. Check that the requested index is within the sample-set size.
  3. Preserve . as the VCF missing-value marker rather than silently turning it into an unbound result.
  4. Treat tab as the vector separator and keep commas inside a value such as AD=30,12.
  5. Validate that the vector contains exactly one item per declared sample.
  6. Expose the function through a GraphDB plugin only when GraphDB-specific deployment is acceptable; otherwise provide a small application-side decoder or a materialized per-value projection.

There is also a performance question. With 1,000 variants, 2,500 samples, and three FORMAT keys, condensed mode stores 3,000 vectors but each vector contains 2,500 values. Expanding every vector in a SPARQL query would produce up to

1,000 × 3 × 2,500 = 7,500,000 query result rows

before filtering to a particular sample. For a targeted lookup, a function that extracts one position is preferable to spif:split over the entire vector. For large cohort-wide analyses, an external decoder or a precomputed analytical table may be more appropriate than forcing a graph database to emit millions of token bindings.

7.5 Overall conclusion

The GraphDB work is useful as a source of implementation ideas, but not as a drop-in GeoSPARQL solution:

  • GeoSPARQL geometry functions: no, they should not be used for genotype vectors.
  • spif:split: yes, as a prototype tokenizer and validation aid.
  • spif:split plus spif:for: potentially useful building blocks, but not sufficient for a trustworthy sample/value join without an ordinal pairing mechanism.
  • Custom vector extraction function or application decoder: yes, this is the most promising future direction.

The condensed model should therefore remain as it is for storage efficiency, while future query support should be added as a separate decoding/projection layer rather than changing genotype vectors into geometries or RDF Lists.

8. Practical choice

Use this short rule of thumb:

Small or single-sample VCF + simple RDF queries  → expanded
Large multi-sample cohort + storage/scale focus  → condensed

If a consumer is unsure which graph it received, inspect the file’s vcfc:representationProfile value before querying. That profile is the explicit signal that tells the consumer whether sample values are represented as individual resources or as ordered vectors.

Glossary

  • Variant record: one row describing a genomic position and its alleles.
  • Sample: one individual or biological sample represented by a VCF column.
  • FORMAT key: a field name such as GT, DP, or AD describing one part of a sample’s value.
  • RDF resource: a named graph object that can have properties, such as one SampleCall or one FormatFieldValue.
  • Vector: an ordered list of values. In condensed mode, values are stored as one tab-separated VCFTextVector literal in sampleIndex order.
  • Graph shape: which resources and relationships are present, independent of the underlying biological values.

See also