Part of the VCF-RDFizer documentation. How these triples are emitted:
conversion.md. Where section 7's assessment leads:
roadmap.md.
This guide explains the two ways VCF-RDFizer can represent sample genotype information in RDF:
- expanded, the default representation;
- condensed, selected with
--sample-representation condensed.
The word representation here means the shape of the RDF graph. It does not mean that the input VCF is changed or that genotype values are discarded.
A VCF is naturally a wide table. A small part of a multi-sample VCF might look like this:
#CHROM POS REF ALT FORMAT SAMPLE_A SAMPLE_B SAMPLE_C
20 100 A G GT:DP:AD 0/1:42:30,12 0/0:18:18,0 1/1:9:9,0
There is one variant row, followed by one sample column for every individual.
The FORMAT column says that each sample value has three fields:
GT: genotype;DP: read depth;AD: allele depths.
For a real cohort, the numbers can be much larger. For example, a file with 10,000 variants, 2,500 samples, and three FORMAT fields contains:
10,000 × 2,500 × 3 = 75,000,000 individual FORMAT values
RDF can describe each of those values separately, but doing so creates a very large number of RDF resources and relationships. The two modes let users choose between a graph that is easy to query one value at a time and a graph that is much more economical for large cohorts.
Expanded mode preserves the original, explicit vocabulary model. For every
variant/sample pair, VCF-RDFizer creates a vcfc:SampleCall. For every FORMAT
field in that pair, it creates a vcfc:FormatFieldValue.
For the example above, the conceptual graph includes resources like these (the full output has one line per RDF statement):
<file://cohort.vcf#call/1>
vcfc:hasSampleCall <file://cohort.vcf#sample/1/SAMPLE_A> .
<file://cohort.vcf#sample/1/SAMPLE_A>
a vcfc:SampleCall ;
vcfc:sampleId "SAMPLE_A" ;
vcfc:forSample <file://cohort.vcf#samples/SAMPLE_A> ;
vcfc:hasFormatValue <file://cohort.vcf#sample/1/SAMPLE_A/fmt/GT> ;
vcfc:hasGenotype <file://cohort.vcf#sample/1/SAMPLE_A/genotype> .
<file://cohort.vcf#sample/1/SAMPLE_A/fmt/GT>
a vcfc:FormatFieldValue ;
vcfc:declaredBy <file://cohort.vcf#header/line/8> ;
vcfc:fieldValue "0/1" .
<file://cohort.vcf#sample/1/SAMPLE_A/genotype>
a vcfc:Genotype ;
vcfc:genotypeString "0/1"^^vcfc:GenotypeString ;
vcfc:ploidy 2 ;
vcfc:phasingStatus vcfc:Unphased ;
vcfc:hasAlleleCall <file://cohort.vcf#sample/1/SAMPLE_A/genotype/call/0> ,
<file://cohort.vcf#sample/1/SAMPLE_A/genotype/call/1> .
<file://cohort.vcf#sample/1/SAMPLE_A/genotype/call/1>
a vcfc:GenotypeAlleleCall ;
vcfc:callIndex 1 ;
vcfc:isNoCall false ;
vcfc:calledAllele <file://cohort.vcf#record/1/allele/1> .
The same structure is repeated for DP and AD, and then repeated again for
SAMPLE_B and SAMPLE_C.
Expanded mode makes each value easy to address directly:
- a consumer can find one sample call without decoding a vector;
- a consumer can ask for one FORMAT field as one RDF resource, and follow
vcfc:declaredByto the##FORMATdeclaration that defines it; - the sample identifier is attached to every sample-call resource, and
vcfc:forSampleadditionally gives it a reusable, file-scoped identity, so a sample is a resource rather than a repeated literal; - genotypes are queryable directly.
GTis parsed into an orderedvcfc:Genotype: a consumer can filter onvcfc:phasingStatus, countvcfc:ploidy, or joinvcfc:calledAllelestraight to the record's allele resources — without string-matching"0/1"or"0|1"in SPARQL.FT,PS/PSL,LAA, the copy-number and haplotype keys, and theM/DPM/ADMbase modifications get their own resources on the same principle.
The VCF file is marked with
vcfc:representationProfile vcfc:ExpandedRepresentation, so a consumer can
identify the graph shape explicitly. The SPARQL SHACL profile enforces that a
graph never carries both profiles at once.
Let:
V= number of variant records;S= number of samples;F= number of FORMAT keys on a record;p= ploidy of a genotype, andrthe positions in it that resolve to an allele (r ≤ p).
Both profiles begin with the same file-level block, which is why it cancels out of every comparison below:
B = 3 + 5S + D
one statement for vcfc:representationProfile, two for the vcfc:SampleSet,
five per sample column (hasSample, rdf:type, sampleName, sampleIndex,
and the #CHROM line's hasGenotypeColumns), and D for the FORMAT
definitions — one statement per distinct key declared by a ##FORMAT line,
since the header emitter owns its attributes, or six for an undeclared key
the sample emitter has to synthesize. With every key declared, D = F.
Expanded mode then emits, per record and sample:
T_expanded = B + V × S × (4 + Σ c_k)
Four statements for the vcfc:SampleCall — hasSampleCall, rdf:type,
sampleId, forSample — plus c_k for each FORMAT key k:
Contribution to c_k |
Statements | When |
|---|---|---|
hasFormatValue, rdf:type, declaredBy |
3 | always |
fieldValue |
1 | the cell is not empty |
fieldValueInteger / fieldValueDecimal |
1 | single-valued, declared Integer/Float, and parses |
vcfc:FieldValueItem |
5 per item | the declared Number makes positions meaningful (A, R, LA, LR, G, LG, P, and the version's tuple keys); a sixth per item for a tuple key's tupleArity |
| the genotype layer | 5 + 5p + r |
k is GT |
The genotype layer is five statements for the vcfc:Genotype
(hasGenotype, rdf:type, genotypeString, ploidy, phasingStatus), five
per position for the vcfc:GenotypeAlleleCall (hasAlleleCall, rdf:type,
callIndex, isNoCall, phaseIndicator), and one calledAllele per position
that names an allele the record actually has. A fully resolvable diploid gives
5 + 10 + 2 = 17; a ./. gives 15, since a no-call calls no allele. GT is a
String, so it earns no typed companion and costs 4 + 17 = 21 in total.
In the simplest uniform case — F single-valued Integer/Float keys, no
GT — every c_k is 5 and the whole expression collapses to:
T_expanded = 3 + 5S + F + V × S × (4 + 5F)
For the worked example (V=1, S=3, FORMAT=GT:DP:AD on a biallelic record):
| Expanded item | Count |
|---|---|
SampleCall resources |
3 |
FormatFieldValue resources |
9 |
Genotype resources |
3 |
GenotypeAlleleCall resources |
6 |
FieldValueItem resources (AD is Number=R) |
6 |
Per-sample statements (3 × 44) |
132 |
File-level SampleSet + VCFSample statements |
17 |
| FORMAT definition statements | 3 |
| Representation-profile statement | 1 |
| Total | 153 |
Expanded mode buys direct queryability at a cost that multiplies by both samples and FORMAT fields.
Condensed mode avoids creating a separate RDF resource for every scalar sample value. It describes the sample columns once, then stores the values for each variant and FORMAT key in one ordered vector.
For the same example, the graph first declares a reusable sample set:
<file://cohort.vcf#samples>
a vcfc:SampleSet ;
vcfc:hasSample <file://cohort.vcf#samples/SAMPLE_A> ;
vcfc:hasSample <file://cohort.vcf#samples/SAMPLE_B> ;
vcfc:hasSample <file://cohort.vcf#samples/SAMPLE_C> .
<file://cohort.vcf#samples/SAMPLE_A>
a vcfc:VCFSample ;
vcfc:sampleName "SAMPLE_A" ;
vcfc:sampleIndex "1" .
The variant then has one cohort matrix. The matrix points back to the sample set and has one vector for each FORMAT key:
<file://cohort.vcf#call/1/matrix>
a vcfc:CohortCallMatrix ;
vcfc:appliesToSampleSet <file://cohort.vcf#samples> ;
vcfc:hasFormatValueVector <file://cohort.vcf#call/1/matrix/fmt/GT> .
<file://cohort.vcf#call/1/matrix/fmt/GT>
a vcfc:FormatValueVector ;
vcfc:valueEncoding vcfc:VCFTextVector ;
vcfc:encodedValues "0/1\t0/0\t1/1" .
The three tab-separated items in encodedValues are in sampleIndex order:
sampleIndex |
Sample | GT value in the vector |
|---|---|---|
| 1 | SAMPLE_A |
0/1 |
| 2 | SAMPLE_B |
0/0 |
| 3 | SAMPLE_C |
1/1 |
The DP vector is "42\t18\t9", and the AD vector is
"30,12\t18,0\t9,0". A comma inside one VCF value remains inside that item;
it is not treated as a sample separator. If a value is absent, the vector
uses . to keep every sample position aligned.
The file is marked with
vcfc:representationProfile vcfc:CondensedRepresentation.
- Sample resources are declared once per file, not once per variant.
- Each variant has one
CohortCallMatrix, rather than one sample-call resource for every sample. - Each variant/FORMAT-key pair has one
FormatValueVector, rather than oneFormatFieldValueresource per sample. - The original sample order and all lexical values remain available.
- A vector-aware consumer can reconstruct sample
iby looking up the sample withsampleIndex=iand selecting itemifrom the corresponding vector.
Condensed mode therefore moves repetition from RDF structure into the literal payload of a vector. It is a storage and graph-shape optimization, not a loss of genotype information.
What it does not carry. Condensed mode stops at the vector: it derives no
vcfc:Genotype, vcfc:PhaseSet, vcfc:LocalAlleleSet or
vcfc:BaseModification. Those are one resource per sample per record, which is
precisely the growth this profile exists to prevent. The values are still all
there — inside the vectors — but a SPARQL engine cannot filter on them without
a decoder or an application-level vector function. If you need to query
genotypes directly, that is the argument for expanded mode.
Everything not per-sample is identical in both profiles: the header layer, the allele layer, structured INFO with its value items, the SV carriers, and the record-level ID/FILTER/FORMAT-key decompositions are all emitted the same way.
Condensed mode emits the same file-level block B, then three statements per
record for the vcfc:CohortCallMatrix — hasCallMatrix, rdf:type,
appliesToSampleSet — and five per vector: hasFormatValueVector, rdf:type,
declaredBy, valueEncoding, encodedValues.
T_condensed = B + V × (3 + 5F)
= 3 + 5S + F + V × (3 + 5F)
with every key declared. --sample-data-raw adds one more statement per record.
Two properties follow directly, and they are the point of the profile:
Sappears only in the file-level term. TheSvalues for a variant and key live inside oneencodedValuesliteral, not asSseparate statements.- The count is independent of ploidy, of
GT, and of how many values a FORMAT key holds. The genotype layer and thevcfc:FieldValueItemdecomposition are expanded-only, so aNumber=Rfield costs a condensed graph exactly what a scalar does.
For the worked example (V=1, S=3, FORMAT=GT:DP:AD):
| Condensed item | Count |
|---|---|
Reusable VCFSample resources |
3 |
CohortCallMatrix resources |
1 |
FormatValueVector resources |
3 |
Per-record statements (3 + 5 × 3) |
18 |
File-level SampleSet + VCFSample statements |
17 |
| FORMAT definition statements | 3 |
| Representation-profile statement | 1 |
| Total | 39 |
Because a condensed graph has fixed metadata for the sample set, matrix, and FORMAT definitions, it can be similar in size to—or even slightly larger than— an expanded graph for a tiny file. Its advantage appears when many variants reuse the same sample columns.
Assume a cohort has:
V = 1,000 variants
S = 2,500 samples
F = 3 FORMAT fields
Subtracting the two closed forms, the shared file-level block cancels and the whole difference is per-record:
ΔT = T_expanded − T_condensed = V × [ S × (4 + Σ c_k) − (3 + 5F) ]
and the ratio has a limit that depends only on the cohort width:
T_expanded / T_condensed ──V→∞──▶ S × (4 + Σ c_k) / (3 + 5F) ──F→∞──▶ S
Condensation removes exactly the per-sample dimension, so the saving in
statements approaches the number of samples. Convergence in V is slow while
S is large, because the 5S term dominates T_condensed until V ≫ S.
Expanded totals depend on the FORMAT shape, so three are given. Condensed is
30,506 in every one of them — it does not depend on Σ c_k at all:
| FORMAT shape | per sample call | Expanded | Condensed | Ratio |
|---|---|---|---|---|
3 single-valued typed scalars, no GT |
19 | 47,512,506 | 30,506 | 1,557× |
GT + 2 typed scalars, diploid |
35 | 87,512,506 | 30,506 | 2,869× |
GT:DP:AD, diploid, biallelic |
44 | 110,012,506 | 30,506 | 3,606× |
The limit for the first shape is 2,500 × 19/18 ≈ 2,639×, which the ratio at
V = 1,000 has not yet reached.
This does not mean the compressed file shrinks by the same factor. The same
scalar values still exist inside the vector literals, and their text length,
escaping, dictionary compression, and the rest of the VCF graph affect the
final .nt, .hdt, or .cottas size. The large saving is the removal of
millions of repeated RDF nodes, predicates, and type statements.
Every constant in these formulas is asserted against the emitters' real output
in test/test_representation_growth_unit.py,
so a change to any per-resource triple count fails a test that names the
resource it belongs to.
The modes serve different downstream needs. Combining them into one universal graph would either make small files unnecessarily complicated or make large cohorts unnecessarily expensive.
Expanded mode is a good fit when:
- the VCF has one sample or a small number of samples;
- downstream software already expects
SampleCallandFormatFieldValueresources; - queries need to match one sample value directly, without splitting a vector;
- maximum vocabulary-level transparency is more important than graph size.
For example, a consumer can navigate directly from a call to the sample whose
sampleId is SAMPLE_A, then to its GT FormatFieldValue. Every value is a
normal RDF resource and literal.
Condensed mode is a good fit when:
- there are hundreds or thousands of samples;
- the same sample columns repeat across many variant rows;
- materializing one RDF object per sample/variant/FORMAT combination causes excessive memory, disk, or indexing work;
- downstream software can use
sampleIndexand decode tab-separated vectors.
For a 2,500-sample cohort, expanded mode creates 2,500 sample-call resources for every variant. Condensed mode creates the 2,500 sample descriptions once and reuses them for all variants.
| Question | Expanded | Condensed |
|---|---|---|
| Where is each sample value? | Its own RDF resource | An item in an ordered vector literal |
| How often are samples declared? | Repeated per variant/sample call | Once per file/sample set |
| Simple one-value RDF query | Easier | Requires vector lookup/decoding |
| Large-cohort graph size | Grows with V × S × F |
Structure grows roughly with S + V × F |
| All source values retained? | Yes | Yes |
| Vocabulary profile | ExpandedRepresentation |
CondensedRepresentation |
Neither mode is universally “more correct.” They expose the same underlying VCF information through different graph contracts.
The command-line choice is explicit:
# Expanded is the default
vcf-rdfizer --mode full --input cohort.vcf.gz \
--sample-representation expanded \
--rdf-storage-mode plain --out results
# Condensed is intended for large multi-sample cohorts
vcf-rdfizer --mode full --input cohort.vcf.gz \
--sample-representation condensed \
--rdf-storage-mode space-optimized \
--representations hdt --out resultsThere is no automatic “switch at N samples” threshold. This is deliberate: the same command should produce the same graph contract every time, and a downstream consumer should not unexpectedly receive a different RDF shape just because a file happened to contain more samples.
For the built-in rules, both modes read the wide sample payload in
records.tsv and stream the selected RDF representation directly. The current
implementation does not first materialize the enormous helper
tables for the built-in maps. That implementation optimization is separate
from the semantic choice:
- expanded still emits the expanded
SampleCall/FormatFieldValuegraph; - condensed still emits the condensed
SampleSet/CohortCallMatrix/FormatValueVectorgraph.
Custom rules that consume sample_calls.tsv or
sample_format_values.tsv retain the materialized helper-table behavior. Such
custom helper-table mappings are rejected in condensed mode, because running them
alongside the condensed emitter would produce both graph shapes and recreate
the expansion that condensed mode is intended to avoid.
GraphDB's SPARQL extensions reference contains several different kinds of functionality. They should not all be treated as ways to read a condensed VCF vector.
The standard geof: functions in the GraphDB reference operate on a
geomLiteral, such as a literal with the datatype geo:wktLiteral or
geo:gmlLiteral. They perform geometry operations such as distance, buffer,
intersection, union, envelope, and spatial relationship tests. GraphDB's
additional geoext: functions provide operations such as area, geometry
validity, simplification, and Hausdorff distance. These functions are intended
for points, lines, polygons, and other spatial objects, not arbitrary ordered
text values. See the GeoSPARQL function table
and GraphDB GeoSPARQL extensions.
Our condensed value is instead a normal RDF string associated with
vcfc:VCFTextVector, for example:
"0/1\t0/0\t1/1"
It is not a WKT geometry.
The analogy with WKT remains useful at the design level: both pack a sequence into one literal. The GeoSPARQL operations themselves, however, are not reusable for genotype-vector extraction.
The reference also documents the GraphDB/SPIN magic predicate spif:split.
It takes a string and a regular expression and produces one result row for
each split item. The implementation is described as using Java's
String.split() method. A prototype query for all tokens in condensed vectors
could look like this:
PREFIX vcfc: <https://w3id.org/vcf-core/vocab#>
PREFIX spif: <http://spinrdf.org/spif#>
SELECT ?vector ?token WHERE {
?vector a vcfc:FormatValueVector ;
vcfc:valueEncoding vcfc:VCFTextVector ;
vcfc:encodedValues ?encoded .
?token spif:split (STR(?encoded) "\t") .
}For a GT vector, this would expose 0/1, 0/0, and 1/1 as separate query
bindings. It is useful for experimentation, validation, or exporting values
to an application.
There is an important limitation: splitting gives the values, but the query
also needs the ordinal position of each value in order to identify the sample.
The condensed model says that the first token belongs to sampleIndex=1, the
second to sampleIndex=2, and so on. The documented spif:split pattern does
not itself return that ordinal. GraphDB also documents spif:for, which can
generate a sequence of integers, but the reference does not provide a direct
“zip these generated indexes with the split results” operation. This means
that spif:split alone is not a complete, reliable sample lookup mechanism.
The page also lists RDF-list functions such as list:index and
list:length. Those operate on RDF Collections. vcfc:encodedValues is a
single string literal, not an RDF Collection, so these functions do not apply
unless the data model is changed to materialize every vector item as list
structure. That would reintroduce much of the per-item RDF overhead that
condensed mode was designed to avoid. GraphDB's helper:tuple and
helper:iterate functions similarly operate on internal query-time lists;
they do not automatically parse a persisted VCF vector literal.
| GraphDB feature | Usefulness for VCFTextVector |
Assessment |
|---|---|---|
geof:* GeoSPARQL functions |
Geometry calculations | Not applicable to genotype text vectors |
geoext:* geometry extensions |
Geometry validity and transformations | Not applicable unless a future dataset contains real sample geometries |
spif:split |
Split tab-separated vector text | Useful prototype, but loses the token ordinal |
spif:for |
Generate expected sample positions | Helpful support function, but does not pair positions with split tokens |
list:index / list:length |
Index RDF Collections | Not applicable to the current string encoding |
helper:tuple / helper:iterate |
Work with internal query lists | Not a persisted-vector parser |
A custom vcfc: vector function |
Return value at a sample index | Most direct future SPARQL integration |
| Application-side parsing | Decode one selected vector after retrieval | Best portable near-term approach |
GraphDB explicitly labels these as extensions beyond the W3C SPARQL
specification, so a query using spif:* or a custom function would be tied to
GraphDB (or to another engine that implements the same extension). That can be
acceptable for a GraphDB deployment, but it should not be presented as a
portable SPARQL solution. These are query-time GraphDB operations; using them
would not change VCF-RDFizer's native HDT creation or indexing path, but their
memory and latency behavior would still need to be benchmarked inside the
GraphDB server.
If query-time sample lookup becomes an important use case, the most useful addition would be an extractor that understands the VCF-RDFizer metadata and returns both the position and the value. Conceptually, it could have a function or magic-predicate contract like:
vcfc:vectorValue(?vector, ?sampleIndex) → ?value
or:
?vector vcfc:valueAt (?sampleIndex ?value)
For a sample-aware form, the function could accept the vcfc:VCFSample IRI
instead of an integer, resolve that resource's sampleIndex, and then extract
the matching token. Returning the index as well as the value would make it
possible to validate that a vector has the expected number of positions.
A production implementation should also consider these safeguards:
- Verify that the vector belongs to the expected
SampleSetthrough itsCohortCallMatrix. - Check that the requested index is within the sample-set size.
- Preserve
.as the VCF missing-value marker rather than silently turning it into an unbound result. - Treat tab as the vector separator and keep commas inside a value such as
AD=30,12. - Validate that the vector contains exactly one item per declared sample.
- Expose the function through a GraphDB plugin only when GraphDB-specific deployment is acceptable; otherwise provide a small application-side decoder or a materialized per-value projection.
There is also a performance question. With 1,000 variants, 2,500 samples, and three FORMAT keys, condensed mode stores 3,000 vectors but each vector contains 2,500 values. Expanding every vector in a SPARQL query would produce up to
1,000 × 3 × 2,500 = 7,500,000 query result rows
before filtering to a particular sample. For a targeted lookup, a function that
extracts one position is preferable to spif:split over the entire vector. For
large cohort-wide analyses, an external decoder or a precomputed analytical
table may be more appropriate than forcing a graph database to emit millions
of token bindings.
The GraphDB work is useful as a source of implementation ideas, but not as a drop-in GeoSPARQL solution:
- GeoSPARQL geometry functions: no, they should not be used for genotype vectors.
spif:split: yes, as a prototype tokenizer and validation aid.spif:splitplusspif:for: potentially useful building blocks, but not sufficient for a trustworthy sample/value join without an ordinal pairing mechanism.- Custom vector extraction function or application decoder: yes, this is the most promising future direction.
The condensed model should therefore remain as it is for storage efficiency, while future query support should be added as a separate decoding/projection layer rather than changing genotype vectors into geometries or RDF Lists.
Use this short rule of thumb:
Small or single-sample VCF + simple RDF queries → expanded
Large multi-sample cohort + storage/scale focus → condensed
If a consumer is unsure which graph it received, inspect the file’s
vcfc:representationProfile value before querying. That profile is the
explicit signal that tells the consumer whether sample values are represented
as individual resources or as ordered vectors.
- Variant record: one row describing a genomic position and its alleles.
- Sample: one individual or biological sample represented by a VCF column.
- FORMAT key: a field name such as
GT,DP, orADdescribing one part of a sample’s value. - RDF resource: a named graph object that can have properties, such as one
SampleCallor oneFormatFieldValue. - Vector: an ordered list of values. In condensed mode, values are stored
as one tab-separated
VCFTextVectorliteral insampleIndexorder. - Graph shape: which resources and relationships are present, independent of the underlying biological values.
- Conversion — how these triples are emitted, and why not through RML
- VCF coverage matrix — which genotype elements are validated
- Representations — why condensed matters at cohort scale
- Roadmap — where section 7 leads