Skip to content

Variant identity

Altar uses one portable join key for scores, named details, and annotations:

chr:positive-one-based-position:UPPERCASE-REF:UPPERCASE-ALT

For example, 1:100:a:g normalizes to chr1:100:A:G. Every field is ASCII; surrounding whitespace is trimmed, and any other whitespace or non-ASCII text is rejected rather than folded.

  • Chromosome: the chr prefix is optional and case-insensitive. Primary chromosomes are normalized in any case, including chr01 to chr1 and MT to chrM. Other contigs keep their exact spelling after the prefix (CHRUn_KI270302v1 becomes chrUn_KI270302v1), because reference contig names are case-sensitive and the key must name the sequence in the FASTA; chrun_KI270302v1 is a different key. Names use the VCF contig-name characters, which exclude :.
  • Position: a positive integer. In text (a TSV cell or a VCF POS) it is ASCII digits only, so +5, 1_000, 1e3, and non-ASCII digits are rejected.
  • Alleles: uppercased, then one or more of A, C, G, T, N, the VCF base alphabet. Symbolic (<DEL>), spanning-deletion (*), missing (.), breakend, IUPAC-ambiguity, and - alleles have no key, and REF must differ from ALT. A key may contain N; reference validation decides whether the variant can be scored.

Input files treat a row without a key differently by format. VCF loading drops and counts such ALT alleles, because symbolic and spanning-deletion ALTs are ordinary VCF records. A variant TSV has no such records, so read_variants_frame and load_variants treat a TSV row without a key as an input error: they reject the file and name the offending 1-based row. A position that is not ASCII digits is an input error in either format.

Use VariantKey to parse or validate keys and canonical_variant_id() to construct one from locus fields. Both are available from altar.models, altar.sources, and altar.variants, and are defined in altar-identity, which model runtimes import so they build the same keys. New scoring and materialization boundaries reject opaque or non-canonical keys, and result rows that carry both variant_id and decomposed locus fields must agree.

Source identifiers

The optional fifth TSV column and the VCF ID field are source metadata, not join keys. Streaming preprocessing records them as source_variant_id so values such as rs123, cohort labels, and missing IDs remain available for occurrence-aware export. Model runtimes always derive their output variant_id from chr, pos, ref, and alt. The public altar.variants DataFrame loaders do the same; use streaming occurrence outputs when source-ID provenance is required.

Genome builds and indels

Genome build is deliberately not embedded in the text key. It travels separately in source queries and model-run provenance, so identical-looking coordinates from different assemblies are not treated as proven equivalents.

Canonical spelling also does not left-align or trim indels. Normalize variants against the selected reference before key construction when equivalent indel representations must join.

See Prepare VCF input for the occurrence and source-identifier outputs produced by streaming preprocessing.