Variant identity¶
Altar uses one portable join key for scores, named details, and annotations:
For example, 1:100:a:g normalizes to chr1:100:A:G. Every field is ASCII; surrounding whitespace is
trimmed, and any other whitespace or non-ASCII text is rejected rather than folded.
- Chromosome: the
chrprefix is optional and case-insensitive. Primary chromosomes are normalized in any case, includingchr01tochr1andMTtochrM. Other contigs keep their exact spelling after the prefix (CHRUn_KI270302v1becomeschrUn_KI270302v1), because reference contig names are case-sensitive and the key must name the sequence in the FASTA;chrun_KI270302v1is a different key. Names use the VCF contig-name characters, which exclude:. - Position: a positive integer. In text (a TSV cell or a VCF
POS) it is ASCII digits only, so+5,1_000,1e3, and non-ASCII digits are rejected. - Alleles: uppercased, then one or more of
A,C,G,T,N, the VCF base alphabet. Symbolic (<DEL>), spanning-deletion (*), missing (.), breakend, IUPAC-ambiguity, and-alleles have no key, and REF must differ from ALT. A key may containN; reference validation decides whether the variant can be scored.
Input files treat a row without a key differently by format. VCF loading drops and counts such ALT alleles,
because symbolic and spanning-deletion ALTs are ordinary VCF records. A variant TSV has no such records, so
read_variants_frame and load_variants treat a TSV row without a key as an input error: they reject the file
and name the offending 1-based row. A position that is not ASCII digits is an input error in either format.
Use VariantKey to parse or validate keys and canonical_variant_id() to construct one from locus fields.
Both are available from altar.models, altar.sources, and altar.variants, and are defined in
altar-identity, which model runtimes import so they build the same keys. New
scoring and materialization boundaries reject opaque or non-canonical keys, and result rows that carry both
variant_id and decomposed locus fields must agree.
Source identifiers¶
The optional fifth TSV column and the VCF ID field are source metadata, not join keys. Streaming
preprocessing records them as source_variant_id so values such as rs123, cohort labels, and missing IDs
remain available for occurrence-aware export. Model runtimes always derive their output variant_id from
chr, pos, ref, and alt. The public altar.variants DataFrame loaders do the same; use streaming
occurrence outputs when source-ID provenance is required.
Genome builds and indels¶
Genome build is deliberately not embedded in the text key. It travels separately in source queries and model-run provenance, so identical-looking coordinates from different assemblies are not treated as proven equivalents.
Canonical spelling also does not left-align or trim indels. Normalize variants against the selected reference before key construction when equivalent indel representations must join.
See Prepare VCF input for the occurrence and source-identifier outputs produced by streaming preprocessing.