Skip to content

Variants

Canonical variant ingestion and validation are exported from altar.variants, together with the published region-annotation contract, REGION_ANNOTATION_CONTRACT, and its portable implementation, RegionAnnotationSource (see Region annotation source). Names other than the identity primitives load lazily, so importing the contract loads no dataframe or interval-tree library.

variants

Canonical variant ingestion, validation, annotation, and export primitives.

The altar.variants namespace is Altar's shared scientific boundary. Model bindings consume these representations rather than defining architecture-specific variant formats.

The region-annotation contract, REGION_ANNOTATION_CONTRACT, and its portable implementation, RegionAnnotationSource, are exported here as well.

Names other than the identity primitives resolve lazily (PEP 562): importing the package or the contract loads neither the dataframe stack (pandas, pyfaidx) nor the region lookup stack (intervaltree, ncls). The defining module loads when one of its names is first accessed.

REGION_ANNOTATION_CONTRACT module-attribute

REGION_ANNOTATION_CONTRACT = AnnotationContract(
    source_id=REGION_ANNOTATION_SOURCE_ID,
    schema_version="1.0.0",
    columns=(
        AnnotationColumn(
            name="region_type",
            dtype="str",
            label="Region Type",
        ),
        AnnotationColumn(
            name="region_labels",
            dtype="str",
            label="Region Labels",
            repeated=True,
        ),
        AnnotationColumn(
            name="region_gene_ids",
            dtype="str",
            label="Region Gene IDs",
            repeated=True,
        ),
        AnnotationColumn(
            name="is_coding", dtype="bool", label="Coding"
        ),
        AnnotationColumn(
            name="in_promoter",
            dtype="bool",
            label="In Promoter",
        ),
        AnnotationColumn(
            name="nearest_genes",
            dtype="struct",
            label="Nearest Genes",
            repeated=True,
            fields=(
                AnnotationColumn(
                    name="gene_name",
                    dtype="str",
                    label="Gene Name",
                ),
                AnnotationColumn(
                    name="gene_id",
                    dtype="str",
                    label="Gene ID",
                ),
                AnnotationColumn(
                    name="distance",
                    dtype="int",
                    label="Distance",
                ),
                AnnotationColumn(
                    name="type",
                    dtype="str",
                    label="Gene Type",
                ),
            ),
        ),
        AnnotationColumn(
            name="gene_within_100kb",
            dtype="bool",
            label="Gene Within 100 kb",
        ),
        AnnotationColumn(
            name="ccre_id", dtype="str", label="cCRE ID"
        ),
        AnnotationColumn(
            name="ccre_group",
            dtype="str",
            label="cCRE Group",
        ),
    ),
)

The region-annotation columns, version 1.0.0. ChromBPNet's prioritization depends on region_type.

VariantIdentityError

Bases: ValueError

A variant key is malformed, non-canonical, or disagrees with its locus.

VariantKey dataclass

VariantKey(
    chromosome: str,
    position: int,
    reference_allele: str,
    alternate_allele: str,
)

Canonical one-based biallelic key serialized as chr:pos:ref:alt.

Fields follow the rules in this module's documentation. The class normalizes spelling but does not normalize biological representation. Callers must left-align and trim indels against the selected reference before constructing a key when equivalence across representations matters.

variant_id property

variant_id: str

Return the portable score, detail, and annotation join key.

parse classmethod

parse(variant_id: str) -> VariantKey

Parse and normalize a four-field textual key.

require_canonical classmethod

require_canonical(variant_id: str) -> VariantKey

Parse variant_id and reject aliases instead of silently rewriting them.

from_fields classmethod

from_fields(
    chromosome: str,
    position: int,
    reference_allele: str,
    alternate_allele: str,
) -> VariantKey

Construct a canonical key from decomposed one-based locus fields.

RegionAnnotationSource

RegionAnnotationSource(
    data_dir: str | PathLike[str] | None = None,
    genome_build: str = REGION_GENOME_BUILD,
)

Bases: AnnotationSource

Compute the org.kundajelab.altar.annotation.regions columns for each variant.

data_dir is the annotation data directory, resolved once at construction: the argument, else $ALTAR_VARIANTS_DATA_DIR, else the installed package's data directory. It must hold region_annotations.parquet, ccres.dnatree, and region_provenance.json; a directory without its own gene_df.tsv uses the shipped one. The data is GRCh38, so genome_build must be "hg38".

identity derives the release from the inputs and logic version region_provenance.json records, after checking each file's SHA-256 against it, and annotate runs that check before it loads anything. The cCRE tree is unpickled only from bytes with the recorded digest. These checks catch drift and accidents; they do not protect against someone who can write the data directory. A variant on a chromosome without protein-coding genes in the gene table, such as an unplaced contig, has no row.

identity

Return the contract identity, with the release named by the verified provenance manifest.

ValidationErrorReason

Bases: Enum

Reasons why a variant validation might fail.

Generic validation never emits ALT_LENGTH_MISMATCH or ALT_MISMATCH. They described a model input window that no longer belongs to core validation and remain defined so persisted reason values keep parsing.

ValidationResult dataclass

ValidationResult(
    is_valid: bool,
    chro: Optional[str] = None,
    error_reason: Optional[ValidationErrorReason] = None,
    error_message: Optional[str] = None,
)

Result of validating a single variant.

canonical_chromosome

canonical_chromosome(chromosome: str) -> str

Return the canonical chr-prefixed spelling used in portable keys.

Primary chromosomes accept common aliases (1/chr1/CHR01 and M/MT/chrM) in any case. Other reference contigs keep their exact spelling after an optional case-insensitive chr prefix is normalized, so CHRUn_KI270302v1 becomes chrUn_KI270302v1 but chrun_KI270302v1 stays as written. Non-ASCII text, whitespace inside the name, and characters outside the VCF contig-name set (including :) raise VariantIdentityError. This is a naming rule only; it does not claim that contigs from different assemblies are interchangeable.

canonical_variant_id

canonical_variant_id(
    chromosome: str,
    position: int,
    reference_allele: str,
    alternate_allele: str,
) -> str

Return the canonical text key for decomposed one-based locus fields.

load_variants

load_variants(variants_loc: str) -> DataFrame

Load a headerless TSV and derive canonical score keys from its locus fields.

An optional fifth input column is a source identifier. The occurrence-aware streaming pipeline preserves it as source_variant_id; this compatibility DataFrame returns the canonical five-column scoring schema instead.

load_variants_vcf

load_variants_vcf(path: str) -> DataFrame

Parse VCF ALT occurrences and derive canonical locus-based score keys.

read_variants_frame

read_variants_frame(
    path: str, fmt: str = "auto"
) -> DataFrame

Load variants into the canonical table, dispatching on file format.

This is the format-neutral entry point; load_variants (TSV) and load_variants_vcf (VCF/gVCF) are the per-format adapters it composes. It loads the whole file into memory; altar_identity.read_variants, by contrast, streams VariantKey values from Altar's container variant file.

Parameters:

Name Type Description Default
path str

Input variants file.

required
fmt str

One of "auto", "tsv", "vcf". "auto" routes .vcf/.vcf.gz (and .gvcf) to the VCF parser and everything else to load_variants; .bcf raises a clear error.

'auto'

Returns:

Type Description
DataFrame

DataFrame with columns chr, pos, ref, alt, variant_id.

validate_variant

validate_variant(
    chro: str, pos: int, ref: str, alt: str, genome: Fasta
) -> ValidationResult

Check one variant against the reference genome, independent of any model.

The chromosome must be a primary chromosome (1-22, X, Y or M, under any alias that canonical_chromosome accepts) and present in the FASTA under its canonical name. REF and ALT must be non-empty uppercase ACGT and differ, REF must lie inside the contig, and REF must equal the reference bases at pos. Soft-masked (lowercase) reference bases match. Whether a model's input window fits around the variant is that model's concern, not part of this check.

Parameters:

Name Type Description Default
chro str

Chromosome

required
pos int

Position (0-based)

required
ref str

Reference allele

required
alt str

Alternate allele

required
genome Fasta

Open genome file handle

required

Returns:

Type Description
ValidationResult

ValidationResult with the canonical chromosome, and error details if invalid

validate_variants

validate_variants(
    variants_df: DataFrame, genome: Fasta
) -> Tuple[DataFrame, DataFrame]

Validate variants with validate_variant and return valid and invalid dataframes.

Parameters:

Name Type Description Default
variants_df DataFrame

DataFrame with variant information

required
genome Fasta

Open genome file handle

required

Returns:

Type Description
DataFrame

Tuple of (valid_variants_df, invalid_variants_df)

DataFrame

Invalid variants have 'error_reason' and 'error_message' columns added.