Variants¶
Canonical variant ingestion and validation are exported from altar.variants, together with the published
region-annotation contract, REGION_ANNOTATION_CONTRACT, and its portable implementation,
RegionAnnotationSource (see Region annotation source). Names other
than the identity primitives load lazily, so importing the contract loads no dataframe or interval-tree library.
variants
¶
Canonical variant ingestion, validation, annotation, and export primitives.
The altar.variants namespace is Altar's shared scientific boundary. Model
bindings consume these representations rather than defining architecture-specific
variant formats.
The region-annotation contract, REGION_ANNOTATION_CONTRACT, and its portable
implementation, RegionAnnotationSource, are exported here as well.
Names other than the identity primitives resolve lazily (PEP 562): importing the package or the contract loads neither the dataframe stack (pandas, pyfaidx) nor the region lookup stack (intervaltree, ncls). The defining module loads when one of its names is first accessed.
REGION_ANNOTATION_CONTRACT
module-attribute
¶
REGION_ANNOTATION_CONTRACT = AnnotationContract(
source_id=REGION_ANNOTATION_SOURCE_ID,
schema_version="1.0.0",
columns=(
AnnotationColumn(
name="region_type",
dtype="str",
label="Region Type",
),
AnnotationColumn(
name="region_labels",
dtype="str",
label="Region Labels",
repeated=True,
),
AnnotationColumn(
name="region_gene_ids",
dtype="str",
label="Region Gene IDs",
repeated=True,
),
AnnotationColumn(
name="is_coding", dtype="bool", label="Coding"
),
AnnotationColumn(
name="in_promoter",
dtype="bool",
label="In Promoter",
),
AnnotationColumn(
name="nearest_genes",
dtype="struct",
label="Nearest Genes",
repeated=True,
fields=(
AnnotationColumn(
name="gene_name",
dtype="str",
label="Gene Name",
),
AnnotationColumn(
name="gene_id",
dtype="str",
label="Gene ID",
),
AnnotationColumn(
name="distance",
dtype="int",
label="Distance",
),
AnnotationColumn(
name="type",
dtype="str",
label="Gene Type",
),
),
),
AnnotationColumn(
name="gene_within_100kb",
dtype="bool",
label="Gene Within 100 kb",
),
AnnotationColumn(
name="ccre_id", dtype="str", label="cCRE ID"
),
AnnotationColumn(
name="ccre_group",
dtype="str",
label="cCRE Group",
),
),
)
The region-annotation columns, version 1.0.0. ChromBPNet's prioritization depends on region_type.
VariantIdentityError
¶
Bases: ValueError
A variant key is malformed, non-canonical, or disagrees with its locus.
VariantKey
dataclass
¶
Canonical one-based biallelic key serialized as chr:pos:ref:alt.
Fields follow the rules in this module's documentation. The class normalizes spelling but does not normalize biological representation. Callers must left-align and trim indels against the selected reference before constructing a key when equivalence across representations matters.
parse
classmethod
¶
parse(variant_id: str) -> VariantKey
Parse and normalize a four-field textual key.
require_canonical
classmethod
¶
require_canonical(variant_id: str) -> VariantKey
Parse variant_id and reject aliases instead of silently rewriting them.
from_fields
classmethod
¶
from_fields(
chromosome: str,
position: int,
reference_allele: str,
alternate_allele: str,
) -> VariantKey
Construct a canonical key from decomposed one-based locus fields.
RegionAnnotationSource
¶
RegionAnnotationSource(
data_dir: str | PathLike[str] | None = None,
genome_build: str = REGION_GENOME_BUILD,
)
Bases: AnnotationSource
Compute the org.kundajelab.altar.annotation.regions columns for each variant.
data_dir is the annotation data directory, resolved once at construction: the argument, else
$ALTAR_VARIANTS_DATA_DIR, else the installed package's data directory. It must hold
region_annotations.parquet, ccres.dnatree, and region_provenance.json; a directory without its
own gene_df.tsv uses the shipped one. The data is GRCh38, so genome_build must be "hg38".
identity derives the release from the inputs and logic version region_provenance.json records,
after checking each file's SHA-256 against it, and annotate runs that check before it loads anything.
The cCRE tree is unpickled only from bytes with the recorded digest. These checks catch drift and accidents;
they do not protect against someone who can write the data directory. A variant on a chromosome without
protein-coding genes in the gene table, such as an unplaced contig, has no row.
identity
¶
identity() -> AnnotationSourceIdentity
Return the contract identity, with the release named by the verified provenance manifest.
ValidationErrorReason
¶
Bases: Enum
Reasons why a variant validation might fail.
Generic validation never emits ALT_LENGTH_MISMATCH or ALT_MISMATCH. They described a model input
window that no longer belongs to core validation and remain defined so persisted reason values keep parsing.
ValidationResult
dataclass
¶
ValidationResult(
is_valid: bool,
chro: Optional[str] = None,
error_reason: Optional[ValidationErrorReason] = None,
error_message: Optional[str] = None,
)
Result of validating a single variant.
canonical_chromosome
¶
Return the canonical chr-prefixed spelling used in portable keys.
Primary chromosomes accept common aliases (1/chr1/CHR01 and
M/MT/chrM) in any case. Other reference contigs keep their exact
spelling after an optional case-insensitive chr prefix is normalized, so
CHRUn_KI270302v1 becomes chrUn_KI270302v1 but chrun_KI270302v1
stays as written. Non-ASCII text, whitespace inside the name, and characters
outside the VCF contig-name set (including :) raise
VariantIdentityError. This is a naming rule only; it does not claim that
contigs from different assemblies are interchangeable.
canonical_variant_id
¶
canonical_variant_id(
chromosome: str,
position: int,
reference_allele: str,
alternate_allele: str,
) -> str
Return the canonical text key for decomposed one-based locus fields.
load_variants
¶
Load a headerless TSV and derive canonical score keys from its locus fields.
An optional fifth input column is a source identifier. The occurrence-aware
streaming pipeline preserves it as source_variant_id; this compatibility
DataFrame returns the canonical five-column scoring schema instead.
load_variants_vcf
¶
Parse VCF ALT occurrences and derive canonical locus-based score keys.
read_variants_frame
¶
Load variants into the canonical table, dispatching on file format.
This is the format-neutral entry point; load_variants (TSV) and
load_variants_vcf (VCF/gVCF) are the per-format adapters it composes.
It loads the whole file into memory; altar_identity.read_variants, by
contrast, streams VariantKey values from Altar's container variant file.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
path
|
str
|
Input variants file. |
required |
fmt
|
str
|
One of |
'auto'
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
DataFrame with columns |
validate_variant
¶
validate_variant(
chro: str, pos: int, ref: str, alt: str, genome: Fasta
) -> ValidationResult
Check one variant against the reference genome, independent of any model.
The chromosome must be a primary chromosome (1-22, X, Y or M, under any alias that
canonical_chromosome accepts) and present in the FASTA under its canonical name. REF and
ALT must be non-empty uppercase ACGT and differ, REF must lie inside the contig, and REF must equal the
reference bases at pos. Soft-masked (lowercase) reference bases match. Whether a model's
input window fits around the variant is that model's concern, not part of this check.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
chro
|
str
|
Chromosome |
required |
pos
|
int
|
Position (0-based) |
required |
ref
|
str
|
Reference allele |
required |
alt
|
str
|
Alternate allele |
required |
genome
|
Fasta
|
Open genome file handle |
required |
Returns:
| Type | Description |
|---|---|
ValidationResult
|
ValidationResult with the canonical chromosome, and error details if invalid |
validate_variants
¶
Validate variants with validate_variant and return valid and invalid dataframes.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
variants_df
|
DataFrame
|
DataFrame with variant information |
required |
genome
|
Fasta
|
Open genome file handle |
required |
Returns:
| Type | Description |
|---|---|
DataFrame
|
Tuple of (valid_variants_df, invalid_variants_df) |
DataFrame
|
Invalid variants have 'error_reason' and 'error_message' columns added. |