Region Classification¶
Altar classifies each variant into genomic region categories (CDS, UTR, splice site/region, intron, promoter, non-coding-gene body, intergenic) by pure interval overlap against an Ensembl gene model. The labels are used to route variants to the right scorer — coding variants to AlphaMissense, splice-region variants to SpliceAI — and for reporting.
Design¶
- Interval overlap only — no FASTA, no codon math. The hard part of variant consequence (codon-level missense/stop calling, splice-disruption prediction) is already owned by the dedicated scorers (AlphaMissense, SpliceAI). The region classifier only needs to categorize and route, and every category here is derivable from interval overlap against a rich enough annotation.
- Multi-label. A variant can legitimately be several things at once (a coding base that is also a canonical splice site; exonic in one transcript and intronic in another). The classifier returns the full set of labels, not a single collapsed string.
- Strand-aware. Donor vs acceptor splice sites and 5′ vs 3′ UTR are resolved using transcript strand.
Setup¶
Region classification reads region_annotations.parquet, which the altar wheel does not include. The
variants runtime image includes it. Outside the image, build it into your annotation data directory. See
Annotation data files for where Altar looks for it.
-
Download the Ensembl gene-model GFF3 (default release 116, GRCh38):
-
Parse it into a flat interval table (parquet). The output defaults to
$ALTAR_VARIANTS_DATA_DIR/region_annotations.parquet:
The output is a flat interval list — one row per labelled genomic interval
(chrom, start, end, strand, feature, gene_id, gene_name, biotype), 1-based and
endpoint-inclusive. exon, cds, five_prime_utr, and three_prime_utr come
straight from the GFF3; intron, splice_donor, splice_acceptor,
splice_region, and promoter are derived from per-transcript exon
structure so that lookup is a single overlap join. The table (~32M intervals) is
gitignored.
Window definitions are constants in altar.variants.scripts.construct_region_annotations:
| Feature | Window |
|---|---|
splice_donor / splice_acceptor |
2 bp intronic (canonical, VEP definition) |
splice_region |
8 bp intronic + 3 bp exonic around each exon boundary |
promoter |
500 bp upstream / 100 bp downstream of the TSS |
Labels¶
Returned labels, ordered most → least severe (the order used for the single-label collapse):
splice_site, cds, splice_region, five_prime_utr, three_prime_utr,
exonic, noncoding_gene, intronic, promoter, intergenic
cdsdrives coding routing — a variant is coding iff it overlaps a CDS (RegionAnnotation.is_coding), and those go to AlphaMissense.- An
exonin a protein-coding gene →exonic; anexonin a non-coding gene (lncRNA, miRNA, …) →noncoding_gene, so non-coding loci are no longer mislabelled asintergenic. intergenicis the empty case (no overlap) and is excluded from both the coding and non-coding scorer outputs.
Python API¶
Experimental implementation API
Region-classification callables are not yet part of the stable façade. The deep import below may change;
only names re-exported directly by altar.variants carry the compatibility commitment today.
from altar.variants.annotation.regions import (
region_annotations, # single interval -> RegionAnnotation
region_annotations_batch, # vectorized over many intervals (preferred at scale)
region_type, # back-compat shim -> single most-severe label (str)
)
ann = region_annotations("chr1", 69094, 69094)
ann.labels # e.g. ["cds", "exonic"] (severity-ordered)
ann.gene_ids # overlapping gene IDs (multi-gene aware)
ann.primary # "cds" -> most-severe single label
ann.is_coding # True -> overlaps a CDS (routes to AlphaMissense)
ann.in_promoter # False -> overlaps a promoter window (used by prioritization)
# Annotate a whole variant set in one vectorized pass:
anns = region_annotations_batch(df["chr"], df["pos"], df["pos"])
region_type() is retained only for back-compatibility (returns .primary);
new code should use region_annotations[_batch] to get the full label set.
Routing variants to scorers¶
altar.variants.preprocessing.region_filter uses these
labels to split a variant file into per-category TSVs, one per region kind, so
each can feed the right scorer:
Routing is membership-based and overlapping — a variant is written to every
category whose labels it carries, so it can appear in several files at once (a
CDS-edge SNV lands in coding, exonic, splice, splice_site,
splice_region, and genic). Each file is a bare, headerless variant TSV,
i.e. a drop-in input to the per-scorer entrypoints.
Categories written (all except intergenic by default; restrict with
--categories):
| File | Membership | Typical scorer |
|---|---|---|
coding.tsv |
overlaps a CDS | AlphaMissense |
exonic.tsv |
protein-coding exon (cds ∪ UTRs) | — |
five_prime_utr.tsv / three_prime_utr.tsv |
UTR | — |
splice.tsv |
splice site or region | SpliceAI |
splice_site.tsv |
canonical ±2 bp donor/acceptor | SpliceAI |
splice_region.tsv |
±8 intron / ±3 exon window | SpliceAI |
intronic.tsv |
intron body | chromatin models |
promoter.tsv |
promoter window | chromatin models |
noncoding_gene.tsv |
lncRNA/miRNA/etc exon | — |
genic.tsv |
any gene-body label (excludes promoter, intergenic) | SpliceAI (broad) |
intergenic.tsv |
no annotated overlap (opt-in) | — |
There is intentionally no noncoding category: "not coding" is ambiguous
under multi-transcript union (a variant can be CDS in one transcript and
intronic/UTR in another), and it is the wrong organizing principle for chromatin models like
ChromBPNet, which score any locus — coding included. Feed those models the
full variant set, not a region subset; use the categories above only to
route the scorers that need it (CDS → AlphaMissense, splice → SpliceAI).
Implementation notes¶
- Index: a per-chromosome NCLS
(C-backed nested containment list) queried with a vectorized
all_overlaps_bothjoin — annotate the whole DataFrame at once instead of aniterrowsloop. NCLS is the core thatpyrangeswraps; it's used directly here becausepyranges1.x needs Python ≥3.10 (this project is on 3.9 + pandas 2). - Storage: parquet, not a pickled tree — the NCLS index rebuilds cheaply from sorted arrays on load, and parquet is inspectable, compressed, and consistent with the AlphaMissense/SpliceAI datasets.
- cCRE / regulatory: still annotated on the existing
DNATreepath (ccre_overlap); folding the Ensembl Regulatory Build into aregulatorylabel is a planned follow-up.