Skip to content

Region Classification

Altar classifies each variant into genomic region categories (CDS, UTR, splice site/region, intron, promoter, non-coding-gene body, intergenic) by pure interval overlap against an Ensembl gene model. The labels are used to route variants to the right scorer — coding variants to AlphaMissense, splice-region variants to SpliceAI — and for reporting.

Design

  • Interval overlap only — no FASTA, no codon math. The hard part of variant consequence (codon-level missense/stop calling, splice-disruption prediction) is already owned by the dedicated scorers (AlphaMissense, SpliceAI). The region classifier only needs to categorize and route, and every category here is derivable from interval overlap against a rich enough annotation.
  • Multi-label. A variant can legitimately be several things at once (a coding base that is also a canonical splice site; exonic in one transcript and intronic in another). The classifier returns the full set of labels, not a single collapsed string.
  • Strand-aware. Donor vs acceptor splice sites and 5′ vs 3′ UTR are resolved using transcript strand.

Setup

Region classification reads region_annotations.parquet, which the altar wheel does not include. The variants runtime image includes it. Outside the image, build it into your annotation data directory. See Annotation data files for where Altar looks for it.

  1. Download the Ensembl gene-model GFF3 (default release 116, GRCh38):

    export ALTAR_VARIANTS_DATA_DIR="$HOME/altar-variants-data"
    bash altar/altar/variants/scripts/download_ensembl_gff.sh  # ENSEMBL_RELEASE=116 by default
    

  2. Parse it into a flat interval table (parquet). The output defaults to $ALTAR_VARIANTS_DATA_DIR/region_annotations.parquet:

    python -m altar.variants.scripts.construct_region_annotations \
        -i "$ALTAR_VARIANTS_DATA_DIR/raw/Homo_sapiens.GRCh38.116.gff3.gz"
    

The output is a flat interval list — one row per labelled genomic interval (chrom, start, end, strand, feature, gene_id, gene_name, biotype), 1-based and endpoint-inclusive. exon, cds, five_prime_utr, and three_prime_utr come straight from the GFF3; intron, splice_donor, splice_acceptor, splice_region, and promoter are derived from per-transcript exon structure so that lookup is a single overlap join. The table (~32M intervals) is gitignored.

Window definitions are constants in altar.variants.scripts.construct_region_annotations:

Feature Window
splice_donor / splice_acceptor 2 bp intronic (canonical, VEP definition)
splice_region 8 bp intronic + 3 bp exonic around each exon boundary
promoter 500 bp upstream / 100 bp downstream of the TSS

Labels

Returned labels, ordered most → least severe (the order used for the single-label collapse):

splice_site, cds, splice_region, five_prime_utr, three_prime_utr,
exonic, noncoding_gene, intronic, promoter, intergenic
  • cds drives coding routing — a variant is coding iff it overlaps a CDS (RegionAnnotation.is_coding), and those go to AlphaMissense.
  • An exon in a protein-coding gene → exonic; an exon in a non-coding gene (lncRNA, miRNA, …) → noncoding_gene, so non-coding loci are no longer mislabelled as intergenic.
  • intergenic is the empty case (no overlap) and is excluded from both the coding and non-coding scorer outputs.

Python API

Experimental implementation API

Region-classification callables are not yet part of the stable façade. The deep import below may change; only names re-exported directly by altar.variants carry the compatibility commitment today.

from altar.variants.annotation.regions import (
    region_annotations,  # single interval -> RegionAnnotation
    region_annotations_batch,  # vectorized over many intervals (preferred at scale)
    region_type,  # back-compat shim -> single most-severe label (str)
)

ann = region_annotations("chr1", 69094, 69094)
ann.labels  # e.g. ["cds", "exonic"]  (severity-ordered)
ann.gene_ids  # overlapping gene IDs (multi-gene aware)
ann.primary  # "cds"  -> most-severe single label
ann.is_coding  # True   -> overlaps a CDS (routes to AlphaMissense)
ann.in_promoter  # False  -> overlaps a promoter window (used by prioritization)

# Annotate a whole variant set in one vectorized pass:
anns = region_annotations_batch(df["chr"], df["pos"], df["pos"])

region_type() is retained only for back-compatibility (returns .primary); new code should use region_annotations[_batch] to get the full label set.

Routing variants to scorers

altar.variants.preprocessing.region_filter uses these labels to split a variant file into per-category TSVs, one per region kind, so each can feed the right scorer:

uv run python -m altar.variants.preprocessing.region_filter -v variants.tsv -o regions/

Routing is membership-based and overlapping — a variant is written to every category whose labels it carries, so it can appear in several files at once (a CDS-edge SNV lands in coding, exonic, splice, splice_site, splice_region, and genic). Each file is a bare, headerless variant TSV, i.e. a drop-in input to the per-scorer entrypoints.

Categories written (all except intergenic by default; restrict with --categories):

File Membership Typical scorer
coding.tsv overlaps a CDS AlphaMissense
exonic.tsv protein-coding exon (cds ∪ UTRs) —
five_prime_utr.tsv / three_prime_utr.tsv UTR —
splice.tsv splice site or region SpliceAI
splice_site.tsv canonical ±2 bp donor/acceptor SpliceAI
splice_region.tsv ±8 intron / ±3 exon window SpliceAI
intronic.tsv intron body chromatin models
promoter.tsv promoter window chromatin models
noncoding_gene.tsv lncRNA/miRNA/etc exon —
genic.tsv any gene-body label (excludes promoter, intergenic) SpliceAI (broad)
intergenic.tsv no annotated overlap (opt-in) —

There is intentionally no noncoding category: "not coding" is ambiguous under multi-transcript union (a variant can be CDS in one transcript and intronic/UTR in another), and it is the wrong organizing principle for chromatin models like ChromBPNet, which score any locus — coding included. Feed those models the full variant set, not a region subset; use the categories above only to route the scorers that need it (CDS → AlphaMissense, splice → SpliceAI).

Implementation notes

  • Index: a per-chromosome NCLS (C-backed nested containment list) queried with a vectorized all_overlaps_both join — annotate the whole DataFrame at once instead of an iterrows loop. NCLS is the core that pyranges wraps; it's used directly here because pyranges 1.x needs Python ≥3.10 (this project is on 3.9 + pandas 2).
  • Storage: parquet, not a pickled tree — the NCLS index rebuilds cheaply from sorted arrays on load, and parquet is inspectable, compressed, and consistent with the AlphaMissense/SpliceAI datasets.
  • cCRE / regulatory: still annotated on the existing DNATree path (ccre_overlap); folding the Ensembl Regulatory Build into a regulatory label is a planned follow-up.