Skip to content

Region annotation source

Altar publishes the region annotations (region class, overlapping genes, nearest protein-coding genes, and the overlapping SCREEN cCRE) as an annotation contract, org.kundajelab.altar.annotation.regions. Any source that serves these columns with this contract's identity can feed materialization and prioritization: the portable RegionAnnotationSource, or a warehouse table loaded from the annotate command's output. ChromBPNet's prioritization policy depends on this contract's region_type column.

from altar.variants import REGION_ANNOTATION_CONTRACT, RegionAnnotationSource

Importing the contract loads no lookup data, and neither pandas nor the interval-tree libraries.

The contract

REGION_ANNOTATION_CONTRACT is version 1.0.0. Its columns have the same names and logical types as the fields altar.variants.annotation.annotate writes to each JSONL row, so those rows load directly into a table declared with these columns.

Column Logical type Meaning
region_type str The most severe region label, such as cds, splice_site, promoter, or intergenic
region_labels repeated str Every overlapping region label, most severe first
region_gene_ids repeated str Ensembl IDs of the genes whose features the variant overlaps
is_coding bool The variant overlaps a CDS
in_promoter bool The variant overlaps a promoter window of any gene
nearest_genes repeated struct The five nearest protein-coding genes: gene_name (str), gene_id (str), distance (int, 0 inside the gene), type (str)
gene_within_100kb bool The nearest protein-coding gene lies within 100 kb
ccre_id str, null when none Accession of an overlapping SCREEN cCRE
ccre_group str, null when none That cCRE's group, such as PLS or dELS

Region classification defines the labels and windows.

The columns come from two gene annotations. region_type, region_labels, region_gene_ids, is_coding, and in_promoter come from the Ensembl GFF3 the region index was built from (Ensembl 116 by default). nearest_genes and gene_within_100kb come from the shipped gene table, GENCODE 47 (Ensembl 113). A gene ID can therefore appear in one and not the other, and the two can name a gene differently.

Each JSONL row also carries fields outside the contract:

  • ref_length, alt_length, variant_length, and variant_type derive from the alleles alone. They have no data release, and any consumer can recompute them from the variant ID.
  • in_ot and the af_* population frequencies come from Open Targets' gnomAD-derived variant table. That is a different dataset with its own release, so it belongs to its own contract.

A test checks the contract against the annotate command's row model, so a field added to the rows must be either added to the contract in a new contract version or listed as outside it.

Identity and data release

A source implementing the contract reports an identity: the contract's source_id and schema_version, the data release, and the genome_build. The region data is GRCh38, so the build is always hg38.

The release follows the science, not the files. It names the three inputs the values come from and the version of the logic that turns them into columns, which the data directory's provenance manifest, region_provenance.json, records (see Annotation data files):

ensembl-116.ccre-v4.genes-gencode-47.logic-1
  • ensembl-116: the Ensembl GFF3 that region_annotations.parquet was built from.
  • ccre-v4: the SCREEN cCRE Registry BED that ccres.dnatree was built from.
  • genes-gencode-47: the nearest-gene table.
  • logic-1: REGION_LOGIC_VERSION, which covers the build scripts' windows and the lookups' semantics. A change to either that changes any value bumps it, and a golden-output test fails until it does.

Altar pins the SHA-256 of the Ensembl 116 GFF3, the SCREEN v4 BED, and the GENCODE 47 gene table. When every input has its pinned digest, the string above is the whole release. Otherwise, for example for another Ensembl release or your own gene table, . and the first 12 hex characters of a digest over the three input digests follow.

The derived files' bytes are not part of the release. A parquet written by another pandas or pyarrow version names the same release, while different inputs or different logic name a different one. The manifest still records each derived file's SHA-256, and annotation checks the files against it before it reads them.

Portable path: RegionAnnotationSource

RegionAnnotationSource(data_dir=None, genome_build="hg38") computes the contract columns in Python with the same function the annotate command uses for its rows. Use it for tests, notebooks, and small batches.

from altar.sources import fetch_annotations
from altar.variants import RegionAnnotationSource

source = RegionAnnotationSource(data_dir="/data/altar-variants")
source.identity()  # the contract identity, with the release from region_provenance.json
annotations = await fetch_annotations([source], ["chr1:1014143:C:T"])

data_dir resolves once, at construction, in the usual order: the argument, then ALTAR_VARIANTS_DATA_DIR, then the installed package's data directory. The directory must hold region_annotations.parquet, ccres.dnatree, and region_provenance.json. Before the source loads any file, it checks each file's SHA-256 against the manifest, and it unpickles ccres.dnatree only from bytes it read once and checked. These checks catch drift and accidents, such as a rebuilt file without a rebuilt manifest. They do not protect against someone who can write the data directory, who can rewrite the manifest too. Load data only from directories you trust.

A variant on a chromosome without protein-coding genes in the gene table, such as an unplaced contig, has no row. Each variant costs a scan of its chromosome's genes, and the first call loads the full region index and cCRE tree into memory.

Warehouse path: the annotate container and a host table

At warehouse scale, run the annotate command in the variants runtime image (see runtimes/variants in the repository) and load its JSONL into a table the host owns:

python -m altar.variants.annotation.annotate -v variants.tsv -o annotated.jsonl

Give each input row its canonical variant ID in the fifth column, because the table is keyed by variant_id.

Next to annotated.jsonl, the command writes annotated.jsonl.identity.json after the last row:

{
  "format": 1,
  "identity": {
    "source_id": "org.kundajelab.altar.annotation.regions",
    "schema_version": "1.0.0",
    "release": "ensembl-116.ccre-v4.genes-gencode-47.logic-1",
    "genome_build": "hg38"
  },
  "identity_hash": "…",
  "columns": ["region_type", "region_labels", "…"],
  "provenance": {"format": 2, "genome_build": "hg38", "logic_version": 1, "inputs": {"…": "…"}, "derived": {"…": "…"}},
  "row_count": 2
}

The JSONL rows themselves are unchanged. A data directory without a provenance manifest still annotates, with a warning and no sidecar, and a run removes any sidecar an earlier run left at the same path.

The host declares its table with the contract's columns and the identity from the sidecar. Any generic source works; for BigQuery:

import json
from pathlib import Path

from altar.sources import AnnotationSourceIdentity, BigQueryAnnotationSource
from altar.variants import REGION_ANNOTATION_CONTRACT

sidecar = json.loads(Path("annotated.jsonl.identity.json").read_text())

regions = BigQueryAnnotationSource.connect(
    "my-project.annotations.variant_regions",
    REGION_ANNOTATION_CONTRACT.columns,
    name="regions",
    genome_default="GRCh38",
    identity=AnnotationSourceIdentity.from_dict(sidecar["identity"]),
    identity_scoped=True,
)

identity_scoped=True makes a writable or multi-release table store, deduplicate, and read rows by identity, so rows loaded under one release never answer for another. Use it whenever Altar writes the table or the table holds more than one release. Every source that reads the table must declare it the same way. A table built before identity scoping must be migrated first; see Migrate an existing BigQuery lake. SqliteAnnotationSource takes the same columns, identity, and identity_scoped.

The identity, not the source class, is what satisfies ChromBPNet's dependency on org.kundajelab.altar.annotation.regions version 1.0.0, so a host table declared this way and RegionAnnotationSource are interchangeable.

An existing table

The release does not depend on the derived files, so a table built earlier can declare it without a sidecar. A table built by equivalent logic from the same inputs (the Ensembl 116 GFF3, SCREEN v4, and the shipped GENCODE 47 gene table, with the windows and lookups of region logic 1) declares:

identity = REGION_ANNOTATION_CONTRACT.identity(
    release="ensembl-116.ccre-v4.genes-gencode-47.logic-1", genome_build="hg38"
)

Declare this only when it is true. A table built from other inputs or by other logic serves other values and needs its own release.