Region annotation source¶
Altar publishes the region annotations (region class, overlapping genes, nearest protein-coding genes, and the
overlapping SCREEN cCRE) as an annotation contract, org.kundajelab.altar.annotation.regions. Any source that
serves these columns with this contract's identity can feed materialization and prioritization: the portable
RegionAnnotationSource, or a warehouse table loaded from the annotate command's output. ChromBPNet's
prioritization policy depends on this contract's region_type column.
Importing the contract loads no lookup data, and neither pandas nor the interval-tree libraries.
The contract¶
REGION_ANNOTATION_CONTRACT is version 1.0.0. Its columns have the same names and logical types as the fields
altar.variants.annotation.annotate writes to each JSONL row, so those rows load directly into a table declared
with these columns.
| Column | Logical type | Meaning |
|---|---|---|
region_type |
str |
The most severe region label, such as cds, splice_site, promoter, or intergenic |
region_labels |
repeated str |
Every overlapping region label, most severe first |
region_gene_ids |
repeated str |
Ensembl IDs of the genes whose features the variant overlaps |
is_coding |
bool |
The variant overlaps a CDS |
in_promoter |
bool |
The variant overlaps a promoter window of any gene |
nearest_genes |
repeated struct |
The five nearest protein-coding genes: gene_name (str), gene_id (str), distance (int, 0 inside the gene), type (str) |
gene_within_100kb |
bool |
The nearest protein-coding gene lies within 100 kb |
ccre_id |
str, null when none |
Accession of an overlapping SCREEN cCRE |
ccre_group |
str, null when none |
That cCRE's group, such as PLS or dELS |
Region classification defines the labels and windows.
The columns come from two gene annotations. region_type, region_labels, region_gene_ids, is_coding, and
in_promoter come from the Ensembl GFF3 the region index was built from (Ensembl 116 by default).
nearest_genes and gene_within_100kb come from the shipped gene table, GENCODE 47 (Ensembl 113). A gene ID can
therefore appear in one and not the other, and the two can name a gene differently.
Each JSONL row also carries fields outside the contract:
ref_length,alt_length,variant_length, andvariant_typederive from the alleles alone. They have no data release, and any consumer can recompute them from the variant ID.in_otand theaf_*population frequencies come from Open Targets' gnomAD-derived variant table. That is a different dataset with its own release, so it belongs to its own contract.
A test checks the contract against the annotate command's row model, so a field added to the rows must be either added to the contract in a new contract version or listed as outside it.
Identity and data release¶
A source implementing the contract reports an identity: the contract's source_id and schema_version, the
data release, and the genome_build. The region data is GRCh38, so the build is always hg38.
The release follows the science, not the files. It names the three inputs the values come from and the version
of the logic that turns them into columns, which the data directory's provenance manifest,
region_provenance.json, records (see Annotation data files):
ensembl-116: the Ensembl GFF3 thatregion_annotations.parquetwas built from.ccre-v4: the SCREEN cCRE Registry BED thatccres.dnatreewas built from.genes-gencode-47: the nearest-gene table.logic-1:REGION_LOGIC_VERSION, which covers the build scripts' windows and the lookups' semantics. A change to either that changes any value bumps it, and a golden-output test fails until it does.
Altar pins the SHA-256 of the Ensembl 116 GFF3, the SCREEN v4 BED, and the GENCODE 47 gene table. When every
input has its pinned digest, the string above is the whole release. Otherwise, for example for another Ensembl
release or your own gene table, . and the first 12 hex characters of a digest over the three input digests
follow.
The derived files' bytes are not part of the release. A parquet written by another pandas or pyarrow version names the same release, while different inputs or different logic name a different one. The manifest still records each derived file's SHA-256, and annotation checks the files against it before it reads them.
Portable path: RegionAnnotationSource¶
RegionAnnotationSource(data_dir=None, genome_build="hg38") computes the contract columns in Python with the
same function the annotate command uses for its rows. Use it for tests, notebooks, and small batches.
from altar.sources import fetch_annotations
from altar.variants import RegionAnnotationSource
source = RegionAnnotationSource(data_dir="/data/altar-variants")
source.identity() # the contract identity, with the release from region_provenance.json
annotations = await fetch_annotations([source], ["chr1:1014143:C:T"])
data_dir resolves once, at construction, in the usual order: the argument, then ALTAR_VARIANTS_DATA_DIR,
then the installed package's data directory. The directory must hold region_annotations.parquet,
ccres.dnatree, and region_provenance.json. Before the source loads any file, it checks each file's SHA-256
against the manifest, and it unpickles ccres.dnatree only from bytes it read once and checked. These checks
catch drift and accidents, such as a rebuilt file without a rebuilt manifest. They do not protect against someone
who can write the data directory, who can rewrite the manifest too. Load data only from directories you trust.
A variant on a chromosome without protein-coding genes in the gene table, such as an unplaced contig, has no row. Each variant costs a scan of its chromosome's genes, and the first call loads the full region index and cCRE tree into memory.
Warehouse path: the annotate container and a host table¶
At warehouse scale, run the annotate command in the variants runtime image (see runtimes/variants in the
repository) and load its JSONL into a table the host owns:
Give each input row its canonical variant ID in the fifth column, because the table is keyed by variant_id.
Next to annotated.jsonl, the command writes annotated.jsonl.identity.json after the last row:
{
"format": 1,
"identity": {
"source_id": "org.kundajelab.altar.annotation.regions",
"schema_version": "1.0.0",
"release": "ensembl-116.ccre-v4.genes-gencode-47.logic-1",
"genome_build": "hg38"
},
"identity_hash": "…",
"columns": ["region_type", "region_labels", "…"],
"provenance": {"format": 2, "genome_build": "hg38", "logic_version": 1, "inputs": {"…": "…"}, "derived": {"…": "…"}},
"row_count": 2
}
The JSONL rows themselves are unchanged. A data directory without a provenance manifest still annotates, with a warning and no sidecar, and a run removes any sidecar an earlier run left at the same path.
The host declares its table with the contract's columns and the identity from the sidecar. Any generic source works; for BigQuery:
import json
from pathlib import Path
from altar.sources import AnnotationSourceIdentity, BigQueryAnnotationSource
from altar.variants import REGION_ANNOTATION_CONTRACT
sidecar = json.loads(Path("annotated.jsonl.identity.json").read_text())
regions = BigQueryAnnotationSource.connect(
"my-project.annotations.variant_regions",
REGION_ANNOTATION_CONTRACT.columns,
name="regions",
genome_default="GRCh38",
identity=AnnotationSourceIdentity.from_dict(sidecar["identity"]),
identity_scoped=True,
)
identity_scoped=True makes a writable or multi-release table store, deduplicate, and read rows by identity, so
rows loaded under one release never answer for another. Use it whenever Altar writes the table or the table holds
more than one release. Every source that reads the table must declare it the same way. A table built before
identity scoping must be migrated first; see
Migrate an existing BigQuery lake.
SqliteAnnotationSource takes the same columns, identity, and identity_scoped.
The identity, not the source class, is what satisfies ChromBPNet's dependency on
org.kundajelab.altar.annotation.regions version 1.0.0, so a host table declared this way and
RegionAnnotationSource are interchangeable.
An existing table¶
The release does not depend on the derived files, so a table built earlier can declare it without a sidecar. A table built by equivalent logic from the same inputs (the Ensembl 116 GFF3, SCREEN v4, and the shipped GENCODE 47 gene table, with the windows and lookups of region logic 1) declares:
identity = REGION_ANNOTATION_CONTRACT.identity(
release="ensembl-116.ccre-v4.genes-gencode-47.logic-1", genome_build="hg38"
)
Declare this only when it is true. A table built from other inputs or by other logic serves other values and needs its own release.