Use your own annotation tables¶
A host application can supply annotations from its own tables or services and compose them with the sources
Altar ships, such as AlphaMissense. A host source is an ordinary AnnotationSource instance that you pass to
materialize. It does not need a package, a registry entry, or an entry point.
Altar checks, joins, and prioritizes on what a source declares: its columns from annotation_columns() and
its identity from identity(). It never looks at the source's class or registry name.
Contracts and identity¶
An AnnotationContract is a published, versioned column set. Its source_id is a namespaced name, such as
org.kundajelab.altar.annotation.alphamissense. Its schema_version is a semantic version. Two versions with
the same major version have compatible columns, and a minor version only adds columns.
An AnnotationSourceIdentity names the contract a concrete source implements, plus the data release and
genome_build it serves. Several sources can implement one contract. For example, a warehouse table you
maintain can implement the same contract as a shipped binding.
from altar.sources import AnnotationColumn, AnnotationContract
REGULATORY = AnnotationContract(
source_id="org.example.host.annotation.regulatory",
schema_version="1.0.0",
columns=(
AnnotationColumn("regulatory_class", "str", "Regulatory class"),
AnnotationColumn("regulatory_score", "float", "Regulatory score"),
),
)
identity = REGULATORY.identity(release="2026-09", genome_build="hg38")
Change release whenever the values the source returns could change, such as a new upstream download.
identity.identity_hash is a deterministic hash of the four identity fields. Spell genome_build the same
way as the genome build in your models' run identity, such as hg38.
Shipped bindings publish their contracts as module constants: ALPHAMISSENSE_ANNOTATIONS in
altar_alphamissense, SPLICEAI_ANNOTATIONS in altar_spliceai, and GPNSTAR_ANNOTATIONS in altar_gpnstar.
Write a source class¶
Subclass AnnotationSource and return the identity from identity(). The source's columns must include the
contract's columns.
from altar.sources import AnnotationSource
class RegulatoryTable(AnnotationSource):
name = "regulatory"
def __init__(self, client, *, release: str, genome_build: str):
self.client = client
self._identity = REGULATORY.identity(release=release, genome_build=genome_build)
def annotation_columns(self):
return list(REGULATORY.columns)
def identity(self):
return self._identity
async def annotate(self, variant_ids):
rows = await self.client.lookup(variant_ids)
return {row["variant_id"]: {"regulatory_class": row["class"], "regulatory_score": row["score"]} for row in rows}
annotate() omits variants the table has no row for. See Add an annotation source
for the rest of the AnnotationSource contract.
Configure a generic source over your table¶
When the table already has one row per variant_id, configure BigQueryAnnotationSource or
SqliteAnnotationSource instead of writing a class. Pass identity= to declare the contract the table
implements.
from altar.sources import BigQueryAnnotationSource
regulatory = BigQueryAnnotationSource.connect(
"my-project.annotations.regulatory_v2026_09",
list(REGULATORY.columns),
name="regulatory",
identity=REGULATORY.identity(release="2026-09", genome_build="hg38"),
)
On its own, the identity is metadata. It does not change how a source writes, deduplicates, or reads rows. To
make it decide which rows a lake serves and counts, also pass identity_scoped=True. See
Keep a writable lake current across releases.
identity.genome_build and genome_default are separate settings. identity.genome_build is the build label
that dependency resolution compares with your score runs, such as hg38. genome_default is how a writable
lake spells the build in its stored genome column. Altar's VCF ingest writes GRCh38 there. A lake whose rows
say GRCh38 keeps genome_default="GRCh38" and can still declare an hg38 identity. Neither setting fills in
the other.
Keep a writable lake current across releases¶
A writable lake is a table that Altar writes annotations into and then reads back as a cache, such as the
region labels your annotation jobs compute. Pass identity_scoped=True with identity= to make the lake keep
each release's rows apart. Both generic sources take the flag. A scoped BigQueryAnnotationSource stores the
identity's hash on every row and uses only rows with that hash. A scoped SqliteAnnotationSource records one
identity per file and refuses to open the file under another. In both, changing release makes every variant
a cache miss, and your pipeline annotates them again. Without the flag, identity= is metadata on both.
A scoped BigQueryAnnotationSource works like this:
ensure_schema()adds a nullable_altar_annotation_identity STRINGcolumn. The name is exported asaltar.sources.ANNOTATION_IDENTITY_COLUMN.add_annotationsandadd_annotations_framewrite the identity'sidentity_hashinto that column on every row. A row that already carries a different hash raisesAnnotationIdentityError.get_unannotated_variantsandexport_unannotatedcount a variant as annotated only when it has a row with that hash.annotateand materialization read only rows with that hash. A lake that holds several releases still joins one row per variant.
regions = BigQueryAnnotationSource.connect(
"my-project.annotations.variant_regions",
list(REGIONS.columns),
name="regions",
genome_default="GRCh38",
identity=REGIONS.identity(release="gencode-v47-screen-v4", genome_build="hg38"),
identity_scoped=True,
)
await regions.ensure_schema()
missing = await regions.get_unannotated_variants([], variants_relation="`my-project.jobs.variants_123`")
Scope every source object that reads or writes the lake, with the same identity. A reader you configure
separately from the writer needs identity_scoped=True too. Otherwise it reads every release's rows, and a
variant can be served an arbitrary release's values.
An unscoped BigQueryAnnotationSource checks for this. It looks up its table's schema, a metadata request that
runs no query and bills nothing. A table with the _altar_annotation_identity column is a scoped lake, and the
source then:
- refuses to write to it.
add_annotationsandadd_annotations_frameraiseAnnotationIdentityError, because rows written without the identity hash are neither served nor counted by the lake's scoped sources. A SQLite lake refuses an unscoped write the same way. - warns before its first read or deduplication. Materialization, the BigQuery exports,
annotate(),get_unannotated_variantsandexport_unannotatedemitUnscopedAnnotationReadWarning(altar.sources.UnscopedAnnotationReadWarning) and then run as before. An unscoped deduplication counts a variant annotated under any release as annotated. Each source object warns at most once.
If the lookup itself fails, for example for lack of permission, Altar logs it at debug level and the operation proceeds unchecked. To make the warning fail a job instead, turn it into an error:
import warnings
from altar.sources import UnscopedAnnotationReadWarning
warnings.simplefilter("error", UnscopedAnnotationReadWarning)
Relations over a lake¶
Altar cannot add a filter to SQL you own. A scoped source's relation= must contain the {identity_filter}
token, which expands to _altar_annotation_identity = '<identity_hash>'. A scoped source whose relation lacks
the token raises ValueError. An unscoped source whose relation uses it raises ValueError too.
reader = BigQueryAnnotationSource.connect(
"my-project.annotations.variant_regions",
list(REGIONS.columns),
name="regions",
relation=(
"(SELECT variant_id, region_type FROM `my-project.annotations.variant_regions` "
"WHERE {identity_filter} AND {variant_filter})"
),
identity=REGIONS.identity(release="gencode-v47-screen-v4", genome_build="hg38"),
identity_scoped=True,
)
A reader needs no genome_default. Only writes and deduplication use it.
Altar does not check reads through an unscoped relation=. Only your SQL names the tables it reads, and
Altar would have to run that SQL to find out. Writes and deduplication always use the source's bare table, so
Altar checks those. A relation over a scoped lake without {identity_filter} reads every
release's rows, with no warning. Review each relation that reads a scoped lake: scope its source and add the
token.
SQLite lakes¶
A scoped SqliteAnnotationSource records its identity, with its hash and readable fields, in the file's
_altar_annotation_identity table the first time it opens a file with no rows. A scoped source with a
different identity cannot open the file and raises AnnotationIdentityError, so a new release starts a new
file. An unscoped source can read any file, but it cannot write to a file that records an identity.
Migrate an existing BigQuery lake¶
A lake written before identities has no _altar_annotation_identity column, so a scoped source's queries
against it fail until ensure_schema() adds the column. After that, the existing rows have a NULL identity.
A scoped source neither serves nor counts them. Until you run the backfill below, deduplication reports every
variant as unannotated, and your next job annotates the whole lake again.
These steps migrate a lake with the columns variant_id, genome, chr, pos, ref, alt plus annotation columns,
clustered on genome, chr, variant_id:
- Find the release that built the existing rows. This is usually the gene annotation and cCRE releases your pipeline used when it wrote them. If rows from several releases are mixed in the table, you cannot adopt them under one identity. Start a new table instead.
- Pause every job that writes to the table.
- Configure the writer with that old release,
identity_scoped=True, and the lake's stored build asgenome_default. Callensure_schema(). It adds the nullable_altar_annotation_identity STRINGcolumn and leaves existing rows and clustering as they are. - If any row has a NULL
genome, callbackfill_identity_columns()first. It fillsgenomeand the locus columns. - Call
backfill_annotation_identity(). It writes the identity hash into every row ofgenome_default's build that has none. It keeps one row per variant, and deletes duplicates and rows already superseded by a row with that hash. Among duplicates it keeps the newest bycreated_atwhen the table has that column, then the first by the row's JSON text, so a rerun picks the same row. It returns the counts asIdentityBackfill(stamped=..., deleted=...). A second call changes nothing. -
Check the result. For the
genome_defaultbuild, this query should show only your identity's hash and no NULL row: -
Give every other source object that reads the lake the same old-release identity and
identity_scoped=True. Add{identity_filter}to each reader'srelation=. - Make any load path that bypasses
add_annotationsandadd_annotations_frame, such as a load job from a file URI, writeidentity.identity_hashinto_altar_annotation_identity. A row loaded without the hash is never counted, so every job annotates that variant again and appends another row. - Resume ingest.
legacy = BigQueryAnnotationSource.connect(
"my-project.annotations.variant_regions",
list(REGIONS.columns),
name="regions",
genome_default="GRCh38",
identity=REGIONS.identity(release="gencode-v46-screen-v4", genome_build="hg38"),
identity_scoped=True,
)
await legacy.ensure_schema()
print(await legacy.backfill_annotation_identity())
The backfill states that every legacy row came from the configured release. Review that before you run it. If you adopt rows under a newer release than the one that built them, the lake serves old values under the new identity.
The backfill runs as one BigQuery transaction. It copies the legacy rows to a temporary table, deletes them, and inserts them back with the identity. When most of the table is legacy, it rewrites all of the table's storage. On-demand pricing bills about three times the table's logical size. For example, a 25-million-row lake of 10 GB bills about 30 GB. BigQuery re-clusters the rewritten rows in the background at no charge. The replaced storage stays in time travel and fail-safe for their configured windows. Datasets on physical storage billing pay for that storage.
identity_hash cannot be reversed into the release it names. To find which hash a release produced, build
its identity and read identity.identity_hash.
For a SQLite lake, open the existing file with the old release's identity= and identity_scoped=True. The
source refuses to read or write it until you call backfill_annotation_identity(). That call records the
identity in the file and returns IdentityBackfill(stamped=<rows adopted>, deleted=0).
Roll out a new release¶
Bump the writer and the readers in two steps:
- Change
releaseon the writer. Its next deduplication reports every variant as unannotated. Run jobs until it reports none for the variants you serve. The readers keep serving the old release meanwhile. - Then change
releaseon every reader.
A reader bumped before re-annotation finishes serves nothing for the variants not yet annotated under the new
release. Their annotation columns are empty, so a prioritization rule that reads them, such as ChromBPNet's
region_type, silently leaves those variants unprioritized.
Compose with shipped sources¶
Pass host sources and binding sources together. Column names must be unique across all sources.
async for result in store.materialize(
variant_ids=variant_ids,
models=models,
annotation_sources=[regulatory, alphamissense],
):
consume(result)
Annotation dependencies¶
A model's prioritization rule can read annotation columns. Its manifest declares them as
AnnotationDependency values. Each dependency names a contract by source_id, a minimum schema_version, and
the columns the rule reads. ChromBPNet depends on region_type from org.kundajelab.altar.annotation.regions,
which altar.variants publishes as REGION_ANNOTATION_CONTRACT.
Before it evaluates any rule, materialize checks each dependency against the configured sources with
resolve_annotation_dependencies. The BigQuery exports that compute prioritized run the same check.
- If a source's identity names the dependency's contract, that source must declare every dependency column.
Its contract version must have the same major version as the dependency and be no older. Only one source may
declare that contract. When Altar knows the contract, each dependency column must also have the contract's
logical type: the same
dtype,repeatedflag and nestedfields. Display labels may differ. OtherwiseAnnotationDependencyErroris raised. The message names the column, the type the source declares, and the type the contract defines. - If no source declares the contract, but other sources declare every dependency column, the columns are used
by name. Altar emits an
UndeclaredAnnotationContractWarningthat names the models and the contract. - If any dependency column is missing from every source,
AnnotationDependencyErroris raised. It names the contract and the missing columns.
Every source that declares an identity must declare the same genome_build. A model whose run identity
records a genome build must match it. Otherwise AnnotationGenomeBuildError is raised.
Altar knows the contracts it publishes, such as REGION_ANNOTATION_CONTRACT. It does not know a binding's
contract or yours, so for those the check in step 1 covers column names only. To check the column types of
your own contract too, pass it when you call the resolver yourself:
resolve_annotation_dependencies(models, sources, contracts=[REGULATORY]). Materialization does not take
contracts. To check a source against every column of its contract, use
altar.testing.assert_implements_contract (see Test a host source).
To clear the warning, declare the contract on the source that supplies the columns: pass identity= to a
generic source, or override identity().
For ChromBPNet's dependency, declare the published region contract on the table that holds the region columns. Region annotation source explains which release to name:
from altar.sources import BigQueryAnnotationSource
from altar.variants import REGION_ANNOTATION_CONTRACT
regions = BigQueryAnnotationSource.connect(
"my-project.annotations.variant_regions",
REGION_ANNOTATION_CONTRACT.columns,
name="regions",
genome_default="GRCh38",
identity=REGION_ANNOTATION_CONTRACT.identity(
release="ensembl-116.ccre-v4.genes-gencode-47.logic-1", genome_build="hg38"
),
)
For your own contracts, name a release that changes whenever the values can change, such as the upstream
releases the table was built from. If Altar also writes this table, add identity_scoped=True so a release
change re-annotates it, and migrate the existing table first. Call
resolve_annotation_dependencies(models, sources) yourself to check a configuration before you submit work. It
returns one ResolvedAnnotationDependency per model
dependency, with the sources that satisfied it.
Test a host source¶
Run the same conformance suite that shipped sources run. Return the contract from the contract fixture to
also check that the source implements it. A source you also write to can run
altar.testing.WritableAnnotationSourceContract as well; see Validate an implementation.
import pytest
from altar.testing import AnnotationSourceContract
class TestRegulatoryTable(AnnotationSourceContract):
@pytest.fixture
def source(self, fake_client):
return RegulatoryTable(fake_client, release="test", genome_build="hg38")
@pytest.fixture
def known_variant_ids(self):
return ["chr1:100:A:G"]
@pytest.fixture
def contract(self):
return REGULATORY
The contract check requires the source's identity to name a compatible version of the contract. Each contract
column must appear with the same dtype, repeated flag, and fields. Display labels may differ. Call
altar.testing.assert_implements_contract(source, contract) directly to check a configured generic source.