Skip to content

Use your own annotation tables

A host application can supply annotations from its own tables or services and compose them with the sources Altar ships, such as AlphaMissense. A host source is an ordinary AnnotationSource instance that you pass to materialize. It does not need a package, a registry entry, or an entry point.

Altar checks, joins, and prioritizes on what a source declares: its columns from annotation_columns() and its identity from identity(). It never looks at the source's class or registry name.

Contracts and identity

An AnnotationContract is a published, versioned column set. Its source_id is a namespaced name, such as org.kundajelab.altar.annotation.alphamissense. Its schema_version is a semantic version. Two versions with the same major version have compatible columns, and a minor version only adds columns.

An AnnotationSourceIdentity names the contract a concrete source implements, plus the data release and genome_build it serves. Several sources can implement one contract. For example, a warehouse table you maintain can implement the same contract as a shipped binding.

from altar.sources import AnnotationColumn, AnnotationContract

REGULATORY = AnnotationContract(
    source_id="org.example.host.annotation.regulatory",
    schema_version="1.0.0",
    columns=(
        AnnotationColumn("regulatory_class", "str", "Regulatory class"),
        AnnotationColumn("regulatory_score", "float", "Regulatory score"),
    ),
)
identity = REGULATORY.identity(release="2026-09", genome_build="hg38")

Change release whenever the values the source returns could change, such as a new upstream download. identity.identity_hash is a deterministic hash of the four identity fields. Spell genome_build the same way as the genome build in your models' run identity, such as hg38.

Shipped bindings publish their contracts as module constants: ALPHAMISSENSE_ANNOTATIONS in altar_alphamissense, SPLICEAI_ANNOTATIONS in altar_spliceai, and GPNSTAR_ANNOTATIONS in altar_gpnstar.

Write a source class

Subclass AnnotationSource and return the identity from identity(). The source's columns must include the contract's columns.

from altar.sources import AnnotationSource


class RegulatoryTable(AnnotationSource):
    name = "regulatory"

    def __init__(self, client, *, release: str, genome_build: str):
        self.client = client
        self._identity = REGULATORY.identity(release=release, genome_build=genome_build)

    def annotation_columns(self):
        return list(REGULATORY.columns)

    def identity(self):
        return self._identity

    async def annotate(self, variant_ids):
        rows = await self.client.lookup(variant_ids)
        return {row["variant_id"]: {"regulatory_class": row["class"], "regulatory_score": row["score"]} for row in rows}

annotate() omits variants the table has no row for. See Add an annotation source for the rest of the AnnotationSource contract.

Configure a generic source over your table

When the table already has one row per variant_id, configure BigQueryAnnotationSource or SqliteAnnotationSource instead of writing a class. Pass identity= to declare the contract the table implements.

from altar.sources import BigQueryAnnotationSource

regulatory = BigQueryAnnotationSource.connect(
    "my-project.annotations.regulatory_v2026_09",
    list(REGULATORY.columns),
    name="regulatory",
    identity=REGULATORY.identity(release="2026-09", genome_build="hg38"),
)

On its own, the identity is metadata. It does not change how a source writes, deduplicates, or reads rows. To make it decide which rows a lake serves and counts, also pass identity_scoped=True. See Keep a writable lake current across releases.

identity.genome_build and genome_default are separate settings. identity.genome_build is the build label that dependency resolution compares with your score runs, such as hg38. genome_default is how a writable lake spells the build in its stored genome column. Altar's VCF ingest writes GRCh38 there. A lake whose rows say GRCh38 keeps genome_default="GRCh38" and can still declare an hg38 identity. Neither setting fills in the other.

Keep a writable lake current across releases

A writable lake is a table that Altar writes annotations into and then reads back as a cache, such as the region labels your annotation jobs compute. Pass identity_scoped=True with identity= to make the lake keep each release's rows apart. Both generic sources take the flag. A scoped BigQueryAnnotationSource stores the identity's hash on every row and uses only rows with that hash. A scoped SqliteAnnotationSource records one identity per file and refuses to open the file under another. In both, changing release makes every variant a cache miss, and your pipeline annotates them again. Without the flag, identity= is metadata on both.

A scoped BigQueryAnnotationSource works like this:

  • ensure_schema() adds a nullable _altar_annotation_identity STRING column. The name is exported as altar.sources.ANNOTATION_IDENTITY_COLUMN.
  • add_annotations and add_annotations_frame write the identity's identity_hash into that column on every row. A row that already carries a different hash raises AnnotationIdentityError.
  • get_unannotated_variants and export_unannotated count a variant as annotated only when it has a row with that hash.
  • annotate and materialization read only rows with that hash. A lake that holds several releases still joins one row per variant.
regions = BigQueryAnnotationSource.connect(
    "my-project.annotations.variant_regions",
    list(REGIONS.columns),
    name="regions",
    genome_default="GRCh38",
    identity=REGIONS.identity(release="gencode-v47-screen-v4", genome_build="hg38"),
    identity_scoped=True,
)
await regions.ensure_schema()
missing = await regions.get_unannotated_variants([], variants_relation="`my-project.jobs.variants_123`")

Scope every source object that reads or writes the lake, with the same identity. A reader you configure separately from the writer needs identity_scoped=True too. Otherwise it reads every release's rows, and a variant can be served an arbitrary release's values.

An unscoped BigQueryAnnotationSource checks for this. It looks up its table's schema, a metadata request that runs no query and bills nothing. A table with the _altar_annotation_identity column is a scoped lake, and the source then:

  • refuses to write to it. add_annotations and add_annotations_frame raise AnnotationIdentityError, because rows written without the identity hash are neither served nor counted by the lake's scoped sources. A SQLite lake refuses an unscoped write the same way.
  • warns before its first read or deduplication. Materialization, the BigQuery exports, annotate(), get_unannotated_variants and export_unannotated emit UnscopedAnnotationReadWarning (altar.sources.UnscopedAnnotationReadWarning) and then run as before. An unscoped deduplication counts a variant annotated under any release as annotated. Each source object warns at most once.

If the lookup itself fails, for example for lack of permission, Altar logs it at debug level and the operation proceeds unchecked. To make the warning fail a job instead, turn it into an error:

import warnings
from altar.sources import UnscopedAnnotationReadWarning

warnings.simplefilter("error", UnscopedAnnotationReadWarning)

Relations over a lake

Altar cannot add a filter to SQL you own. A scoped source's relation= must contain the {identity_filter} token, which expands to _altar_annotation_identity = '<identity_hash>'. A scoped source whose relation lacks the token raises ValueError. An unscoped source whose relation uses it raises ValueError too.

reader = BigQueryAnnotationSource.connect(
    "my-project.annotations.variant_regions",
    list(REGIONS.columns),
    name="regions",
    relation=(
        "(SELECT variant_id, region_type FROM `my-project.annotations.variant_regions` "
        "WHERE {identity_filter} AND {variant_filter})"
    ),
    identity=REGIONS.identity(release="gencode-v47-screen-v4", genome_build="hg38"),
    identity_scoped=True,
)

A reader needs no genome_default. Only writes and deduplication use it.

Altar does not check reads through an unscoped relation=. Only your SQL names the tables it reads, and Altar would have to run that SQL to find out. Writes and deduplication always use the source's bare table, so Altar checks those. A relation over a scoped lake without {identity_filter} reads every release's rows, with no warning. Review each relation that reads a scoped lake: scope its source and add the token.

SQLite lakes

A scoped SqliteAnnotationSource records its identity, with its hash and readable fields, in the file's _altar_annotation_identity table the first time it opens a file with no rows. A scoped source with a different identity cannot open the file and raises AnnotationIdentityError, so a new release starts a new file. An unscoped source can read any file, but it cannot write to a file that records an identity.

Migrate an existing BigQuery lake

A lake written before identities has no _altar_annotation_identity column, so a scoped source's queries against it fail until ensure_schema() adds the column. After that, the existing rows have a NULL identity. A scoped source neither serves nor counts them. Until you run the backfill below, deduplication reports every variant as unannotated, and your next job annotates the whole lake again.

These steps migrate a lake with the columns variant_id, genome, chr, pos, ref, alt plus annotation columns, clustered on genome, chr, variant_id:

  1. Find the release that built the existing rows. This is usually the gene annotation and cCRE releases your pipeline used when it wrote them. If rows from several releases are mixed in the table, you cannot adopt them under one identity. Start a new table instead.
  2. Pause every job that writes to the table.
  3. Configure the writer with that old release, identity_scoped=True, and the lake's stored build as genome_default. Call ensure_schema(). It adds the nullable _altar_annotation_identity STRING column and leaves existing rows and clustering as they are.
  4. If any row has a NULL genome, call backfill_identity_columns() first. It fills genome and the locus columns.
  5. Call backfill_annotation_identity(). It writes the identity hash into every row of genome_default's build that has none. It keeps one row per variant, and deletes duplicates and rows already superseded by a row with that hash. Among duplicates it keeps the newest by created_at when the table has that column, then the first by the row's JSON text, so a rerun picks the same row. It returns the counts as IdentityBackfill(stamped=..., deleted=...). A second call changes nothing.
  6. Check the result. For the genome_default build, this query should show only your identity's hash and no NULL row:

    SELECT genome, _altar_annotation_identity, COUNT(*) AS n
    FROM `my-project.annotations.variant_regions`
    GROUP BY 1, 2
    
  7. Give every other source object that reads the lake the same old-release identity and identity_scoped=True. Add {identity_filter} to each reader's relation=.

  8. Make any load path that bypasses add_annotations and add_annotations_frame, such as a load job from a file URI, write identity.identity_hash into _altar_annotation_identity. A row loaded without the hash is never counted, so every job annotates that variant again and appends another row.
  9. Resume ingest.
legacy = BigQueryAnnotationSource.connect(
    "my-project.annotations.variant_regions",
    list(REGIONS.columns),
    name="regions",
    genome_default="GRCh38",
    identity=REGIONS.identity(release="gencode-v46-screen-v4", genome_build="hg38"),
    identity_scoped=True,
)
await legacy.ensure_schema()
print(await legacy.backfill_annotation_identity())

The backfill states that every legacy row came from the configured release. Review that before you run it. If you adopt rows under a newer release than the one that built them, the lake serves old values under the new identity.

The backfill runs as one BigQuery transaction. It copies the legacy rows to a temporary table, deletes them, and inserts them back with the identity. When most of the table is legacy, it rewrites all of the table's storage. On-demand pricing bills about three times the table's logical size. For example, a 25-million-row lake of 10 GB bills about 30 GB. BigQuery re-clusters the rewritten rows in the background at no charge. The replaced storage stays in time travel and fail-safe for their configured windows. Datasets on physical storage billing pay for that storage.

identity_hash cannot be reversed into the release it names. To find which hash a release produced, build its identity and read identity.identity_hash.

For a SQLite lake, open the existing file with the old release's identity= and identity_scoped=True. The source refuses to read or write it until you call backfill_annotation_identity(). That call records the identity in the file and returns IdentityBackfill(stamped=<rows adopted>, deleted=0).

Roll out a new release

Bump the writer and the readers in two steps:

  1. Change release on the writer. Its next deduplication reports every variant as unannotated. Run jobs until it reports none for the variants you serve. The readers keep serving the old release meanwhile.
  2. Then change release on every reader.

A reader bumped before re-annotation finishes serves nothing for the variants not yet annotated under the new release. Their annotation columns are empty, so a prioritization rule that reads them, such as ChromBPNet's region_type, silently leaves those variants unprioritized.

Compose with shipped sources

Pass host sources and binding sources together. Column names must be unique across all sources.

async for result in store.materialize(
    variant_ids=variant_ids,
    models=models,
    annotation_sources=[regulatory, alphamissense],
):
    consume(result)

Annotation dependencies

A model's prioritization rule can read annotation columns. Its manifest declares them as AnnotationDependency values. Each dependency names a contract by source_id, a minimum schema_version, and the columns the rule reads. ChromBPNet depends on region_type from org.kundajelab.altar.annotation.regions, which altar.variants publishes as REGION_ANNOTATION_CONTRACT.

Before it evaluates any rule, materialize checks each dependency against the configured sources with resolve_annotation_dependencies. The BigQuery exports that compute prioritized run the same check.

  1. If a source's identity names the dependency's contract, that source must declare every dependency column. Its contract version must have the same major version as the dependency and be no older. Only one source may declare that contract. When Altar knows the contract, each dependency column must also have the contract's logical type: the same dtype, repeated flag and nested fields. Display labels may differ. Otherwise AnnotationDependencyError is raised. The message names the column, the type the source declares, and the type the contract defines.
  2. If no source declares the contract, but other sources declare every dependency column, the columns are used by name. Altar emits an UndeclaredAnnotationContractWarning that names the models and the contract.
  3. If any dependency column is missing from every source, AnnotationDependencyError is raised. It names the contract and the missing columns.

Every source that declares an identity must declare the same genome_build. A model whose run identity records a genome build must match it. Otherwise AnnotationGenomeBuildError is raised.

Altar knows the contracts it publishes, such as REGION_ANNOTATION_CONTRACT. It does not know a binding's contract or yours, so for those the check in step 1 covers column names only. To check the column types of your own contract too, pass it when you call the resolver yourself: resolve_annotation_dependencies(models, sources, contracts=[REGULATORY]). Materialization does not take contracts. To check a source against every column of its contract, use altar.testing.assert_implements_contract (see Test a host source).

To clear the warning, declare the contract on the source that supplies the columns: pass identity= to a generic source, or override identity().

For ChromBPNet's dependency, declare the published region contract on the table that holds the region columns. Region annotation source explains which release to name:

from altar.sources import BigQueryAnnotationSource
from altar.variants import REGION_ANNOTATION_CONTRACT

regions = BigQueryAnnotationSource.connect(
    "my-project.annotations.variant_regions",
    REGION_ANNOTATION_CONTRACT.columns,
    name="regions",
    genome_default="GRCh38",
    identity=REGION_ANNOTATION_CONTRACT.identity(
        release="ensembl-116.ccre-v4.genes-gencode-47.logic-1", genome_build="hg38"
    ),
)

For your own contracts, name a release that changes whenever the values can change, such as the upstream releases the table was built from. If Altar also writes this table, add identity_scoped=True so a release change re-annotates it, and migrate the existing table first. Call resolve_annotation_dependencies(models, sources) yourself to check a configuration before you submit work. It returns one ResolvedAnnotationDependency per model dependency, with the sources that satisfied it.

Test a host source

Run the same conformance suite that shipped sources run. Return the contract from the contract fixture to also check that the source implements it. A source you also write to can run altar.testing.WritableAnnotationSourceContract as well; see Validate an implementation.

import pytest
from altar.testing import AnnotationSourceContract


class TestRegulatoryTable(AnnotationSourceContract):
    @pytest.fixture
    def source(self, fake_client):
        return RegulatoryTable(fake_client, release="test", genome_build="hg38")

    @pytest.fixture
    def known_variant_ids(self):
        return ["chr1:100:A:G"]

    @pytest.fixture
    def contract(self):
        return REGULATORY

The contract check requires the source's identity to name a compatible version of the contract. Each contract column must appear with the same dtype, repeated flag, and fields. Display labels may differ. Call altar.testing.assert_implements_contract(source, contract) directly to check a configured generic source.