Skip to content

Add an annotation source

Use AnnotationSource for precomputed per-variant evidence.

1. Declare logical columns

from altar.sources import AnnotationColumn, AnnotationSource


class ConservationSource(AnnotationSource):
    name = "conservation"

    def annotation_columns(self):
        return [AnnotationColumn("phylo_score", "float", "Phylogenetic score")]

    async def annotate(self, variant_ids):
        rows = await self.backend.fetch(variant_ids)
        return {row.variant_id: {"phylo_score": row.score} for row in rows}

Omit variants with no source row. Return an observed zero as zero. Validate that a backend never returns an unrequested variant or duplicate logical record.

Repeated and structured columns should be declared as such even if a simple physical backend serializes them as JSON.

2. Publish a contract and declare the identity

Publish the columns as an AnnotationContract constant so other sources, including a host's own tables, can implement the same contract. Return the contract's identity from identity(), with the data release and genome build the source serves.

from altar.sources import AnnotationColumn, AnnotationContract

CONSERVATION_ANNOTATIONS = AnnotationContract(
    source_id="org.example.annotation.conservation",
    schema_version="1.0.0",
    columns=(AnnotationColumn("phylo_score", "float", "Phylogenetic score"),),
)


class ConservationSource(AnnotationSource):
    ...

    def annotation_columns(self):
        return list(CONSERVATION_ANNOTATIONS.columns)

    def identity(self):
        return CONSERVATION_ANNOTATIONS.identity(release=self.release, genome_build=self.genome_build)

A model's AnnotationDependency names a contract by source_id. Materialization resolves each dependency against the configured sources' identities; see Use your own annotation tables. Adding columns is a minor version. Removing or redefining a column is a major version.

3. Add a rule only when justified

Return None from prioritize_predicate() for passive evidence. If the source has a documented policy, return a predicate from altar.predicates and expose configurable thresholds where appropriate.

from altar.predicates import Col, Ge, Lit


def prioritize_predicate(self):
    return Ge(Col("phylo_score"), Lit(self.threshold))

4. Separate science from storage

The source should normalize scientific meaning. For source-native tables keyed by chromosome, position, reference, and alternate allele, accept the shared AllelicRecordBackend rather than defining a new backend protocol and a new Parquet query class for every source. The binding creates an AllelicRecordQuery, decodes the projected native fields, and owns all aggregation. The backend only retrieves records.

from altar.sources import AllelicKeyColumns, ParquetAllelicRecordBackend

backend = ParquetAllelicRecordBackend(
    "./conservation-records",
    key_columns=AllelicKeyColumns(
        chromosome="chrom",
        position="pos",
        reference_allele="ref",
        alternate_allele="alt",
    ),
)
source = ConservationSource(backend)

The same source can receive a PostgreSQL, BigQuery, or service implementation of AllelicRecordBackend. Conversely, the same Parquet adapter serves AlphaMissense, SpliceAI, and other allelic record collections. This is the intended M scientific sources + N storage adapters composition; do not add SourceNamePostgresBackend, SourceNameBigQueryBackend, and similar cross-product classes.

Altar also ships SqliteAllelicRecordBackend(database, table=..., key_columns=...) and BigQueryAllelicRecordBackend(client, table_id, key_columns=..., filters=...) against the identical contract. Parquet, SQLite, and BigQuery are conformance-tested with projection, request scoping, missing records, and one-to-many matches.

BigQueryScoreStore needs SQL to join a source into its materialize and export queries. For a table that only needs columns selected, renamed, or aggregated, configure a BigQueryAnnotationSource with a relation= subquery. For a sparse source whose binding computes its annotations in Python, open staged_annotation_source with the job's variants and pass the source it yields to the store inside the async with block. Pass the BigQueryAllelicRecordBackend as coverage= so that only variants with records are annotated. See backends for when staging fits.

For an interval-overlap relation, do not define another source-specific backend. Normalize rows into VariantGeneEvidenceRecord, persist them through VariantGeneLinkStore, and expose them through StoredVariantGeneLinkSource. The in-memory and SQLite stores share generation visibility, filtering, and seek-pagination semantics; external PostgreSQL, Parquet, BigQuery, and service implementations run the same VariantGeneLinkStoreContract.

5. Register and verify

A distributed binding registers under altar.annotation_sources so hosts can look it up by name. Registration is optional: materialization accepts any AnnotationSource instance.

Subclass AnnotationSourceContract and return the published contract from its contract fixture. Add tests for missing rows, zero values, malformed backend responses, repeated detail, rule behavior, and registry discovery.

Use VariantGeneLinkSource instead when one variant can produce multiple gene relations with link-specific context and provenance.