Add an annotation source¶
Use AnnotationSource for precomputed per-variant evidence.
1. Declare logical columns¶
from altar.sources import AnnotationColumn, AnnotationSource
class ConservationSource(AnnotationSource):
name = "conservation"
def annotation_columns(self):
return [AnnotationColumn("phylo_score", "float", "Phylogenetic score")]
async def annotate(self, variant_ids):
rows = await self.backend.fetch(variant_ids)
return {row.variant_id: {"phylo_score": row.score} for row in rows}
Omit variants with no source row. Return an observed zero as zero. Validate that a backend never returns an unrequested variant or duplicate logical record.
Repeated and structured columns should be declared as such even if a simple physical backend serializes them as JSON.
2. Publish a contract and declare the identity¶
Publish the columns as an AnnotationContract constant so other sources, including a host's own tables, can
implement the same contract. Return the contract's identity from identity(), with the data release and
genome build the source serves.
from altar.sources import AnnotationColumn, AnnotationContract
CONSERVATION_ANNOTATIONS = AnnotationContract(
source_id="org.example.annotation.conservation",
schema_version="1.0.0",
columns=(AnnotationColumn("phylo_score", "float", "Phylogenetic score"),),
)
class ConservationSource(AnnotationSource):
...
def annotation_columns(self):
return list(CONSERVATION_ANNOTATIONS.columns)
def identity(self):
return CONSERVATION_ANNOTATIONS.identity(release=self.release, genome_build=self.genome_build)
A model's AnnotationDependency names a contract by source_id. Materialization resolves each dependency
against the configured sources' identities; see
Use your own annotation tables. Adding columns is a minor
version. Removing or redefining a column is a major version.
3. Add a rule only when justified¶
Return None from prioritize_predicate() for passive evidence. If the source has a documented policy, return
a predicate from altar.predicates and expose configurable thresholds where appropriate.
from altar.predicates import Col, Ge, Lit
def prioritize_predicate(self):
return Ge(Col("phylo_score"), Lit(self.threshold))
4. Separate science from storage¶
The source should normalize scientific meaning. For source-native tables keyed by chromosome, position,
reference, and alternate allele, accept the shared AllelicRecordBackend rather than defining a new backend
protocol and a new Parquet query class for every source. The binding creates an AllelicRecordQuery, decodes
the projected native fields, and owns all aggregation. The backend only retrieves records.
from altar.sources import AllelicKeyColumns, ParquetAllelicRecordBackend
backend = ParquetAllelicRecordBackend(
"./conservation-records",
key_columns=AllelicKeyColumns(
chromosome="chrom",
position="pos",
reference_allele="ref",
alternate_allele="alt",
),
)
source = ConservationSource(backend)
The same source can receive a PostgreSQL, BigQuery, or service implementation of AllelicRecordBackend.
Conversely, the same Parquet adapter serves AlphaMissense, SpliceAI, and other allelic record collections.
This is the intended M scientific sources + N storage adapters composition; do not add
SourceNamePostgresBackend, SourceNameBigQueryBackend, and similar cross-product classes.
Altar also ships SqliteAllelicRecordBackend(database, table=..., key_columns=...) and
BigQueryAllelicRecordBackend(client, table_id, key_columns=..., filters=...) against the identical contract.
Parquet, SQLite, and BigQuery are conformance-tested with projection, request scoping, missing records, and
one-to-many matches.
BigQueryScoreStore needs SQL to join a source into its materialize and export queries. For a table that only
needs columns selected, renamed, or aggregated, configure a BigQueryAnnotationSource with a relation=
subquery. For a sparse source whose binding computes its annotations in Python, open staged_annotation_source
with the job's variants and pass the source it yields to the store inside the async with block. Pass the
BigQueryAllelicRecordBackend as coverage= so that only variants with records are annotated. See
backends for when staging fits.
For an interval-overlap relation, do not define another source-specific backend. Normalize rows into
VariantGeneEvidenceRecord, persist them through VariantGeneLinkStore, and expose them through
StoredVariantGeneLinkSource. The in-memory and SQLite stores share generation visibility, filtering, and
seek-pagination semantics; external PostgreSQL, Parquet, BigQuery, and service implementations run the same
VariantGeneLinkStoreContract.
5. Register and verify¶
A distributed binding registers under altar.annotation_sources so hosts can look it up by name. Registration
is optional: materialization accepts any AnnotationSource instance.
Subclass AnnotationSourceContract and return the published contract from its contract fixture. Add tests
for missing rows, zero values, malformed backend responses, repeated detail, rule behavior, and registry
discovery.
Use VariantGeneLinkSource instead when one variant can produce multiple gene relations with link-specific
context and provenance.