Skip to content

Execution and storage integrations

Infrastructure implementations consume the same public contracts as external extensions.

Execution

Name Implementation Intended use
local LocalDockerExecutionBackend local development and single-host runs
modal ModalExecutionBackend managed elastic container execution
kubernetes KubernetesExecutionBackend existing clusters with deployment-owned storage and policy

Model bindings emit the same ContainerTaskSpec regardless of backend. Provider-specific image handling, staging, polling, and capacity reporting stay in the backend.

A ContainerTaskSpec with timeout_s=None asks for no time limit. Local Docker and Kubernetes run it without one. Modal sandboxes live at most 24 hours, so the Modal backend requests that maximum instead of the SDK's 5-minute default, and rejects an explicit timeout_s longer than 24 hours rather than shortening it.

Score storage

Name Implementation Intended use
sqlite SqliteScoreStore local analyses, tests, and reference deployments
bigquery BigQueryScoreStore warehouse-scale persistence and SQL result assembly

SQLite deduplicates its primary key on insert. BigQuery's append path is not intrinsically idempotent; callers must use get_unscored_variants() before scoring rather than relying on repeated inserts.

Named one-to-many model results use the independent altar.detail_stores axis. Altar ships SqliteDetailStore for local persistence and BigQueryDetailStore for warehouse persistence. Both accept every binding-declared DetailSchema, so adding another detail producer does not require another storage integration. The BigQuery adapter keeps complete schemas, plugin run identities, and rows as canonical JSON behind fixed metadata/key columns. New logical fields therefore do not trigger physical schema migrations; staged MERGE writes make repeated logical keys idempotent, including keys containing null values.

Annotation and relation storage

SQLite and BigQuery annotation sources provide writable reference implementations. Scientific bindings such as AlphaMissense, SpliceAI, and GPN-Star inject the same AllelicRecordBackend; Altar provides generic Parquet, SQLite, and BigQuery implementations. GPN-Star's binding-local asset verifier checks its published shard identity independently of lookup, rather than combining one scientific source with one storage engine.

A writable BigQueryAnnotationSource stores each row with its genome and chr/pos/ref/alt beside the canonical variant_id, and get_unannotated_variants only counts rows of the source's genome_default build. A table written before those columns existed gains them, empty, from ensure_schema. Pause ingest into the table, then run backfill_identity_columns() once before the next deduplication, so its earlier rows are stamped with genome_default rather than annotated a second time. It runs as one BigQuery transaction. Only rows with a canonical variant_id are stamped, one per variant; a legacy duplicate, or a legacy row whose variant already has a row of that build, is deleted. It returns the stamped and deleted counts.

A writable source constructed with identity= and identity_scoped=True keeps each data release's rows apart. On BigQuery it writes the identity's hash into an _altar_annotation_identity column, and deduplication and reads use only rows with that hash, so a new data release annotates every variant again. Rows written before the column existed are not used until backfill_annotation_identity() adopts them under the release that built them. That backfill rewrites the table's storage in one transaction. A scoped SQLite file records one identity and refuses to open under another. Without identity_scoped, identity= is metadata only. See Keep a writable lake current across releases for the migration steps and their cost.

BigQueryScoreStore joins annotation sources into its materialize and export queries in SQL. On BigQuery the default is a BigQueryAnnotationSource whose relation= subquery selects, renames, or aggregates the reference table's columns. The whole materialization then stays in one BigQuery query.

BigQuery cannot export nested or repeated data as CSV. The flat TSV download from export_download therefore writes each annotation column declared repeated or struct, such as am_transcript_scores or nearest_genes, as JSON text, and leaves a missing value empty. A column declared repeated or struct must physically be an ARRAY or STRUCT, not a pre-serialized string. A scalar json column, including one a relation= projects with TO_JSON_STRING(...), is exported unchanged. The NDJSON and Parquet exports keep native types. The declared AnnotationColumn decides the encoding, not the column's name.

materialize results carry each declared annotation column as its logical value. A scalar json column that arrives as JSON text, such as a TO_JSON_STRING(...) projection, is decoded, so it equals the value a SQLite source returns. See Store and assemble results.

staged_annotation_source is for the exception: a source whose annotations are computed in Python. Use it only when all three hold:

  1. A binding's Python code is the definition of the annotations, so a SQL relation would be a second copy.
  2. The logic is more than column selection and renaming, for example picking the most damaging of several transcript records.
  3. The source is sparse: only a small share of a job's variants have records, as with missense predictions.

It runs the source for one job's variants and loads the rows into an expiring scratch table. Inside its async with block it yields a source bound to that table, which the store joins like a SQL source, and the table is dropped when the block exits. With coverage= set to a BigQueryAllelicRecordBackend, one query finds the job's variants that have records and reads those records. Only those variants go through Python, and the reference table is read once per job. A dense source staged this way sends most of a job's variants through Python and back; keep such a source as a SQL relation.

VariantGeneLinkStore persists canonical one-to-many relation evidence independently of its upstream format. The in-memory and SQLite reference implementations support idempotent staged generations, atomic publication, indexed filters, and stable seek pagination. ENCODE-rE2G, scE2G, and Open Targets E2G all use these stores without source-specific database classes.

Object storage and authentication

Core ships local filesystem storage and a development authentication provider. Production object storage, authorization, tenancy, and credential policy belong to the hosting application or external extensions. The AuthProvider and JobQueue contracts are provisional until a public production adapter and an Altar consumer exist.