Execution and storage integrations¶
Infrastructure implementations consume the same public contracts as external extensions.
Execution¶
| Name | Implementation | Intended use |
|---|---|---|
local |
LocalDockerExecutionBackend |
local development and single-host runs |
modal |
ModalExecutionBackend |
managed elastic container execution |
kubernetes |
KubernetesExecutionBackend |
existing clusters with deployment-owned storage and policy |
Model bindings emit the same ContainerTaskSpec regardless of backend. Provider-specific image handling,
staging, polling, and capacity reporting stay in the backend.
A ContainerTaskSpec with timeout_s=None asks for no time limit. Local Docker and Kubernetes run it without
one. Modal sandboxes live at most 24 hours, so the Modal backend requests that maximum instead of the SDK's
5-minute default, and rejects an explicit timeout_s longer than 24 hours rather than shortening it.
Score storage¶
| Name | Implementation | Intended use |
|---|---|---|
sqlite |
SqliteScoreStore |
local analyses, tests, and reference deployments |
bigquery |
BigQueryScoreStore |
warehouse-scale persistence and SQL result assembly |
SQLite deduplicates its primary key on insert. BigQuery's append path is not intrinsically idempotent; callers
must use get_unscored_variants() before scoring rather than relying on repeated inserts.
Named one-to-many model results use the independent altar.detail_stores axis. Altar ships
SqliteDetailStore for local persistence and BigQueryDetailStore for warehouse persistence. Both accept
every binding-declared DetailSchema, so adding another detail producer does not require another storage
integration. The BigQuery adapter keeps complete schemas, plugin run identities, and rows as canonical JSON
behind fixed metadata/key columns. New logical fields therefore do not trigger physical schema migrations;
staged MERGE writes make repeated logical keys idempotent, including keys containing null values.
Annotation and relation storage¶
SQLite and BigQuery annotation sources provide writable reference implementations. Scientific bindings such
as AlphaMissense, SpliceAI, and GPN-Star inject the same AllelicRecordBackend; Altar provides generic
Parquet, SQLite, and BigQuery implementations. GPN-Star's binding-local asset verifier checks its published
shard identity independently of lookup, rather than combining one scientific source with one storage engine.
A writable BigQueryAnnotationSource stores each row with its genome and chr/pos/ref/alt beside the
canonical variant_id, and get_unannotated_variants only counts rows of the source's genome_default build.
A table written before those columns existed gains them, empty, from ensure_schema. Pause ingest into the
table, then run backfill_identity_columns() once before the next deduplication, so its earlier rows are
stamped with genome_default rather than annotated a second time. It runs as one BigQuery transaction. Only rows
with a canonical variant_id are stamped, one per variant; a legacy duplicate, or a legacy row whose variant
already has a row of that build, is deleted. It returns the stamped and deleted counts.
A writable source constructed with identity= and identity_scoped=True keeps each data release's rows apart.
On BigQuery it writes the identity's hash into an _altar_annotation_identity column, and deduplication and
reads use only rows with that hash, so a new data release annotates every variant again. Rows written before
the column existed are not used until backfill_annotation_identity() adopts them under the release that
built them. That backfill rewrites the table's storage in one transaction. A scoped SQLite file records one
identity and refuses to open under another. Without identity_scoped, identity= is metadata only. See
Keep a writable lake current across releases
for the migration steps and their cost.
BigQueryScoreStore joins annotation sources into its materialize and export queries in SQL. On BigQuery the
default is a BigQueryAnnotationSource whose relation= subquery selects, renames, or aggregates the reference
table's columns. The whole materialization then stays in one BigQuery query.
BigQuery cannot export nested or repeated data as CSV. The flat TSV download from export_download therefore
writes each annotation column declared repeated or struct, such as am_transcript_scores or
nearest_genes, as JSON text, and leaves a missing value empty. A column declared repeated or struct must
physically be an ARRAY or STRUCT, not a pre-serialized string. A scalar json column, including one a
relation= projects with TO_JSON_STRING(...), is exported unchanged. The NDJSON and Parquet exports keep
native types. The declared AnnotationColumn decides the encoding, not the column's name.
materialize results carry each declared annotation column as its logical value. A scalar json column that
arrives as JSON text, such as a TO_JSON_STRING(...) projection, is decoded, so it equals the value a SQLite
source returns. See Store and assemble results.
staged_annotation_source is for the exception: a source whose annotations are computed in Python. Use it only
when all three hold:
- A binding's Python code is the definition of the annotations, so a SQL
relationwould be a second copy. - The logic is more than column selection and renaming, for example picking the most damaging of several transcript records.
- The source is sparse: only a small share of a job's variants have records, as with missense predictions.
It runs the source for one job's variants and loads the rows into an expiring scratch table. Inside its
async with block it yields a source bound to that table, which the store joins like a SQL source, and the
table is dropped when the block exits. With coverage= set to a BigQueryAllelicRecordBackend, one query
finds the job's variants that have records and reads those records. Only those variants go through Python,
and the reference table is read once per job. A dense source staged this way sends most of a job's variants
through Python and back; keep such a source as a SQL relation.
VariantGeneLinkStore persists canonical one-to-many relation evidence independently of its upstream format.
The in-memory and SQLite reference implementations support idempotent staged generations, atomic publication,
indexed filters, and stable seek pagination. ENCODE-rE2G, scE2G, and Open Targets E2G all use these stores
without source-specific database classes.
Object storage and authentication¶
Core ships local filesystem storage and a development authentication provider. Production object storage,
authorization, tenancy, and credential policy belong to the hosting application or external extensions. The
AuthProvider and JobQueue contracts are provisional until a public
production adapter and an Altar consumer exist.