Store and assemble results¶
A score store persists attributable rows and assembles them with annotations into MaterializedVariant
records.
Rich repeated results use a separate DetailStore. A model binding names and versions each DetailSchema;
the store persists those rows with their parent variant_id, model identity, complete plugin run identity,
and schema provenance. This keeps the scalar materialization contract stable while retaining track-, gene-,
or event-level observations losslessly.
The boundary follows ownership rather than a fixed cardinality cutoff: repeated output from a model run is a detail result, while an externally produced release is an annotation source. Reusable normalized variant–gene evidence uses the separate link contract. See Choose the result shape.
from altar.results import BigQueryDetailStore, SqliteDetailStore
details = SqliteDetailStore("details.sqlite")
# The injected client is a google.cloud.bigquery.Client.
warehouse_details = BigQueryDetailStore(client, "project.dataset.detail_rows")
async for row in details.read_rows(model_id="alphagenome-liver", detail_name="track_scores"):
consume_track(row)
Declare the schema¶
Every model binding returns a finite list of ScoreColumn definitions. Before inserting rows, pass the union
of the participating schemas to ensure_schema().
Reusing a column name is valid only when the complete definition agrees. This allows ChromBPNet and
Cherimoya to share semantically identical fields such as logfc while rejecting accidental collisions.
Write scores¶
await store.add_scores(
model_id="k562-cbp",
model_name="ChromBPNet K562",
plugin_identity=plan.plugin_identity,
rows=rows,
columns=plugin.score_columns(),
)
Rows remain attributable through model_id, model_name, and the exact plugin_identity that defined their
schema, reducer, prioritization expression, runtime, model configuration, output selection, and genome build.
Compact-lineage stores record full metadata in a catalog and persist its compact ID on each row. They can keep
multiple scoring versions under the same model ID and apply an explicit ScoreReusePolicy during cache
lookup and assembly. See Reuse scores across updates for registration, single-column
backfills, and dense miss preparation. The original store mode continues to require scientifically compatible
identities across each model's cache.
Provide the intended current identity and reuse policy to ScoredModel when reading results. Use
get_missing_scores() or prepare_scoring_batches() before executing work; backend write idempotency is not
uniform. The unscoped get_unscored_variants() API is available only in the original store mode.
Assemble variant results¶
async for result in store.materialize(
variant_ids=variant_ids,
models=models,
annotation_sources=sources,
):
consume(result)
Assembly first checks that the annotation sources supply every column the models' rules depend on
(resolve_annotation_dependencies). It then loads annotations, reads each model's scores, evaluates declared
predicates, and emits one result at a time. The streaming interface avoids holding a complete large result set in memory.
Every store yields exactly one result per distinct requested variant, so a repeated ID appears once. Results
come in ascending variant_id order, compared as plain strings, so chr10:… sorts before chr2:…. Request
order is not kept, because a warehouse store streams its rows sorted by variant. A variant that no selected
model scored still appears, with empty model_scores; it is prioritized only when an annotation source's rule
matches it. Missing scores therefore stay visible as missing rather than disappearing from the result.
Each result's annotations maps the declared annotation columns to their values for that variant, merged
across every configured source, so a caller does not need a separate fetch_annotations call. Every store
fills it the same way:
- A column with no value is absent, never
None. A source with no row for the variant, a null cell, and an empty repeated value all count as no value, because BigQuery returns a null array as an empty one. A variant no source annotated gets{}. - Values keep their logical type. A repeated column is a list and a struct column is a dict, or a list of dicts
when repeated, even on SQLite, which stores those columns as JSON text and decodes them on read. On BigQuery a
repeated or struct column must physically be an
ARRAYorSTRUCT. A scalarjsoncolumn may be aJSONvalue or JSON text, such as arelation=that projectsTO_JSON_STRING(...): BigQuery decodes the text, so the value is the same as on SQLite, and'[]'is absent like[].
async for result in store.materialize(variant_ids=variant_ids, models=models, annotation_sources=sources):
region = result.annotations.get("region_type") # None when no source has a value for this variant
The portable driver already holds every requested variant's annotations while it streams. BigQueryScoreStore
reads the values of the sources it joins from the query rows it already returns, and calls annotate() once for
the other sources. Requested variants outside its query (no score and no prioritizing-source row) read the joined
passive sources again. With variants_relation and no inline variant_ids, a source that cannot join into the
query has no ids to annotate, so the store raises ValueError; stage it, leave it out, or also pass the ids.
SQLite and BigQuery¶
SqliteScoreStore and SqliteDetailStore are the local reference implementations. The former performs
portable reads and result assembly in Python. BigQueryScoreStore is intended for warehouse-scale use and
can push compatible joins and predicate evaluation into SQL. BigQueryDetailStore supplies the independent
warehouse detail axis with fixed metadata and canonical JSON-row tables.
Adapters on each axis implement the same public contract. Physical types and write behavior can differ, so portable consumers should rely on the logical schema and conformance contract rather than backend internals.