Skip to content

Store and assemble results

A score store persists attributable rows and assembles them with annotations into MaterializedVariant records.

Rich repeated results use a separate DetailStore. A model binding names and versions each DetailSchema; the store persists those rows with their parent variant_id, model identity, complete plugin run identity, and schema provenance. This keeps the scalar materialization contract stable while retaining track-, gene-, or event-level observations losslessly.

The boundary follows ownership rather than a fixed cardinality cutoff: repeated output from a model run is a detail result, while an externally produced release is an annotation source. Reusable normalized variant–gene evidence uses the separate link contract. See Choose the result shape.

from altar.results import BigQueryDetailStore, SqliteDetailStore

details = SqliteDetailStore("details.sqlite")
# The injected client is a google.cloud.bigquery.Client.
warehouse_details = BigQueryDetailStore(client, "project.dataset.detail_rows")
async for row in details.read_rows(model_id="alphagenome-liver", detail_name="track_scores"):
    consume_track(row)

Declare the schema

Every model binding returns a finite list of ScoreColumn definitions. Before inserting rows, pass the union of the participating schemas to ensure_schema().

Reusing a column name is valid only when the complete definition agrees. This allows ChromBPNet and Cherimoya to share semantically identical fields such as logfc while rejecting accidental collisions.

Write scores

await store.add_scores(
    model_id="k562-cbp",
    model_name="ChromBPNet K562",
    plugin_identity=plan.plugin_identity,
    rows=rows,
    columns=plugin.score_columns(),
)

Rows remain attributable through model_id, model_name, and the exact plugin_identity that defined their schema, reducer, prioritization expression, runtime, model configuration, output selection, and genome build. Compact-lineage stores record full metadata in a catalog and persist its compact ID on each row. They can keep multiple scoring versions under the same model ID and apply an explicit ScoreReusePolicy during cache lookup and assembly. See Reuse scores across updates for registration, single-column backfills, and dense miss preparation. The original store mode continues to require scientifically compatible identities across each model's cache.

Provide the intended current identity and reuse policy to ScoredModel when reading results. Use get_missing_scores() or prepare_scoring_batches() before executing work; backend write idempotency is not uniform. The unscoped get_unscored_variants() API is available only in the original store mode.

Assemble variant results

async for result in store.materialize(
    variant_ids=variant_ids,
    models=models,
    annotation_sources=sources,
):
    consume(result)

Assembly first checks that the annotation sources supply every column the models' rules depend on (resolve_annotation_dependencies). It then loads annotations, reads each model's scores, evaluates declared predicates, and emits one result at a time. The streaming interface avoids holding a complete large result set in memory.

Every store yields exactly one result per distinct requested variant, so a repeated ID appears once. Results come in ascending variant_id order, compared as plain strings, so chr10:… sorts before chr2:…. Request order is not kept, because a warehouse store streams its rows sorted by variant. A variant that no selected model scored still appears, with empty model_scores; it is prioritized only when an annotation source's rule matches it. Missing scores therefore stay visible as missing rather than disappearing from the result.

Each result's annotations maps the declared annotation columns to their values for that variant, merged across every configured source, so a caller does not need a separate fetch_annotations call. Every store fills it the same way:

  • A column with no value is absent, never None. A source with no row for the variant, a null cell, and an empty repeated value all count as no value, because BigQuery returns a null array as an empty one. A variant no source annotated gets {}.
  • Values keep their logical type. A repeated column is a list and a struct column is a dict, or a list of dicts when repeated, even on SQLite, which stores those columns as JSON text and decodes them on read. On BigQuery a repeated or struct column must physically be an ARRAY or STRUCT. A scalar json column may be a JSON value or JSON text, such as a relation= that projects TO_JSON_STRING(...): BigQuery decodes the text, so the value is the same as on SQLite, and '[]' is absent like [].
async for result in store.materialize(variant_ids=variant_ids, models=models, annotation_sources=sources):
    region = result.annotations.get("region_type")  # None when no source has a value for this variant

The portable driver already holds every requested variant's annotations while it streams. BigQueryScoreStore reads the values of the sources it joins from the query rows it already returns, and calls annotate() once for the other sources. Requested variants outside its query (no score and no prioritizing-source row) read the joined passive sources again. With variants_relation and no inline variant_ids, a source that cannot join into the query has no ids to annotate, so the store raises ValueError; stage it, leave it out, or also pass the ids.

SQLite and BigQuery

SqliteScoreStore and SqliteDetailStore are the local reference implementations. The former performs portable reads and result assembly in Python. BigQueryScoreStore is intended for warehouse-scale use and can push compatible joins and predicate evaluation into SQL. BigQueryDetailStore supplies the independent warehouse detail axis with fixed metadata and canonical JSON-row tables.

Adapters on each axis implement the same public contract. Physical types and write behavior can differ, so portable consumers should rely on the logical schema and conformance contract rather than backend internals.