Skip to content

Reuse scores across model updates

Use ScoreReusePolicy when an analysis may reuse explicitly accepted older scores for the same registered model. For example, a publisher may approve results from a previous model release, or an application may retain a historical scoring method while adopting a new one. Altar does not infer scientific equivalence from semantic version numbers.

The row key separates four concerns: model_id, input genome, canonical variant_id, and score_lineage_id. A model's training species is independent of the input genome. One imported lineage can describe a shared scoring method across many models and genomes; their row keys continue to distinguish them. Changing which trained model a model ID denotes remains the registry owner's responsibility.

Configure compact identity storage

from altar.results import BigQueryScoreStore, ScoreLineage, ScoreReusePolicy

store = BigQueryScoreStore(
    client,
    "project.dataset.variant_scores",
    lineage_table_id="project.dataset.score_lineages",
)
await store.ensure_schema(plugin.score_columns())

The canonical score table holds one additional string column, score_lineage_id. Full metadata is recorded once per lineage in the small catalog table, including exact run metadata for new Altar computations. A provenance table beside it, {lineage_table_id}__provenance, records each distinct exact run identity that wrote scores under a lineage (see what decides a lineage). This mode does not add or write per-row plugin_identity_json. Existing constructor calls retain the previous strict, single-compatible-identity behavior; enabling compact lineages is explicit.

For a new local store, use SqliteScoreStore(path, lineages=True). It keeps the catalog in _altar_score_lineages and exact provenance in _altar_score_lineage_provenance. Existing SQLite files using the old primary key require an explicit schema migration before enabling this mode; Altar does not rebuild them implicitly.

Register existing scores

historical = ScoreLineage(
    lineage_id="release-2025",
    plugin_id=plugin.manifest.plugin_id,
    columns=tuple(plugin.score_columns()),  # declare the outputs actually present
    metadata_json='{"description":"Publisher-approved historical scoring method"}',
)
await store.register_score_lineage(historical)

The catalog has no architecture-specific columns. Optional model settings, release labels, provenance, or operator notes live inside its JSON metadata. Imported scores do not need an invented Altar run identity. Registration is immutable and idempotent; changing an existing declaration requires a new ID.

An application can then backfill its existing table with that single ID. When all historical rows share the declared method, the deployment-owned operation is simply:

UPDATE `project.dataset.variant_scores`
SET score_lineage_id = 'release-2025'
WHERE score_lineage_id IS NULL;

Scope the update more narrowly if the historical table contains different scoring methods. ensure_schema only adds the nullable column to an existing BigQuery table; it never performs this backfill. The update is a separate DML operation that can process substantial data and should be planned by the table owner. It changes metadata, not score values. There is no requirement to regenerate historical scores or move them into a separate table. If migrating an older Altar table whose plugin_identity_json is REQUIRED, first relax that old column to NULLABLE so new compact rows can omit it.

Rows with NULL or unaccepted lineage IDs are cache misses. There is no implicit legacy mapping.

What decides a generated lineage

A lineage ID computed by Altar hashes the run's scientific projection, PluginRunIdentity.scientific_dict(): the plugin release, its exact runtimes, the scientific configuration projection, the output selection, the genome build and genome hash, and the primary result schema. That projection also defines the lineage. A store accepts a run under an existing generated lineage when the plugin, the declared output columns, and the projection agree, even if the run's exact identity differs. So these runs reuse each other's scores and write under the same lineage:

  • content-addressed resources served from a new location (same digest, different URI);
  • a configuration-schema migration, or a score-neutral serialization field such as ChromBPNet's peaks_compression, which change only the exact configuration_hash;
  • a later Altar release, which changes altar_api_version or the plugin's supported_altar_api.

A different scientific configuration, plugin release, or runtime image is a different lineage and a cache miss. A republished runtime image is a new computation; accept its predecessor explicitly with ScoreReusePolicy. An imported lineage (no run_identity) matches only an identical declaration, and a generated ID can never be rebound to an imported declaration.

Exact provenance is kept as a set per lineage, not per row. The catalog's run_identity is one deterministic representative of the lineage. Usually that is the first run to register it, but not always: in BigQuery, concurrent first registrations keep the smallest declaration, and a process that registered its own declaration may keep using it until it restarts. StoredScore.lineage and materialized plugin_identities carry that reference. A score row is attributable to its lineage, not to the exact identity that wrote it. Each run that registers to write under a lineage records its exact identity once, keyed by its deterministic hash, in the lineage provenance table. That includes a run whose rows were all skipped as duplicates. Read-only lookups such as get_missing_scores record nothing. Read the set of exact identities, oldest record first:

lineage_id = ScoreLineage.from_run(plan.plugin_identity, plugin.score_columns()).lineage_id
exact_runs = await store.read_lineage_provenance(lineage_id)  # tuple[PluginRunIdentity, ...]

The result always includes the catalog reference, including for a catalog written before the provenance table existed. An imported lineage, or an ID the store never registered, returns an empty tuple.

Choose acceptable results

reuse = ScoreReusePolicy(accepted_lineage_ids=("release-2025",))

missing = await store.get_missing_scores(
    candidate_ids,
    model_id,
    plugin_identity=plan.plugin_identity,
    columns=plugin.score_columns(),
    reuse_policy=reuse,
)

The current run's lineage is always accepted and preferred. Other accepted IDs are tried in declaration order. Acceptance is request-local, explicit, and not transitive. An empty policy accepts only the current run in compact-lineage mode. ScoreLineage.from_run(plan.plugin_identity, plugin.score_columns()).lineage_id gives the automatically generated ID when registering a current run as an accepted predecessor of a future one. These generated IDs describe a run's scientific projection, including its runtime and plugin release; plugin/image upgrades therefore require explicit acceptance to reuse their predecessors. The plugin has one plugin_version, not separate wrapper and scoring-semantics versions. Prioritization policy, expression, annotation dependencies, and the whole-manifest hash are absent from run identity; changing downstream selection alone preserves the lineage and permits reprioritization without inference.

All candidates remain constrained to the requested model and input genome. Accepted lineages must belong to the requested plugin and supply compatible output types. A declared nullable output may legitimately contain NULL; that is distinct from an output missing from the lineage's schema. Requesting a subset of columns from the same run is a projection and does not generate a new identity. An output selection that changes inference or reduction remains part of run identity; this API does not automatically equate different computations.

read_selected_scores takes the same arguments and returns StoredScore objects with both values and the actual lineage. Selection takes an entire row rather than combining individual columns from different versions. BigQuery resolves physical duplicates within a lineage by newest created_at, then a stable whole-row tie-breaker; it does not claim two differing duplicate computations are scientifically identical.

To apply the same policy during assembly, set ScoredModel(..., reuse_policy=reuse). Materialized results include score_lineages per model. plugin_identities contains actual recorded run metadata only; an imported lineage without an Altar run does not acquire fabricated provenance. Compact-mode warehouse exports and staged materialization tables also retain lineage IDs.

Filter the cohort before batching

Pass reuse_policy=reuse to prepare_scoring_batches. It filters the complete cohort, then packs the miss set into executable requests. For 1,000 unique candidates and 601 acceptable cached results, a batch limit of 100 produces 100, 100, 100, and 99 variants. The policy is retained on the returned preparation summary.

Result ingestion always writes the current computation's own identity. Its final deduplication checks that identity, so an older row cannot prevent a deliberately recomputed result from being stored alongside it.

For warehouse-scale cohorts, keep the candidates and misses in BigQuery:

missing_count = await store.prepare_unscored_table(
    dest_table="`project.dataset.job_missing`",
    variants_relation="`project.dataset.job_candidates`",
    model_id=model_id,
    plugin_identity=plan.plugin_identity,
    columns=plugin.score_columns(),
    reuse_policy=reuse,
    max_variants_per_batch=100,
)

for batch in range((missing_count + 99) // 100):
    await store.export_prepared_batch(
        prepared_relation="`project.dataset.job_missing`",
        scoring_batch=batch,
        destination_uri=f"gs://example/job/batch-{batch}/part-*.tsv",
    )

The candidate relation must contain unique canonical loci and variant_id, scoped to the request's genome. Preparation filters accepted cache entries and numbers the global miss set in one CTAS. Batch exports read only that job table, whose default expiry is one day. The caller owns its relation names, lifetime, transfers, and execution requests. Concurrent jobs may prepare overlapping misses; this is a snapshot, not a distributed work reservation. Scores computed after preparation do not change its batch boundaries.

get_missing_scores(..., variants_relation=...) supports a warehouse relation when the returned miss IDs are small enough to fetch. export_unscored supports the same policy and the older frozen-cutoff interface; its compact mode uses dense ranks rather than hash buckets, but recomputes the miss query on each call. Prefer the prepared-table path for many execution batches.

Scope and validation

This contract covers primary scalar scores. Named detail tables have no lineage column: a detail table is bound to one scientific result identity per model ID, so rows from an accepted older lineage cannot be told apart from rows of the current run. prepare_scoring_batches therefore rejects a non-empty ScoreReusePolicy for a binding that declares detail tables (CapabilityNotSupported). With an empty policy it accepts such a binding and counts a variant as cached only when its current-run primary row and a row in every detail table are present; see bindings with named details. Invalid-result retry policy also remains a separate contract: lineage acceptance alone does not distinguish a failed biological prediction from a completed valid score.

SQLite behavior and translated warehouse SELECT semantics are exercised locally, including mixed versions, model/genome isolation, whole-row preference, and dense batching. The shared score-store conformance suite runs the lineage acceptance and provenance rules against SQLite and against the BigQuery adapter's SQL on DuckDB. Recording-client tests cover BigQuery DDL, catalog registration, dataframe/staging writes and export composition. The opt-in real BigQuery test must pass in a scratch dataset before a production cutover; local SQL execution does not establish warehouse cost or performance on a multi-terabyte score table.