Skip to content

Scores, annotations, and links

These data types can coexist in one analysis, but they are not interchangeable.

Model scores

A score belongs to one model instance. The instance supplies biological scope—such as a cell type, tissue, fold set, ontology filter, checkpoint, or artifact version—while the model binding supplies the score schema and interpretation.

Altar binds persisted scores to that scope. The run identity hashes both the complete validated model configuration and its credential-free scientific projection, plus the complete output selection, and records the readable genome build alongside an optional content-addressed genome resource. SHA-256 resource digests drive scientific compatibility independently of storage location; immutable URI+revision is the fallback. A single model_id therefore cannot silently mix checkpoint or FASTA bytes, tissue/ontology filters, output modes, or genome builds. Credential references remain exact provenance but do not change result compatibility. Variant subjects and execution batches are excluded so incremental scoring remains possible.

The run identity records one plugin release and the concrete computation, not downstream prioritization. Changing a prioritization rule alone permits reuse of the same scores. Strict stores compare a role-specific compatibility fingerprint, including the plugin version and exact runtime. Compact-lineage stores can retain multiple releases and reuse explicitly accepted predecessors through ScoreReusePolicy; neither package version numbers nor image republication automatically establish numerical equivalence.

Two model architectures may intentionally use the same semantic column name. Altar can keep those values comparable while preventing collisions by retaining model and schema identity, not by inventing different biological names. Conversely, related models should keep different names when their quantities are not actually identical—for example, current ChromBPNet logfc and Cherimoya counts_log2fc have distinct declared schemas.

Annotations

An annotation is looked up from an existing dataset or service. It may be an observed fact, a published prediction, or a summary over several source records. Examples include precomputed SpliceAI, AlphaMissense, and GPN-Star scores.

Annotations may contribute a prioritization rule, but they do not acquire a model-run identity merely because their upstream dataset was produced by a model.

A variant–gene link is a one-to-many relation, not a scalar annotation. It can preserve:

  • the query variant and linked regulatory element;
  • target gene identifiers;
  • biosample or tissue context;
  • link scores and score direction;
  • genomic distances;
  • method, dataset release, genome build, and source record provenance.

Altar exchanges these typed relations without constructing a knowledge graph or asserting causality.

Variant results

MaterializedVariant is the shared per-variant summary. It contains attributable model_scores, the versioned plugin/run identity for each model, an overall prioritized flag, the exact model or annotation sources whose rules matched, and the variant's annotations keyed by declared annotation column. Rich one-to-many outputs, tracks, matrices, and relations stay in named detail or artifact outputs rather than being flattened into the scalar score table.

Missing is not zero

No row can mean the dataset does not cover the locus, the allele is unsupported, the selected model could not score it, or an upstream artifact was absent. A numeric zero is an observed value. Altar therefore left-merges missing evidence as missing and does not zero-fill it.