Choose the result shape¶
Altar separates results by meaning and lifecycle, not merely by row count. Cardinality still affects the physical adapter a deployment chooses, but it does not decide the public contract by itself.
Decision¶
Use these three contracts:
| Result | Contract | Owner and lifecycle |
|---|---|---|
| Imported or remotely looked-up evidence | AnnotationSource |
An external dataset or service release; no Altar model run occurred |
| Repeated output produced by one model run | DetailSchema and DetailStore |
The model instance and exact PluginRunIdentity that produced it |
| Reusable variant-to-gene evidence relation | VariantGeneLinkSource and VariantGeneLinkStore |
A normalized evidence release intended to compose across analyses |
This means boundedness is not the dividing line. A five-row-per-variant splice prediction and a ten-thousand-track prediction are both model details when an Altar run produced them. Conversely, a precomputed SpliceAI dataset remains an annotation source even if each lookup returns several gene records. A model detail becomes a variant–gene link only through an explicit normalization step whose output has the link contract's provenance and semantics.
Primary scores and details¶
The primary score table has one scalar row per (model_id, variant_id). It is deliberately small enough for
filtering, prioritization, and cross-model materialization. Named detail tables preserve gene-, event-,
track-, cell-, or other repeated observations without flattening or truncating them.
A binding may reduce detail rows into headline primary fields. Those primary fields must declare their own
meaning in ResultSchema; identifiers needed to explain the selected row, such as the top gene or event,
should also be primary fields. The lossless source observations remain in the named detail table. Altar does
not assume that every detail table has a meaningful reducer.
Container output roles¶
Container plans classify native terminal files with ContainerResultFiles. The reserved primary role is
decoded by scoring_result_codec(); every other role must exactly match a manifest-declared
DetailSchema.name and is decoded by detail_result_codec(name).
Classification uses storage-neutral logical paths from the plan. Deployment code may map the file URI, but
it cannot change the logical path or infer a result role from a bucket name. Every terminal output must be
classified exactly once. The scoring engine validates the complete set of roles and paths before launching
compute, then routes decoded batches independently through ScoreStore and DetailStore.
from altar.models import ContainerResultFiles, ContainerScoringPlan
plan = ContainerScoringPlan(
shards=[score_task],
summarize=summarize_task,
ready_when=1,
plugin_identity=identity,
result_files=(
ContainerResultFiles("primary", ("jobs/J/results/M/scores.tsv",)),
ContainerResultFiles("splice_events", ("jobs/J/results/M/splice-events.tsv",)),
),
)
Bindings without named details may omit result_files; all terminal outputs then retain the original
primary-result behavior. A binding that declares details must classify both the primary and every declared
detail result explicitly.
Examples¶
- A live SpliceAI or Pangolin run: a scalar maximum in the primary table and variant × gene × splice-event rows in a named detail table.
- A precomputed SpliceAI VCF: an
AnnotationSource, including any lossless repeated gene records supplied by that release. - AlphaGenome or scooby: headline scalar reductions plus track- or cell-level detail tables.
- ENCODE-rE2G/scE2G atlas records: normalized variant–gene links when they satisfy the reusable link contract.
- Borzoi, Akita, or Orca matrices too large for row-wise detail: a future typed artifact result; do not encode a matrix as thousands of scalar columns.
This boundary keeps each model binding, execution backend, and storage adapter independent: M model integrations plus N backends, rather than M × N combinations.