Plugin ABI and result identity¶
Every model plugin ships an immutable PluginManifest. The manifest is the durable contract between a
wrapper, Altar, and a host that persists runs or results; model_type remains only the discovery key. Entry
points use altar.model_plugins (and the other canonical altar.* groups).
The manifest records:
- a reverse-DNS-style globally namespaced plugin ID, plugin kind, capabilities, and model family;
- one
plugin_versionfor the binding release and the half-open range of Altar API versions it supports; - configuration-schema version;
- an ordered result schema, its version, and a SHA-256 content hash;
- prioritization policy/version and versioned annotation-column dependencies; and
- named runtime roles with an identity policy: record the exact value, require an image digest, or require a fixed declared service identity; and
- optionally, the variant classes and maximum allele length the model can score (
variant_eligibility). An omitted declaration admits every variant and is left out of the canonical JSON, so manifests that don't declare it keep their existing hash.
PluginManifest.to_json() is canonical JSON: object keys are sorted, separators are fixed, non-finite
numbers are forbidden, unordered sets are sorted, and ordered sequences remain ordered. A result-schema hash
is computed over ordered field definitions (name, scalar type, biological meaning, and nullability), not
Python object hashes. It is therefore stable across mapping insertion order, operating processes, and
PYTHONHASHSEED. Display labels and schema-version strings remain in the complete manifest provenance but
do not change the meaning of stored values; changing column order remains intentionally detectable.
Run and result contracts¶
PluginRunIdentity is a plain, transport-safe score contract produced when a plan is built. It carries the
plugin ID and version, supported Altar API range and exact planning API version, result schema hash,
selected reducer, and actual runtime/image identities. It does not contain prioritization policy,
predicate hashes, annotation dependencies, schema release labels, or the complete manifest hash. The latter
would indirectly reintroduce prioritization into score identity. A scoring identity additionally records canonical hashes of both the
complete validated model configuration and its scientific projection, plus a hash of the complete output
selection and the genome build. When ModelRunInputs.genome is present, its structured resource identity is
included in both configuration hashes and recorded separately as genome_hash, while genome_build remains
the readable assembly label. The engine compares both fields with the request before compute or persistence.
The complete hash distinguishes exact provenance. The scientific projection omits fields typed to contain
SecretReference, because authentication changes how a run obtains access—not what the model computes.
Container plans attach the same snapshot to every
ContainerTaskSpec; spec_to_dict preserves it when a host stores a task in a ledger. Inline plans carry it
directly.
Bindings create scoring identities with ModelPlugin.scoring_run_identity(). Put every durable setting that
can change score meaning in ModelConfiguration: checkpoints, resource revisions or digests, ontology and
tissue filters, masking, sequence windows, and similar scientific settings. Prefer immutable
ResourceReference.digest or revision values; hashing a mutable URI makes its spelling reproducible, not
the bytes later served there. SHA-256 digests are validated and take precedence in scientific identity, so
relocating identical content preserves cache compatibility. Without a digest, URI plus immutable revision is
used; subtype selectors such as a filename remain identity-bearing. A digest addresses the resource narrowed
by those selectors, so a multi-file selector needs a bundle/snapshot digest. Every field added by a
ResourceReference subclass is identity-bearing; represent operational location metadata through uri, not
fields such as region or endpoint. Supply ModelRunInputs.genome with the reference FASTA digest whenever
sequence bytes affect scores. scoring_run_identity() accepts the complete validated run_inputs and derives
the genome fields itself. Use SecretReference only in wholly operational fields; Altar excludes any
field whose declared type contains it from ModelConfiguration.scientific_dict(), recursively and even when
the current value is None. Do not combine credentials and scientific alternatives in one union-typed field.
The scoring identity deliberately excludes job and model IDs, subject variants, batch boundaries, paths, and
injected runtime clients, so separate batches and credential rotations can incrementally populate the same
model-instance cache.
Adding a digest to a previously URI/revision-only resource intentionally changes its scientific identity. Configuration schema versions still belong to saved configurations, where they select a decoder/migration; they are not a second plugin release number. Result/detail schemas also retain format versions alongside their content hashes. Those format declarations live in manifests and saved formats, not as additional version counters in score lineage.
run_scoring() passes this identity to ScoreStore.add_scores(). The complete identity is immutable
provenance. Reuse is gated by a smaller, role-specific ResultCompatibility: primary scores and each named
detail table have independent fingerprints. Both include the scientific model-configuration projection,
output selection, genome, reducer, exact runtime, plugin version, and the relevant semantic schema.
Complete configuration provenance, credential references, labels, and descriptions are not scientific cache
gates. Prioritization stays in the plugin's materialization contract, not in score provenance. Reference stores retain
every exact compatible identity in row metadata or a provenance sidecar and return one deterministic
representative through their existing singular metadata API. A conforming store rejects scientifically
incompatible values under the same model ID.
Compact-lineage stores key each lineage on the run's scientific projection,
PluginRunIdentity.scientific_dict(), record every distinct exact identity that writes under it in a
provenance table, and support explicit acceptance of previous lineages through ScoreReusePolicy.
Plugin/image upgrades therefore create new lineages; they are not automatically declared equivalent. Framework
API metadata remains in the exact identity for execution checks, even though it is excluded from
ResultCompatibility and from the lineage. See what decides a generated
lineage.
This is a breaking development contract: wrapper_version and scoring_semantics_version are removed,
not aliases. Serialized manifests, saved configurations, and plans must use plugin_version. Old run
identities containing removed fields are rejected. Existing persisted records are not rewritten or assigned
new provenance automatically. Rebuild development plans/configuration envelopes before executing with this
contract; frozen historical imports without an Altar run identity remain a separate explicit reuse mechanism.
To materialize stored scores, construct ScoredModel with the identity expected by the current model
configuration. The store loads its recorded provenance and validates primary-score compatibility before
decoding rows. The current prioritization policy is then applied to those compatible raw scores, so changing a
threshold requires rematerialization, not rescoring. A scientific mismatch raises
IncompatibleResultError; old rows are never silently assigned a new meaning. Materialized results expose a
plugin_identities mapping so downstream transport retains the provenance without changing scientific score
values.
Altar API compatibility is checked when a model plugin is instantiated, when a run identity is created, when a recorded container task is reconstructed for execution, and again before results are interpreted. The error names the wrapper's supported range and the running API version.
Change rules¶
Treat the following as additive only when old consumers remain correct:
- expanding the supported Altar API range after conformance testing;
- adding a capability that does not change existing plans or results;
- adding an unrelated named detail result or changing labels/descriptions;
- widening
variant_eligibilityso that it admits more variants. Eligibility selects inputs but never changes an eligible variant's score, so it is excluded fromPluginRunIdentityand score lineage; - adding an optional configuration field while keeping its old default and meaning; or
- fixing plugin implementation details without changing produced values. Bump the plugin patch/minor version as appropriate. Configuration changes also bump the configuration-schema version when the accepted serialized shape changes.
Treat these as breaking result changes:
- adding, removing, renaming, or reordering a result column;
- changing a column's scalar type, nullability, order, or biological meaning;
- changing units, normalization, reduction, sign convention, or missing-value interpretation;
- changing how
to_model_scoredecodes a stored value; or - narrowing
variant_eligibilityto exclude variants an earlier release stored valid scores for. Those rows keep the same score identity and remain readable, so bumpplugin_versionor purge them explicitly.
Use plugin_version for each binding release, including value-producing changes. There is no independent
scoring-semantics release counter. Document scientific changes and explicitly approve previous score
lineages when their reuse is acceptable. When the result fields also change,
bump the result-schema major version and ship the new schema hash. Do not
rewrite the old manifest in place. Keep the old wrapper available to decode old rows, or provide an explicit,
host-owned migration that reads with the old identity, transforms the data, validates the new schema, and
atomically records the new identity. A host may accept an added nullable column as an additive migration only
through such an explicit migration; the runtime validator still requires an exact schema hash so a partially
migrated result cannot be mistaken for either schema.
Changing thresholds, predicate structure, policy meaning, or required annotations is a prioritization-policy change. Its declaration and required annotations belong to materialization, not to score lineage. Materialization resolves each required annotation against the configured sources' declared contracts; see Use your own annotation tables. Updating the policy or expression alone leaves the score identity unchanged. Compatible raw scores may be re-prioritized under the new policy; any previously materialized result must be rebuilt. If the change ships with a new plugin release, accept the prior score lineage explicitly rather than assuming all releases are scientifically interchangeable.
Runtime policies are similarly explicit. record_exact accepts a tag or provider identifier but records the
exact value used. Published reproducible container plugins should use digest_required; mutable tags then
fail before submission. The Cherimoya reference binding additionally requires a SHA-256 identity for every
generic weight resource. Model runtimes receive staged local paths and never embed a storage provider.
declared_exact is for a provider-managed service whose stable public identity is declared in the manifest.
Discovery collision check¶
The registry validates every loaded object against its group's ABC and discriminator. A model class must
declare a compatible manifest of kind model, include the capabilities required by the model result
contract, use an entry-point key equal to model_type, and own a globally unique manifest plugin ID.
Metadata collisions and load/validation failures remain separate diagnostics so one broken optional plugin
does not hide healthy plugins.