Skip to content

Plugin ABI and result identity

Every model plugin ships an immutable PluginManifest. The manifest is the durable contract between a wrapper, Altar, and a host that persists runs or results; model_type remains only the discovery key. Entry points use altar.model_plugins (and the other canonical altar.* groups).

The manifest records:

  • a reverse-DNS-style globally namespaced plugin ID, plugin kind, capabilities, and model family;
  • one plugin_version for the binding release and the half-open range of Altar API versions it supports;
  • configuration-schema version;
  • an ordered result schema, its version, and a SHA-256 content hash;
  • prioritization policy/version and versioned annotation-column dependencies; and
  • named runtime roles with an identity policy: record the exact value, require an image digest, or require a fixed declared service identity; and
  • optionally, the variant classes and maximum allele length the model can score (variant_eligibility). An omitted declaration admits every variant and is left out of the canonical JSON, so manifests that don't declare it keep their existing hash.

PluginManifest.to_json() is canonical JSON: object keys are sorted, separators are fixed, non-finite numbers are forbidden, unordered sets are sorted, and ordered sequences remain ordered. A result-schema hash is computed over ordered field definitions (name, scalar type, biological meaning, and nullability), not Python object hashes. It is therefore stable across mapping insertion order, operating processes, and PYTHONHASHSEED. Display labels and schema-version strings remain in the complete manifest provenance but do not change the meaning of stored values; changing column order remains intentionally detectable.

Run and result contracts

PluginRunIdentity is a plain, transport-safe score contract produced when a plan is built. It carries the plugin ID and version, supported Altar API range and exact planning API version, result schema hash, selected reducer, and actual runtime/image identities. It does not contain prioritization policy, predicate hashes, annotation dependencies, schema release labels, or the complete manifest hash. The latter would indirectly reintroduce prioritization into score identity. A scoring identity additionally records canonical hashes of both the complete validated model configuration and its scientific projection, plus a hash of the complete output selection and the genome build. When ModelRunInputs.genome is present, its structured resource identity is included in both configuration hashes and recorded separately as genome_hash, while genome_build remains the readable assembly label. The engine compares both fields with the request before compute or persistence. The complete hash distinguishes exact provenance. The scientific projection omits fields typed to contain SecretReference, because authentication changes how a run obtains access—not what the model computes. Container plans attach the same snapshot to every ContainerTaskSpec; spec_to_dict preserves it when a host stores a task in a ledger. Inline plans carry it directly.

Bindings create scoring identities with ModelPlugin.scoring_run_identity(). Put every durable setting that can change score meaning in ModelConfiguration: checkpoints, resource revisions or digests, ontology and tissue filters, masking, sequence windows, and similar scientific settings. Prefer immutable ResourceReference.digest or revision values; hashing a mutable URI makes its spelling reproducible, not the bytes later served there. SHA-256 digests are validated and take precedence in scientific identity, so relocating identical content preserves cache compatibility. Without a digest, URI plus immutable revision is used; subtype selectors such as a filename remain identity-bearing. A digest addresses the resource narrowed by those selectors, so a multi-file selector needs a bundle/snapshot digest. Every field added by a ResourceReference subclass is identity-bearing; represent operational location metadata through uri, not fields such as region or endpoint. Supply ModelRunInputs.genome with the reference FASTA digest whenever sequence bytes affect scores. scoring_run_identity() accepts the complete validated run_inputs and derives the genome fields itself. Use SecretReference only in wholly operational fields; Altar excludes any field whose declared type contains it from ModelConfiguration.scientific_dict(), recursively and even when the current value is None. Do not combine credentials and scientific alternatives in one union-typed field. The scoring identity deliberately excludes job and model IDs, subject variants, batch boundaries, paths, and injected runtime clients, so separate batches and credential rotations can incrementally populate the same model-instance cache.

Adding a digest to a previously URI/revision-only resource intentionally changes its scientific identity. Configuration schema versions still belong to saved configurations, where they select a decoder/migration; they are not a second plugin release number. Result/detail schemas also retain format versions alongside their content hashes. Those format declarations live in manifests and saved formats, not as additional version counters in score lineage.

run_scoring() passes this identity to ScoreStore.add_scores(). The complete identity is immutable provenance. Reuse is gated by a smaller, role-specific ResultCompatibility: primary scores and each named detail table have independent fingerprints. Both include the scientific model-configuration projection, output selection, genome, reducer, exact runtime, plugin version, and the relevant semantic schema. Complete configuration provenance, credential references, labels, and descriptions are not scientific cache gates. Prioritization stays in the plugin's materialization contract, not in score provenance. Reference stores retain every exact compatible identity in row metadata or a provenance sidecar and return one deterministic representative through their existing singular metadata API. A conforming store rejects scientifically incompatible values under the same model ID.

Compact-lineage stores key each lineage on the run's scientific projection, PluginRunIdentity.scientific_dict(), record every distinct exact identity that writes under it in a provenance table, and support explicit acceptance of previous lineages through ScoreReusePolicy. Plugin/image upgrades therefore create new lineages; they are not automatically declared equivalent. Framework API metadata remains in the exact identity for execution checks, even though it is excluded from ResultCompatibility and from the lineage. See what decides a generated lineage.

This is a breaking development contract: wrapper_version and scoring_semantics_version are removed, not aliases. Serialized manifests, saved configurations, and plans must use plugin_version. Old run identities containing removed fields are rejected. Existing persisted records are not rewritten or assigned new provenance automatically. Rebuild development plans/configuration envelopes before executing with this contract; frozen historical imports without an Altar run identity remain a separate explicit reuse mechanism.

To materialize stored scores, construct ScoredModel with the identity expected by the current model configuration. The store loads its recorded provenance and validates primary-score compatibility before decoding rows. The current prioritization policy is then applied to those compatible raw scores, so changing a threshold requires rematerialization, not rescoring. A scientific mismatch raises IncompatibleResultError; old rows are never silently assigned a new meaning. Materialized results expose a plugin_identities mapping so downstream transport retains the provenance without changing scientific score values.

Altar API compatibility is checked when a model plugin is instantiated, when a run identity is created, when a recorded container task is reconstructed for execution, and again before results are interpreted. The error names the wrapper's supported range and the running API version.

Change rules

Treat the following as additive only when old consumers remain correct:

  • expanding the supported Altar API range after conformance testing;
  • adding a capability that does not change existing plans or results;
  • adding an unrelated named detail result or changing labels/descriptions;
  • widening variant_eligibility so that it admits more variants. Eligibility selects inputs but never changes an eligible variant's score, so it is excluded from PluginRunIdentity and score lineage;
  • adding an optional configuration field while keeping its old default and meaning; or
  • fixing plugin implementation details without changing produced values. Bump the plugin patch/minor version as appropriate. Configuration changes also bump the configuration-schema version when the accepted serialized shape changes.

Treat these as breaking result changes:

  • adding, removing, renaming, or reordering a result column;
  • changing a column's scalar type, nullability, order, or biological meaning;
  • changing units, normalization, reduction, sign convention, or missing-value interpretation;
  • changing how to_model_score decodes a stored value; or
  • narrowing variant_eligibility to exclude variants an earlier release stored valid scores for. Those rows keep the same score identity and remain readable, so bump plugin_version or purge them explicitly.

Use plugin_version for each binding release, including value-producing changes. There is no independent scoring-semantics release counter. Document scientific changes and explicitly approve previous score lineages when their reuse is acceptable. When the result fields also change, bump the result-schema major version and ship the new schema hash. Do not rewrite the old manifest in place. Keep the old wrapper available to decode old rows, or provide an explicit, host-owned migration that reads with the old identity, transforms the data, validates the new schema, and atomically records the new identity. A host may accept an added nullable column as an additive migration only through such an explicit migration; the runtime validator still requires an exact schema hash so a partially migrated result cannot be mistaken for either schema.

Changing thresholds, predicate structure, policy meaning, or required annotations is a prioritization-policy change. Its declaration and required annotations belong to materialization, not to score lineage. Materialization resolves each required annotation against the configured sources' declared contracts; see Use your own annotation tables. Updating the policy or expression alone leaves the score identity unchanged. Compatible raw scores may be re-prioritized under the new policy; any previously materialized result must be rebuilt. If the change ships with a new plugin release, accept the prior score lineage explicitly rather than assuming all releases are scientifically interchangeable.

Runtime policies are similarly explicit. record_exact accepts a tag or provider identifier but records the exact value used. Published reproducible container plugins should use digest_required; mutable tags then fail before submission. The Cherimoya reference binding additionally requires a SHA-256 identity for every generic weight resource. Model runtimes receive staged local paths and never embed a storage provider. declared_exact is for a provider-managed service whose stable public identity is declared in the manifest.

Discovery collision check

The registry validates every loaded object against its group's ABC and discriminator. A model class must declare a compatible manifest of kind model, include the capabilities required by the model result contract, use an entry-point key equal to model_type, and own a globally unique manifest plugin ID. Metadata collisions and load/validation failures remain separate diagnostics so one broken optional plugin does not hide healthy plugins.