Skip to content

Add a model binding

A model binding owns one architecture's versioned result meaning, strict model and run configuration, interpretation policy, and scoring plan. It does not own a compute provider.

1. Define the durable manifest

Every model plugin declares a PluginManifest. The manifest is persisted with plans and results so a newer wrapper cannot silently reinterpret older rows.

from altar.models import (
    ApiVersionRange,
    PluginCapability,
    PluginKind,
    PluginManifest,
    PrioritizationPolicy,
    ResultField,
    ResultSchema,
    RuntimeIdentityPolicy,
    RuntimeKind,
    RuntimeRequirement,
)


EXAMPLE_MANIFEST = PluginManifest(
    plugin_id="org.example.examplenet.model",
    kind=PluginKind.MODEL,
    capabilities=frozenset(
        {
            PluginCapability.SCORE,
            PluginCapability.MATERIALIZE_RESULTS,
        }
    ),
    architecture="ExampleNet",
    plugin_version="1.0.0",
    altar_api=ApiVersionRange("0.1.0", "0.2.0"),
    configuration_schema_version="1.0.0",
    result_schema=ResultSchema(
        version="1.0.0",
        fields=(
            ResultField(
                "effect",
                "float",
                "Predicted effect",
                "Signed alternate-versus-reference effect",
            ),
        ),
    ),
    prioritization=PrioritizationPolicy("examplenet-effect", "1.0.0"),
    runtimes=(
        RuntimeRequirement(
            name="model_runtime",
            kind=RuntimeKind.CONTAINER_IMAGE,
            policy=RuntimeIdentityPolicy.DIGEST_REQUIRED,
        ),
    ),
)

Use a globally namespaced plugin ID. Give result fields exact units, direction, nullability, and biological meaning. Read Versioned plugin ABI before choosing schema, policy, or runtime versions. Use one plugin_version for package releases and document any score-changing behavior in the release notes. Reusing scores from another release is an explicit ScoreReusePolicy decision, not an inference from the version number. Configuration/result schema versions describe saved formats; change them when those formats need different interpretation, not mechanically with every plugin release.

Declare the variants the model can score

If the runtime can score only some variant classes, declare them with variant_eligibility:

from altar.models import VariantClass, VariantEligibility

SNV_ONLY_MANIFEST = replace(
    EXAMPLE_MANIFEST,
    variant_eligibility=VariantEligibility(frozenset({VariantClass.SNV})),
)

VariantClass has four values: snv, mnv, insertion and deletion. It compares REF and ALT lengths as written and does not trim or left-align, so A>AT and A>GT are both insertions. max_allele_length optionally limits the longer allele, in bases. The default admits every variant; leave it unset when the runtime handles every class itself, for example by returning a null row for a variant that doesn't fit its input window.

prepare_scoring_batches() excludes and reports ineligible variants before any cache lookup. The engine rejects result rows for them (see variant eligibility). Eligibility filters inputs; it doesn't change the score of any eligible variant, so it is not part of PluginRunIdentity or score lineage. Keep the runtime's own check as a defense: it still guards a batch that was staged by hand.

Widening eligibility, so that the model admits variants it previously excluded, is additive. Narrowing it is not. Stores keep the rows an earlier release wrote for variants you now exclude, under the same score identity, and read_score() and materialization still return them. If an earlier release stored valid scores for variants the new declaration excludes, bump plugin_version so those rows stop matching, or purge them explicitly. Declaring SNV-only eligibility for Borzoi, Enformer, LegNet, and Sei needed neither, because their runtimes never produced a valid non-SNV score.

2. Declare result interpretation

from altar.models import ModelPlugin
from altar.predicates import Abs, Col, Ge, Lit


class ExamplePlugin(ModelPlugin):
    model_type = "EXAMPLE"
    manifest = EXAMPLE_MANIFEST

    def prioritize_predicate(self):
        return Ge(Abs(Col("effect")), Lit(0.5))

The manifest's result schema is the only column declaration. The inherited score_columns() derives the ordered ScoreColumn list from it, and ModelResultContract requires the two to match exactly. The inherited to_model_score() builds the model_scores entry by strictly decoding each declared field as its dtype: a value is converted only when that cannot change its meaning (an integer to float, a whole float to int, SQLite's 0/1 to bool), and anything else raises an error naming the column. A missing, None, or NaN value stays None for every dtype, so an unreported measurement never becomes False, 0, or an empty string. The per-model prioritized flag is the materialized decision: an unknown predicate result becomes False, matching the warehouse export. Do not copy either method into a binding; override to_model_score() only when a stored value genuinely needs a model-specific decoding, because changing how stored values decode is a breaking result change.

Document supported variants, biological scope, missingness, and the scientific status of every threshold.

3. Declare model configuration and run inputs

Durable model resources and per-run biological inputs are separate strict models. Unknown fields are rejected. Use resource references for scientific identity and SecretReference for a durable credential pointer; never put a resolved credential, SDK client, or callable in these models.

from pydantic import model_validator

from altar.models import (
    ConfigurationModel,
    ContainerImageDigest,
    ModelConfiguration,
    ModelRunInputs,
    OutputSelection,
    ResourceReference,
    RuntimeContext,
    require_resource_digests,
)


class ExampleConfiguration(ModelConfiguration):
    weights: ResourceReference
    image: ContainerImageDigest

    @model_validator(mode="after")
    def require_content_addressed_resources(self) -> "ExampleConfiguration":
        require_resource_digests(self, owner="ExampleNet")
        return self


class ExampleVariant(ConfigurationModel):
    variant_id: str
    chromosome: str
    position: int
    reference: str
    alternate: str


class ExampleRunInputs(ModelRunInputs):
    variants: tuple[ExampleVariant, ...]


class ExampleRuntime(RuntimeContext):
    __slots__ = ("client",)

    def __init__(self, client: object) -> None:
        self.client = client


class ExamplePlugin(ModelPlugin):
    # manifest and result methods omitted here
    configuration_schema_id = "org.example.examplenet.model.configuration"
    configuration_model = ExampleConfiguration
    run_input_model = ExampleRunInputs
    default_output = OutputSelection(reducer="mean")

ContainerImageDigest accepts only an immutable name@sha256:... image, which a digest_required runtime needs. The name follows the OCI reference grammar, [registry[:port]/]path[:tag], with lowercase path components such as ghcr.io/org/model:1.2; the type and the digest_required check use the same pattern. A field default is not validated, so test that a binding's pinned default image passes. require_resource_digests() rejects any ResourceReference field, or tuple of them, that lacks a SHA-256 digest, naming each one (weights[2]). Use both when the runtime verifies the bytes it stages, and keep model-specific checks, such as distinct fold digests, in the same validator.

configuration.to_dict() is the complete saved configuration. configuration.scientific_dict() is the cache-compatibility projection. Altar recursively excludes any field whose declared type contains SecretReference, including its None state, so rotating or removing credentials does not force rescoring. Treat such a field as wholly operational—do not use one union-typed field for both a credential and a scientific resource.

The semantic configuration version comes from manifest.configuration_schema_version. A generic caller can render help without importing a model SDK:

plugin = ExamplePlugin()
configuration_schema = plugin.configuration_schema()
run_input_schema = plugin.run_input_schema()

Both methods return JSON Schema. save_configuration() attaches plugin, wrapper, schema, and version identity; load_configuration() verifies it. Override migrate_configuration() when a future schema version needs an explicit migration. ExampleRuntime is intentionally non-durable: its representation is redacted and any serialization attempt fails.

For existing callers, implement convert_legacy_artifact() only for the exact legacy shape you support. Reject unknown keys and invalid values rather than accepting a free-form compatibility bag.

4. Build a plan and record runtime identity

Implement build_scoring_plan() with one of these forms:

  • InlineScoringPlan for a hosted API or lightweight in-process scorer;
  • ContainerScoringPlan for one or more container tasks and an optional dependent summary task.

Validate request.configuration, request.run_inputs, and request.output first. Use ScoringRequest.layout and resolver for paths; declare every input, output, resource request, and structured task label. ScoringLayout is model-neutral: it names the staged genome, the variant input, per-member intermediate scores, and the primary and named-detail result files. Name model-owned files with layout.model_resource_file(model_id, basename) and choose the basename from declared configuration, such as a weight-format field, never from a resource URI's suffix; a content-addressed URI need not have an extension. Architecture-specific paths, such as derived calibration artifacts or an interpretation directory tree, belong in the binding. InterpretationRequest and PreprocessingRequest carry no layout; the binding owns those paths entirely. Build the plan's plugin_identity with self.scoring_run_identity({"model_runtime": actual_runtime_identity}, configuration=configuration, output=output, run_inputs=run_inputs). Altar derives both the readable build label and structured genome from the validated inputs, so a binding cannot accidentally drop genome content identity. This binds persisted values to the complete validated model configuration, output selection, and content-addressed genome when supplied, while still allowing different jobs, subjects, and batches to populate one model-instance cache. Supply ResourceReference(uri=..., digest="sha256:...") for the actual reference bytes. A digest-required container runtime must receive an immutable name@sha256:... value.

Treat uri as location only. A container binding declares the resource as an input Transfer, preserving its digest, and passes the backend-resolved logical path to the runtime. resource_transfer(resource, logical_path, locality_key=model_id) builds that transfer from a configured resource, so a binding does not copy the URI and digest by hand. The storage adapter retrieves the URI; the runtime must not import a bucket SDK, call a model registry, or accept provider-specific repository arguments. This keeps model × storage composition M + N rather than M × N, and a digest lets callers relocate identical bytes without changing scientific cache identity.

Inline plans contain only an executor name and JSON payload. Implement execute_inline(plan, runtime) to combine that durable descriptor with a separately supplied RuntimeContext; do not close over a client or function. scoring_plan_to_json() and scoring_plan_from_json() must work across a process restart.

Preparation and interpretation are separate optional operations. Training is outside Altar: register trained weights and other scientific resources in the configuration. If a model must derive reusable artifacts from those resources before scoring, implement build_preprocessing_plan() and return a ModelWorkflowPlan whose operation is "preparation". If it supports contribution scores, motif analysis, plots, or another post-score analysis, implement build_interpretation_plan() with operation "interpretation".

Both operations use named WorkflowStage values. Tasks inside one stage run in parallel; needs names the stages that must succeed first. Stage names and topology belong to the binding, so core does not assume that a model has folds, motifs, or one preparation job. Declare terminal artifact transfers in plan.outputs, then serialize with workflow_plan_to_json() or execute the complete DAG with run_model_workflow().

Preparation provenance is output metadata, not custody enforcement. A binding may emit a manifest and a caller may retain or verify it, but scoring must not require that optional metadata unless the model itself scientifically requires it. Trusted externally managed artifacts remain valid inputs. Preparation, scoring, and interpretation are independently advertised capabilities; a model may implement any supported subset.

Do not import Docker, Kubernetes, Modal, a heavyweight model SDK, or a host application at module import time. If scoring needs a heavyweight framework, ship its CLI and dependency lock as a separate runtime project.

5. Decode native results

Container bindings own the translation from runtime files to their versioned ResultSchema. The default scoring_result_codec() expects one or more headered TSV files containing variant_id plus exactly the manifest field names. Override it only when the runtime file differs:

from altar.scoring import DelimitedResultCodec


class ExamplePlugin(ModelPlugin):
    # manifest and plan methods omitted

    def scoring_result_codec(self):
        return DelimitedResultCodec(
            self.manifest.result_schema,
            source_names={"effect": "native_delta"},
            row_identity_fields={"chrom": "str", "position": "int"},
            allowed_extra_fields=("transport_only",),
        )

Mappings rename runtime syntax, not scientific meaning. row_identity_fields preserves typed identity columns needed by a store; allowed_extra_fields lists only deliberately ignored transport fields. An unexpected new column is an error. The codec streams bounded batches and the common engine validates the decoded row again before calling a generic ScoreStore. Do not emit a decomposed locus only because a store needs one: the engine derives the portable chr/pos/ref/alt fields a store requires from variant_id.

When a runtime also emits repeated observations, declare a DetailSchema in the manifest and classify the terminal files in ContainerScoringPlan.result_files. The engine sends only the primary files to scoring_result_codec() and sends each named detail file set to detail_result_codec(name). The default detail codec expects exact-schema TSV files; override it for native names with DelimitedDetailResultCodec. See Choose the result shape before deciding between a model detail, imported annotation, variant–gene link, or large artifact.

6. Register and test

[project.entry-points."altar.model_plugins"]
EXAMPLE = "altar_example:ExamplePlugin"
import pytest
from altar.testing import ContainerScorerContract, ModelResultContract


class TestExampleResults(ModelResultContract):
    @pytest.fixture
    def plugin(self):
        return ExamplePlugin()


class TestExampleContainerPlan(ContainerScorerContract):
    @pytest.fixture
    def plugin(self):
        return ExamplePlugin()

    @pytest.fixture
    def scoring_request(self):
        return make_synthetic_scoring_request()

Use InlineScorerContract instead for an inline plugin, and add PluginDiscoveryContract in the separately installed distribution test. The capability contract validates schemas, task labels, transfers, resources, declared terminal outputs, serialization, and runtime identity. Also golden-test exact task commands, schema hashes, native-result decoding, missing values, and fold behavior. Include one test that calls altar.scoring.run_scoring() and reads the row back from a SqliteScoreStore; this proves the binding fits the public end-to-end path. Use synthetic inputs for unit tests; keep real GPU/image tests as a separate integration gate.