Skip to content

Models and plans

Model binding authors should import from altar.models. This stable façade exposes model manifests, strict configuration and run-input bases, typed resource and credential references, saved configuration envelopes, durable plan serializers, model plugins, and logical path layouts without importing a model SDK. Importing it also loads no storage adapter or concrete execution backend; those come from altar.results, altar.sources, and altar.execution.

models

Stable model-binding contracts.

External model bindings should import from this module. Names re-exported here are covered by Altar's public compatibility policy; their original altar.plugins.model locations are implementation paths.

VariantIdentityError

Bases: ValueError

A variant key is malformed, non-canonical, or disagrees with its locus.

VariantKey dataclass

VariantKey(
    chromosome: str,
    position: int,
    reference_allele: str,
    alternate_allele: str,
)

Canonical one-based biallelic key serialized as chr:pos:ref:alt.

Fields follow the rules in this module's documentation. The class normalizes spelling but does not normalize biological representation. Callers must left-align and trim indels against the selected reference before constructing a key when equivalence across representations matters.

variant_id property

variant_id: str

Return the portable score, detail, and annotation join key.

parse classmethod

parse(variant_id: str) -> VariantKey

Parse and normalize a four-field textual key.

require_canonical classmethod

require_canonical(variant_id: str) -> VariantKey

Parse variant_id and reject aliases instead of silently rewriting them.

from_fields classmethod

from_fields(
    chromosome: str,
    position: int,
    reference_allele: str,
    alternate_allele: str,
) -> VariantKey

Construct a canonical key from decomposed one-based locus fields.

CapabilityNotSupported

Bases: NotImplementedError

Raised by an optional plugin method that the adapter deliberately does not implement.

It subclasses NotImplementedError, so an existing except NotImplementedError still catches it. It is a named, catchable signal distinct from a genuinely missing abstract method: a caller probing an optional capability can except CapabilityNotSupported without also swallowing real not-yet-implemented bugs.

AnnotationDependency dataclass

AnnotationDependency(
    source_id: str,
    schema_version: str,
    columns: tuple[str, ...],
)

An annotation contract required by the prioritization policy.

ApiVersionRange dataclass

ApiVersionRange(minimum: str, maximum_exclusive: str)

A half-open Altar API interval: minimum <= version < maximum.

DetailSchema dataclass

DetailSchema(
    name: str,
    version: str,
    description: str,
    fields: tuple[ResultField, ...],
    key_fields: tuple[str, ...],
)

One versioned, named one-to-many result table owned by a model binding.

variant_id is the implicit parent key and must not be repeated in fields. key_fields names the binding-declared fields that, together with variant_id, identify one logical detail row. Key values may be null when the upstream science has no value (for example a non-gene-centric track has no gene ID); stores use Altar's canonical logical-key encoding rather than backend-specific NULL equality.

schema_hash property

schema_hash: str

Hash row identity and biological meaning, excluding descriptions and display labels.

IncompatibleAltarApiError

Bases: ManifestError

Raised when a plugin does not support the running Altar API.

IncompatibleResultError

Bases: ManifestError

Raised before results are decoded with an incompatible plugin contract.

IneligibilityReason

Bases: StrEnum

Why a model's :class:VariantEligibility excludes a variant.

ManifestError

Bases: ValueError

Base class for invalid or incompatible plugin ABI data.

PluginManifest dataclass

PluginManifest(
    plugin_id: str,
    kind: PluginKind,
    capabilities: frozenset[PluginCapability],
    architecture: str | None,
    plugin_version: str,
    altar_api: ApiVersionRange,
    configuration_schema_version: str,
    result_schema: ResultSchema,
    prioritization: PrioritizationPolicy,
    runtimes: tuple[RuntimeRequirement, ...],
    detail_schemas: tuple[DetailSchema, ...] = (),
    variant_eligibility: VariantEligibility = VariantEligibility(),
    manifest_schema_version: str = MANIFEST_SCHEMA_VERSION,
)

The immutable declaration shipped by one plugin release.

detail_schema

detail_schema(name: str) -> DetailSchema

Return one declared detail schema by name or fail with the available names.

PluginRunIdentity dataclass

PluginRunIdentity(
    plugin_id: str,
    plugin_kind: PluginKind,
    architecture: str | None,
    plugin_version: str,
    supported_altar_api: ApiVersionRange,
    altar_api_version: str,
    result_schema_hash: str,
    runtimes: tuple[RuntimeIdentity, ...],
    reducer: str | None = None,
    configuration_hash: str | None = None,
    output_selection_hash: str | None = None,
    genome_build: str | None = None,
    scientific_configuration_hash: str | None = None,
    genome_hash: str | None = None,
)

Transport-safe exact identity attached to a run and its result set.

The optional scoring-context fields distinguish concrete model instances, not merely wrapper ABIs. configuration_hash preserves the complete durable configuration for exact provenance, while scientific_configuration_hash may omit typed operational credentials for cache compatibility. genome_hash records the structured genome's scientific projection separately so an engine can verify that a plan was built from the request it is about to execute. These fields are optional when inspecting a plugin before configuring a concrete model, but persistence requires :meth:assert_scoring_context. Prioritization, annotation dependencies, schema release labels, and the full manifest hash deliberately stay out of this identity: changing downstream selection must not invalidate raw scores.

assert_altar_compatible

assert_altar_compatible(
    version: str = ALTAR_API_VERSION,
) -> None

Reject executing this recorded plan under an API its wrapper did not support.

assert_scoring_context

assert_scoring_context() -> None

Reject unconfigured identities that do not identify weights, outputs, and genome.

primary_result_compatibility

primary_result_compatibility() -> ResultCompatibility

Return the scientific cache identity for the primary scalar score table.

scientific_dict

scientific_dict() -> dict[str, Any]

Return the projection that keys and defines a generated score lineage.

A lineage identifies what can change a primary score: the plugin release and its exact runtimes, the scientific configuration projection, the output selection, the genome build and the genome's scientific hash, and the primary result schema. The exact configuration_hash stays out because it also covers resource locations and every declared serialization detail, so a schema migration that adds defaulted fields, or unchanged bytes served from a new location, would otherwise orphan every cached score. The framework API versions stay out for the same reason.

Two runs whose projections are equal share one lineage ID, and a lineage store accepts either run for that lineage (ScoreLineage.same_lineage), so each reuses the other's scores. Exact provenance is not lost: the catalog keeps one deterministic representative identity, and each distinct exact identity that registers to write scores under the lineage is recorded once in the store's lineage provenance table, read back as a set with ScoreStore.read_lineage_provenance.

detail_result_compatibility

detail_result_compatibility(
    schema: DetailSchema,
) -> ResultCompatibility

Return the scientific cache identity for one named detail result table.

ResultCompatibility dataclass

ResultCompatibility(
    role: ResultRole,
    plugin_id: str,
    plugin_kind: PluginKind,
    architecture: str | None,
    plugin_version: str,
    runtimes: tuple[RuntimeIdentity, ...],
    reducer: str | None,
    configuration_hash: str,
    output_selection_hash: str,
    genome_build: str,
    schema_name: str,
    schema_hash: str,
)

Role-specific scientific identity used to decide whether persisted results may be reused.

Plugin releases and exact runtimes distinguish computations. Reusing a previous release requires an explicit score-lineage acceptance; version numbers alone never establish scientific equivalence. Framework compatibility and display metadata are not scientific cache gates. Prioritization is not part of score identity at all.

fingerprint property

fingerprint: str

Return the canonical role-specific compatibility fingerprint.

ResultField dataclass

ResultField(
    name: str,
    dtype: Literal["float", "bool", "int", "str"],
    label: str,
    meaning: str,
    nullable: bool = True,
)

One ordered scalar result field, including its durable biological meaning.

ResultRole

Bases: StrEnum

The persisted result role whose compatibility is being evaluated.

ResultSchema dataclass

ResultSchema(version: str, fields: tuple[ResultField, ...])

A versioned ordered schema with a biological-compatibility hash.

The complete schema, including display labels and its version, remains in the manifest for provenance. schema_hash deliberately covers only fields that change how stored values are interpreted.

schema_hash property

schema_hash: str

Hash ordered storage and biological meaning, excluding presentation labels.

RuntimeIdentityError

Bases: ManifestError

Raised when a run does not satisfy its manifest's runtime identity policy.

RuntimeIdentityPolicy

Bases: StrEnum

How an actual runtime value is checked and recorded for a run.

RuntimeRequirement dataclass

RuntimeRequirement(
    name: str,
    kind: RuntimeKind,
    policy: RuntimeIdentityPolicy,
    declared_identity: str | None = None,
)

One named runtime role and the policy for its per-run identity.

VariantClass

Bases: StrEnum

The allele-length class of one canonical biallelic variant.

Classification compares REF and ALT lengths exactly as written. It does not trim shared bases or left-align, so both an anchored VCF insertion (A>AT) and a delins that lengthens the sequence (A>GT) are insertions: a sequence model substitutes the whole ALT for the whole REF either way.

SNV class-attribute instance-attribute

SNV = 'snv'

One reference base replaced by one different base.

MNV class-attribute instance-attribute

MNV = 'mnv'

REF and ALT of equal length greater than one.

INSERTION class-attribute instance-attribute

INSERTION = 'insertion'

ALT longer than REF.

DELETION class-attribute instance-attribute

DELETION = 'deletion'

REF longer than ALT.

VariantEligibility dataclass

VariantEligibility(
    variant_classes: frozenset[VariantClass] = (
        lambda: frozenset(VariantClass)
    )(),
    max_allele_length: int | None = None,
)

The variants a model can score, declared once in its manifest.

variant_classes lists the admitted :class:VariantClass values. max_allele_length, when set, bounds the longer of REF and ALT in bases. The default admits every class at any length, which is how a manifest that declares nothing behaves. Batch preparation excludes and reports ineligible variants before cache lookup, and the scoring engine rejects result rows for them.

Eligibility filters inputs; it is not a score-changing setting. The score of an eligible variant does not depend on it, so it is excluded from :class:PluginRunIdentity and score lineage.

is_unrestricted property

is_unrestricted: bool

Whether every variant class at every allele length is eligible.

ineligibility

ineligibility(
    key: VariantKey,
) -> IneligibilityReason | None

Return why key is ineligible, or None when the model can score it.

admits

admits(key: VariantKey) -> bool

Return whether the model can score key.

Version dataclass

Version(major: int, minor: int, patch: int)

A deliberately small SemVer core used by the ABI (no prerelease syntax).

ConfigurationError

Bases: ValueError

Base error for durable model configuration.

ConfigurationIdentityError

Bases: ConfigurationError

A saved configuration belongs to a different plugin or schema.

ConfigurationManifest dataclass

ConfigurationManifest(
    plugin_id: str,
    plugin_version: str,
    schema_id: str,
    schema_version: str,
    configuration_model: type[ModelConfiguration],
    run_input_model: type[ModelRunInputs],
)

Plugin and schema identity attached to every saved configuration.

ConfigurationModel

Bases: BaseModel

Strict, frozen, JSON-serializable base for Altar's durable public values.

Unknown fields are forbidden and defaults are validated. The wrapper methods are the supported public serialization/schema API; Pydantic remains an implementation detail behind them.

json_schema classmethod

json_schema() -> dict[str, Any]

Return the machine-readable JSON Schema for this public value.

to_dict

to_dict() -> dict[str, JsonValue]

Return JSON-compatible primitives, rejecting non-durable values first.

to_json

to_json() -> str

Serialize this durable value to JSON.

from_json classmethod

from_json(value: str) -> Self

Rebuild this value from its JSON representation.

ConfigurationVersionError

Bases: ConfigurationError

A saved configuration needs a migration that the plugin does not provide.

EmptyModelConfiguration

Bases: ModelConfiguration

Configuration for a results-only plugin with no model-specific settings.

LegacyArtifactError

Bases: ConfigurationError

A legacy artifact mapping cannot be converted safely.

LegacyScoringConversion dataclass

LegacyScoringConversion(
    configuration: ModelConfiguration,
    run_inputs: ModelRunInputs,
    output: OutputSelection,
    runtime: RuntimeContext | None = None,
)

Explicit result of converting a supported legacy artifact mapping.

ModelConfiguration

Bases: ConfigurationModel

Durable, architecture-specific model configuration.

to_dict() is the complete saved configuration. scientific_dict() is the cache-identity view: fields typed to contain :class:SecretReference are operational and excluded recursively, while every other durable value remains identity-bearing.

scientific_dict

scientific_dict() -> dict[str, JsonValue]

Return the scientific cache-identity projection of this configuration.

A field whose annotation contains SecretReference is wholly operational. This type-driven rule applies even when its current value is None, so adding, removing, or rotating a credential does not change scientific identity. Nested configuration models are traversed recursively.

ModelRunInputs

Bases: ConfigurationModel

Base per-run biological inputs shared by scoring plugins.

genome_build is the readable assembly label used in paths and diagnostics. genome records the optional reference resource for integrations that can identify one. Models that read local reference sequence bytes should inherit :class:ReferenceGenomeRunInputs, which requires a content address.

NonDurableRuntimeError

Bases: TypeError

A caller attempted to serialize runtime-only injected state.

OutputSelection

Bases: ConfigurationModel

Requested named outputs and reducer, separate from model and biological inputs.

ReferenceGenomeRunInputs

Bases: ModelRunInputs

Per-run inputs for models whose results depend on staged reference FASTA bytes.

ResourceReference

Bases: ConfigurationModel

Durable identity for a model, weights, genome, or other resource.

digest is a content address and takes precedence over location metadata in scientific cache identity. It addresses the resource bytes as narrowed by any subtype selectors; a selector covering multiple files therefore needs a bundle or snapshot digest. Every additional subtype field is identity-bearing, so keep operational location metadata in uri rather than adding fields such as region or endpoint. If no digest is available, an immutable provider revision plus URI is the next-best identity. A bare URI remains supported for legacy inputs, but identifies only the name, not the bytes later served there.

scientific_dict

scientific_dict() -> dict[str, JsonValue]

Return the strongest identity plus any subtype-specific selectors.

A digest identifies the selector-narrowed resource without its location. Revision-based identities retain the URI, and subclasses retain identity-bearing selectors such as a filename within a versioned repository.

RuntimeContext

Explicitly non-durable injected clients/functions used only while executing a plan.

to_json

to_json() -> str

Reject serialization because runtime-injected state is never durable.

SavedConfiguration

Bases: ConfigurationModel

Versioned configuration envelope suitable for a database or saved plan.

SecretReference

Bases: ConfigurationModel

Durable pointer to a credential; it never contains the credential value itself.

SubjectSelection

Bases: ConfigurationModel

Which biological subjects within a run should be scored.

ContainerResultFiles dataclass

ContainerResultFiles(
    name: str, logical_paths: tuple[str, ...]
)

Collected native files for the primary table or one named detail table.

name is "primary" or a manifest-declared :class:DetailSchema name. Logical paths refer to terminal task outputs before a deployment maps their storage URIs. Keeping this classification in the portable plan prevents result codecs from guessing meaning from provider-specific paths or filenames.

ContainerScoringPlan dataclass

ContainerScoringPlan(
    shards: list[ContainerTaskSpec],
    plugin_identity: PluginRunIdentity,
    summarize: ContainerTaskSpec | None = None,
    ready_when: int = 1,
    result_files: tuple[ContainerResultFiles, ...] = (),
)

A scoring run expressed as N container shards, plus an optional fan-in summarize shard.

This is the container form of ScoringPlan, used by architectures like ChromBPNet, ProCapNet, and Cherimoya. Each shard is a ContainerTaskSpec that runs on any ExecutionBackend (local, Modal, or Kubernetes), so the same plan scores on any of them. ready_when is the number of shards to wait for before the summarize shard runs; it equals ModelPlugin.num_folds. result_files assigns collected terminal logical paths to the primary table or manifest-declared named details. It may be omitted only for the backward-compatible primary-only shape, where every terminal output is primary.

InlineScoringPlan dataclass

InlineScoringPlan(
    plugin_identity: PluginRunIdentity,
    executor: str,
    payload: dict[str, JsonValue],
)

A durable inline execution descriptor, such as an AlphaGenome API scoring request.

The plan contains only JSON data and an executor discriminator. Runtime clients, secret resolvers, and injected callables are supplied separately to ModelPlugin.execute_inline and never become plan state.

InterpretationRequest dataclass

InterpretationRequest(
    run_id: str,
    model_id: str,
    configuration: ModelConfiguration,
    run_inputs: ModelRunInputs,
    resolver: PathResolver,
)

The inputs build_interpretation_plan needs to turn one (interpret-run, model) pair into an InterpretationPlan. Like ScoringRequest, it is backend-agnostic: the binding names base-relative logical paths and the backend PathResolver supplies the mount root. Interpretation outputs are architecture-specific, so the binding owns its logical layout; core carries no interpretation layout. run_id is the interpretation run's UUID, not a job id.

ModelPlugin

ModelPlugin()

Bases: ModelResults

One model architecture's full behavior: the results methods inherited from ModelResults, plus the optional compute methods num_folds and the build_* plan builders. model_type remains the registry key and model-family discriminator; manifest is the durable versioned identity. This is the class an architecture author subclasses and registers.

model_type equals the model's stored architecture type (for example "CHROMBPNET"). Every other plugin axis keys on a name discriminator; this one keeps model_type because the value is stored and used elsewhere and cannot be renamed freely.

The compute methods are optional. By default they raise CapabilityNotSupported, so a plugin whose scores are produced elsewhere can implement none of them.

Validate Altar API compatibility before this plugin can execute or interpret data.

num_folds property

num_folds: int

Number of container shards one scoring run fans into (ChromBPNet uses 5; a single-shard or inline model uses 1). This sets ContainerScoringPlan.ready_when and the number of shards the completion fan-in waits for. Defaults to 1.

configuration_manifest classmethod

configuration_manifest() -> ConfigurationManifest

Return stable plugin/schema identity plus the two discoverable public models.

configuration_schema classmethod

configuration_schema() -> dict[str, Any]

Return this plugin's durable-configuration JSON Schema without loading a model SDK.

run_input_schema classmethod

run_input_schema() -> dict[str, Any]

Return this plugin's per-run biological-input JSON Schema.

parse_configuration

parse_configuration(
    value: Mapping[str, object] | ModelConfiguration,
) -> ModelConfiguration

Validate new-format durable configuration and reject unknown fields.

parse_run_inputs

parse_run_inputs(
    value: Mapping[str, object] | ModelRunInputs,
) -> ModelRunInputs

Validate new-format per-run inputs and reject unknown fields.

validate_output_selection

validate_output_selection(
    output: OutputSelection,
) -> OutputSelection

Reject output/reducer choices the plugin would otherwise silently ignore.

scoring_result_codec

scoring_result_codec() -> ScoringResultCodec

Return the binding-owned decoder for collected container outputs.

The default accepts a headered TSV containing variant_id plus exactly the manifest's primary result fields. Bindings override this when a runtime uses different source names or emits explicitly ignored transport fields. Storage adapters never learn either runtime-file shape.

detail_result_codec

detail_result_codec(name: str) -> DetailResultCodec

Return the binding-owned decoder for one collected container detail result.

The default accepts exact-schema TSV output. Bindings override this method when their runtime uses native column names, while the engine remains responsible for selecting only the logical files that the plan assigned to name.

save_configuration

save_configuration(
    configuration: ModelConfiguration,
) -> SavedConfiguration

Wrap validated configuration with the plugin/schema identity needed after a restart.

load_configuration

load_configuration(
    saved: SavedConfiguration | str,
) -> ModelConfiguration

Validate identity/version, migrate if needed, and rebuild typed configuration.

migrate_configuration

migrate_configuration(
    from_version: str, configuration: Mapping[str, object]
) -> Mapping[str, object]

Migration seam for older saved schema versions; plugins override when they add one.

convert_legacy_artifact

convert_legacy_artifact(
    artifact: Mapping[str, object], *, genome_build: str
) -> LegacyScoringConversion

Explicit compatibility seam; arbitrary old mappings are never accepted implicitly.

serialize_scoring_request

serialize_scoring_request(request: ScoringRequest) -> str

Serialize a validated request with configuration and plugin identity.

deserialize_scoring_request

deserialize_scoring_request(value: str) -> ScoringRequest

Rebuild a durable request after a process restart.

execute_inline async

execute_inline(
    plan: InlineScoringPlan,
    runtime: RuntimeContext | None = None,
) -> list[dict[str, Any]] | InlineScoringResult

Execute a durable inline descriptor using separately injected runtime state.

run_identity

run_identity(
    runtime_identities: Mapping[str, str],
    *,
    reducer: str | None = None,
) -> PluginRunIdentity

Snapshot plugin/output/runtime identity independently of downstream prioritization.

scoring_run_identity

scoring_run_identity(
    runtime_identities: Mapping[str, str],
    *,
    configuration: ModelConfiguration,
    output: OutputSelection,
    run_inputs: ModelRunInputs,
) -> PluginRunIdentity

Snapshot the exact model instance and scoring semantics for persisted results.

Durable model configuration (including weights/resources), the complete output selection, and genome inputs affect result meaning and therefore the cache identity. The method derives both genome build and structured genome identity from run_inputs so a binding cannot silently omit a supplied genome resource. The exact hash retains resource locations; the scientific projection prefers content digests and excludes typed credential references from compatibility. Per-run subjects, job/model IDs, batch boundaries, path resolvers, and runtime clients deliberately do not prevent incremental scoring into the same model result set.

build_scoring_plan

build_scoring_plan(request: ScoringRequest) -> ScoringPlan

Turn one (model, job) pair into a ScoringPlan — the runtime image, argv, and I/O for each shard.

The model-specific knowledge (the subcommand, the flag names, the per-fold path math) lives here, while the returned ContainerTaskSpec shards run on any ExecutionBackend. Paths come from request.layout (logical paths) combined with request.resolver (the mount root), so the plan is backend-agnostic: one plan runs unchanged locally or on Modal.

The default raises CapabilityNotSupported, since a results-only plugin has no plan. Container architectures override this to return a ContainerScoringPlan; inline ones such as AlphaGenome return an InlineScoringPlan.

build_interpretation_plan

build_interpretation_plan(
    request: InterpretationRequest,
) -> InterpretationPlan

Turn one (interpret-run, model) pair into a staged interpretation workflow.

The binding chooses the named stages and their dependencies. Model-specific commands and resources live in its backend-neutral container tasks; core does not assume folds, motifs, plots, or another particular interpretation shape.

The default raises CapabilityNotSupported; only container architectures with an interpretation pipeline override it.

build_preprocessing_plan

build_preprocessing_plan(
    request: PreprocessingRequest,
) -> PreprocessingPlan

Turn one model into an optional staged preparation workflow.

Preparation derives reusable model artifacts before scoring; it is not model training. The binding chooses its stage graph and optional provenance outputs. Scoring remains free to consume trusted, externally managed artifacts without an Altar-produced provenance manifest.

The default raises CapabilityNotSupported; only container architectures with a preprocessing step override it.

ModelResults

Bases: ABC

The results half of a model architecture: the methods the score store uses to build one wide row per variant.

These are split out from the compute methods so a caller can depend on only this narrower contract. The data-lake driver's ScoredModel.plugin is typed as ModelResults rather than the full plugin, so it can only call these methods and never reach a build_*.

score_columns defines the column schema. prioritize_predicate is the rule that flags a variant as prioritized, rendered to SQL or Python. to_model_score builds one model_scores JSONB entry. Both score_columns and to_model_score default to the manifest's result schema, so prioritize_predicate is the only results method a subclass must implement. prioritize is a plain-boolean convenience over prioritize_predicate; note that core's materialize path does not call it — it evaluates the predicate directly to keep the three-valued (Kleene) result (see prioritize below). A full ModelPlugin adds the compute methods on top of this base.

validate_result_identity

validate_result_identity(
    identity: PluginRunIdentity,
) -> None

Fail before this results implementation decodes an incompatible primary score row.

Prioritization is applied after raw scores are read, so its exact expression remains provenance but is not part of primary-score compatibility. A changed policy can rematerialize the same raw scores. Runtime, configuration, output-selection, reducer, and genome compatibility belong to the concrete persisted result and are therefore checked by the score store, not against this wrapper's defaults.

score_columns

score_columns() -> list[ScoreColumn]

Return the per-variant score columns, derived from the manifest's result schema.

The manifest's ordered ResultSchema is the single source for the column schema. A ScoreColumn is a ResultField without its biological meaning, and ModelResultContract and the scoring engine both require the two to match exactly, so the default is the only correct shape for a manifest-declared model. Override it only when the columns come from somewhere other than manifest.result_schema.

prioritize_predicate abstractmethod

prioritize_predicate() -> Predicate

The prioritization rule as a Predicate. It renders to SQL, so the score store can filter with it, or to Python for the reference store. Its columns may reference both score columns and annotation columns.

to_model_score

to_model_score(
    *, model_id: str, model_name: str, score: VariantRecord
) -> dict[str, Any]

Build one model_scores JSONB entry from this model's per-variant score mapping.

The entry's keys are model_id, model_name, one key per score_columns() entry in order, and prioritized. The default decodes each score value as its column's declared dtype. A missing, None, or NaN value becomes None for every dtype: the default never turns an absent measurement into False, 0, or an empty string, so a store can still tell "not measured" from a measured negative.

Decoding is strict, so a schema mismatch fails instead of producing a plausible wrong value. A NumPy-style scalar (from a pandas frame) is unwrapped first. float accepts integers and floats but not booleans. int accepts integers and whole floats, because a frame stores an integer column with missing values as float, and rejects 3.7. bool accepts booleans and the integers 0 and 1, which is how SQLite stores a flag. str accepts only strings. Any other value raises TypeError or ValueError naming the column.

prioritized follows a different rule because it is the materialized per-model decision, not a score. The materialize path passes the predicate's three-valued result, and a None (a required input was absent) becomes False. That matches prioritize() and the warehouse export's COALESCE(is_prioritized, FALSE), so both materialization paths produce the same entry. A present value is decoded with the strict bool rule.

Override this only when a stored value needs a model-specific decoding. Changing how a stored value is decoded is a breaking result change (see the plugin ABI guide).

prioritize

prioritize(
    score: VariantRecord, annotation: VariantRecord
) -> bool

Evaluate prioritize_predicate against the merged score and annotation values as a plain boolean.

Score and annotation names must be disjoint. A collision raises SchemaCollisionError instead of silently choosing one producer's value.

This is a convenience for callers who want a definite yes/no. It goes through to_python(), which collapses the predicate's three-valued (Kleene) result — a missing/None column yields None — down to False, matching a SQL boolean filter where only a definite True counts.

Core's materialize path does not call this. It evaluates the predicate with prioritize_predicate().evaluate(...) (see data_lake/materialize.py), which returns the raw True/False/None, then threads that value into to_model_score and aggregates the row-level flag with an is True test. The net prioritized outcome matches prioritize() today — to_model_score performs the same None→False collapse, and None is True is False — so the two do not disagree; the difference is only where the collapse happens. Keeping the Kleene value until to_model_score leaves a to_model_score implementation free to treat "the rule did not fire" (False) and "a required input was absent" (None) differently. Prefer evaluate when that distinction matters; prioritize is the shortcut when it does not.

PreprocessingRequest dataclass

PreprocessingRequest(
    model_id: str,
    configuration: ModelConfiguration,
    run_inputs: ModelRunInputs,
    resolver: PathResolver,
)

The inputs build_preprocessing_plan needs to turn one model into a PreprocessingPlan.

Like ScoringRequest, it is backend-agnostic: the binding names base-relative logical paths and the backend PathResolver supplies the mount root. Derived artifacts are architecture-specific, so the binding owns its logical layout; core carries no preprocessing layout. Staged file names come from the typed configuration, with no runtime directory listing.

ScoreColumn dataclass

ScoreColumn(
    name: str,
    dtype: ScalarDtype,
    label: str,
    nullable: bool = True,
)

One per-variant score column a model plugin produces.

Declaring the column here sets four things that must otherwise be kept in sync by hand: the score-store schema (table columns / DDL), the write-time validator, the wide-TSV column names, and the frontend column definitions.

It lives in this shared leaf module — not in model/plugin.py — because it is the schema the score store consumes (ScoreStore.ensure_schema / add_scores); keeping it here lets the data-lake axis name it without importing back into the model axis. model.plugin re-exports it for architecture authors.

ScoringRequest dataclass

ScoringRequest(
    job_id: str,
    model_id: str,
    configuration: ModelConfiguration,
    run_inputs: ModelRunInputs,
    resolver: PathResolver,
    output: OutputSelection,
    scoring_batch: int | None = None,
    scoring_batch_count: int | None = None,
    layout: ScoringLayout = ScoringLayout(),
)

The inputs build_scoring_plan needs to turn one (model, job) pair into a ScoringPlan.

Paths are described independent of any backend, so a plugin can build the plan for any backend: they come from a ScoringLayout (a base-relative logical layout) combined with the backend's PathResolver (the mount root), never a concrete filesystem. configuration holds durable model/resource identity; run_inputs holds per-run biology; output holds reducer/output selection. Runtime-injected clients and callables are supplied separately and cannot enter this request.

genome_label property

genome_label: str

Compatibility spelling for the typed run input's genome build.

DelimitedResultCodec dataclass

DelimitedResultCodec(
    schema: ResultSchema,
    source_names: Mapping[str, str] = dict(),
    row_identity_fields: Mapping[str, str] = dict(),
    allowed_extra_fields: tuple[str, ...] = (),
    identity_field: str = "variant_id",
    delimiter: str = "\t",
    null_tokens: tuple[str, ...] = ("", "NA", "NaN", "nan"),
    optional_extra_fields: tuple[str, ...] = (),
)

Strict decoder for a headered delimited primary-result file.

source_names maps each declared result name to its runtime-file spelling. row_identity_fields maps transport columns such as chr/pos/ref/alt to scalar dtypes and preserves them for storage adapters that need decomposed loci. allowed_extra_fields names required runtime columns that are intentionally ignored; optional_extra_fields names ignored diagnostics accepted when present. Every other undeclared column is rejected. This makes both preservation and projection explicit binding decisions instead of silently dropping new runtime output.

decode async

decode(
    storage: Storage,
    output_uris: Sequence[str],
    *,
    batch_size: int = 10000,
) -> AsyncIterator[ResultBatch]

Stream every collected file in order and validate its complete input schema.

ScoringResultCodec

Bases: Protocol

Decode collected runtime outputs into bounded, schema-shaped primary batches.

decode

decode(
    storage: Storage,
    output_uris: Sequence[str],
    *,
    batch_size: int = 10000,
) -> AsyncIterator[ResultBatch]

Yield decoded batches without loading a complete output into memory.

ModelArtifactLayout dataclass

ModelArtifactLayout()

Paths to shared, model-owned inputs: the staged genome FASTA and files under models/{model_id}.

These live at the storage root under genomes/ and models/, not inside any job directory. Paths are POSIX and relative to the storage root; the backend's PathResolver supplies the mount root. A binding that needs architecture-specific model files (derived calibration data, an interpretation motif set, and so on) names them with model_resource_file or subclasses this layout in its own package.

model_resource_file

model_resource_file(model_id: str, filename: str) -> str

Return a fixed model-owned resource path, models/{model_id}/{filename}.

Bindings use this for content-addressed weight bundles, track catalogs, derived model artifacts, and similar resources. They supply a constant basename chosen from declared configuration rather than a provider path or URI suffix; rejecting path components keeps the logical layout portable and prevents traversal.

fold_model_file

fold_model_file(model_id: str, fold: int, ext: str) -> str

Path to one cross-validation ensemble member's weights, models/{model_id}/fold_{fold}_model{ext}.

ext includes the leading dot, e.g. .h5. Derive it from the binding's declared weight format, never from a resource URI: a content-addressed URI need not carry a file extension.

ScoringLayout dataclass

ScoringLayout()

Bases: ModelArtifactLayout

Model-neutral paths for one container scoring run, one method per file a scoring plan names.

Adds the scoring job's inputs and outputs on top of the shared genome and model-file paths from ModelArtifactLayout. ScoringRequest.layout defaults to this class; the scoring engine and the portable preparation path read the variant input and result paths from it.

ineligible_variants_file

ineligible_variants_file(job_id: str, model_id: str) -> str

Return the report of candidates that batch preparation excluded as ineligible for this model.

fold_score_file

fold_score_file(
    job_id: str,
    model_id: str,
    fold: int,
    scoring_batch: int | None = None,
) -> str

Return one ensemble member's intermediate score file, reduced by the plan's summary task.

scoring_result_file

scoring_result_file(
    job_id: str,
    model_id: str,
    scoring_batch: int | None = None,
) -> str

Return the logical path of the primary result table the scoring engine collects.

detail_result_file

detail_result_file(
    job_id: str,
    model_id: str,
    name: str,
    scoring_batch: int | None = None,
) -> str

Return the logical TSV path for one manifest-declared named detail result.

ModelWorkflowPlan dataclass

ModelWorkflowPlan(
    operation: str,
    stages: tuple[WorkflowStage, ...],
    plugin_identity: PluginRunIdentity,
    outputs: tuple[Transfer, ...] = (),
)

A backend-neutral DAG for one optional model operation.

operation is a stable plugin-facing discriminator such as "preparation" or "interpretation". Each stage runs all of its tasks in parallel after every stage in needs has completed. outputs names the terminal artifacts the caller wants copied back through the backend. Intermediate products may remain on the backend's shared working volume.

tasks property

tasks: tuple[ContainerTaskSpec, ...]

Return every task in declaration order.

stage

stage(name: str) -> WorkflowStage

Return one named stage or raise a clear lookup error.

WorkflowExecutionError

Bases: RuntimeError

A model-workflow task failed on its execution backend.

WorkflowRunResult dataclass

WorkflowRunResult(
    plan: ModelWorkflowPlan,
    completed_stages: tuple[str, ...],
    output_uris: tuple[str, ...],
)

The completed stages and collected terminal artifact locations for one workflow.

WorkflowStage dataclass

WorkflowStage(
    name: str,
    tasks: tuple[ContainerTaskSpec, ...],
    needs: tuple[str, ...] = (),
)

One parallel task set whose declared prerequisite stages must finish first.

canonical_chromosome

canonical_chromosome(chromosome: str) -> str

Return the canonical chr-prefixed spelling used in portable keys.

Primary chromosomes accept common aliases (1/chr1/CHR01 and M/MT/chrM) in any case. Other reference contigs keep their exact spelling after an optional case-insensitive chr prefix is normalized, so CHRUn_KI270302v1 becomes chrUn_KI270302v1 but chrun_KI270302v1 stays as written. Non-ASCII text, whitespace inside the name, and characters outside the VCF contig-name set (including :) raise VariantIdentityError. This is a naming rule only; it does not claim that contigs from different assemblies are interchangeable.

canonical_variant_id

canonical_variant_id(
    chromosome: str,
    position: int,
    reference_allele: str,
    alternate_allele: str,
) -> str

Return the canonical text key for decomposed one-based locus fields.

canonical_json

canonical_json(value: Any) -> str

Return canonical UTF-8 JSON text: sorted object keys, compact separators, no NaN.

classify_variant

classify_variant(key: VariantKey) -> VariantClass

Return the :class:VariantClass of key from its REF and ALT lengths.

Classification trusts that both alleles are literal bases. A placeholder allele such as - or * would count as one base; rejecting those is :class:VariantKey's job, not this function's.

deterministic_hash

deterministic_hash(value: Any) -> str

Return a lowercase SHA-256 identity over :func:canonical_json.

validate_result_identity

validate_result_identity(
    manifest: PluginManifest, identity: PluginRunIdentity
) -> None

Reject a stored primary result before the current plugin assigns it new meaning.

This validates the binding-owned portion of primary-score compatibility. Concrete model configuration, runtime, output, reducer, and genome are compared by stores through :class:ResultCompatibility. Prioritization belongs to materialization rather than score identity.

require_resource_digests

require_resource_digests(
    value: ConfigurationModel, *, owner: str
) -> None

Raise ValueError unless every resource on value carries a SHA-256 digest.

Checks each top-level field whose value is a ResourceReference or a tuple or list of them; an unset optional resource (None) is skipped. Nested configuration models and mappings are not traversed, so a configuration that nests resources must also check the nested model. Missing resources are named in field order, with an index for a sequence (weights[2]), and owner prefixes the message. Call it from a model_validator(mode="after") on a configuration whose runtime verifies staged bytes, so Pydantic reports the failure as a validation error before a plan can be built.

get_model_plugin

get_model_plugin(model_type: str) -> ModelPlugin

Look up and instantiate the ModelPlugin for a model architecture, keyed by its model_type string (the registry key, e.g. "CHROMBPNET").

A call site holds a model's model_type and asks for its plugin without naming a concrete class, which is what lets the scoring, interpretation, and preprocessing flows work with any architecture. Adding an architecture then only requires registering it. The registry stores the plugin class (an entry point loads the class object); this instantiates it, and plugins take no constructor arguments. Raises KeyError, listing the registered names, if model_type has no plugin.

get_model_plugin_registry

get_model_plugin_registry() -> PluginRegistry[
    type[ModelPlugin]
]

Return the shared registry of model-architecture plugins. It stores ModelPlugin subclasses, so .get(model_type) returns a type[ModelPlugin].

resource_transfer

resource_transfer(
    resource: ResourceReference,
    logical_path: str,
    *,
    locality_key: str | None = None,
) -> Transfer

Declare a configured resource as a container input staged at logical_path.

The transfer carries the resource's uri as its location and its digest as the content identity the staging backend verifies, so relocating identical bytes changes neither the plan's meaning nor its cache identity. revision and subtype selectors stay in the configuration's scientific identity; a storage adapter reads only the location and the digest. locality_key is the optional co-scheduling hint described on Transfer, typically the model ID for model-owned resources.

scoring_plan_from_dict

scoring_plan_from_dict(
    data: Mapping[str, object],
) -> ScoringPlan

Rebuild a scoring plan, rejecting unknown envelope fields.

scoring_plan_from_json

scoring_plan_from_json(value: str) -> ScoringPlan

Rebuild a scoring plan after a process restart.

scoring_plan_to_dict

scoring_plan_to_dict(
    plan: ScoringPlan,
) -> dict[str, JsonValue]

Serialize either scoring-plan shape to JSON-compatible values.

scoring_plan_to_json

scoring_plan_to_json(plan: ScoringPlan) -> str

Serialize a durable scoring plan to JSON.

run_model_workflow async

run_model_workflow(
    plan: ModelWorkflowPlan,
    *,
    backend: ExecutionBackend,
    input_transfer: TransferMapper | None = None,
    output_transfer: TransferMapper | None = None,
    spec_mapper: SpecMapper | None = None,
    poll_interval_s: float = 1.0,
    max_ticks_per_wave: int = 3600,
) -> WorkflowRunResult

Execute all ready stages in dependency waves and collect declared terminal outputs.

workflow_plan_from_dict

workflow_plan_from_dict(
    data: Mapping[str, object],
) -> ModelWorkflowPlan

Rebuild and validate a workflow plan from durable values.

workflow_plan_from_json

workflow_plan_from_json(value: str) -> ModelWorkflowPlan

Rebuild a durable model workflow plan from JSON.

workflow_plan_to_dict

workflow_plan_to_dict(
    plan: ModelWorkflowPlan,
) -> dict[str, Any]

Serialize a workflow plan to JSON-compatible values.

workflow_plan_to_json

workflow_plan_to_json(plan: ModelWorkflowPlan) -> str

Serialize a durable model workflow plan to canonical JSON.