Models and plans¶
Model binding authors should import from altar.models. This stable façade exposes model manifests, strict
configuration and run-input bases, typed resource and credential references, saved configuration envelopes,
durable plan serializers, model plugins, and logical path layouts without importing a model SDK. Importing it
also loads no storage adapter or concrete execution backend; those come from altar.results, altar.sources,
and altar.execution.
models
¶
Stable model-binding contracts.
External model bindings should import from this module. Names re-exported here are covered by Altar's
public compatibility policy; their original altar.plugins.model locations are implementation paths.
VariantIdentityError
¶
Bases: ValueError
A variant key is malformed, non-canonical, or disagrees with its locus.
VariantKey
dataclass
¶
Canonical one-based biallelic key serialized as chr:pos:ref:alt.
Fields follow the rules in this module's documentation. The class normalizes spelling but does not normalize biological representation. Callers must left-align and trim indels against the selected reference before constructing a key when equivalence across representations matters.
parse
classmethod
¶
parse(variant_id: str) -> VariantKey
Parse and normalize a four-field textual key.
require_canonical
classmethod
¶
require_canonical(variant_id: str) -> VariantKey
Parse variant_id and reject aliases instead of silently rewriting them.
from_fields
classmethod
¶
from_fields(
chromosome: str,
position: int,
reference_allele: str,
alternate_allele: str,
) -> VariantKey
Construct a canonical key from decomposed one-based locus fields.
CapabilityNotSupported
¶
Bases: NotImplementedError
Raised by an optional plugin method that the adapter deliberately does not implement.
It subclasses NotImplementedError, so an existing except NotImplementedError still catches it. It is a
named, catchable signal distinct from a genuinely missing abstract method: a caller probing an optional
capability can except CapabilityNotSupported without also swallowing real not-yet-implemented bugs.
AnnotationDependency
dataclass
¶
An annotation contract required by the prioritization policy.
ApiVersionRange
dataclass
¶
A half-open Altar API interval: minimum <= version < maximum.
DetailSchema
dataclass
¶
DetailSchema(
name: str,
version: str,
description: str,
fields: tuple[ResultField, ...],
key_fields: tuple[str, ...],
)
One versioned, named one-to-many result table owned by a model binding.
variant_id is the implicit parent key and must not be repeated in fields. key_fields names
the binding-declared fields that, together with variant_id, identify one logical detail row. Key
values may be null when the upstream science has no value (for example a non-gene-centric track has no
gene ID); stores use Altar's canonical logical-key encoding rather than backend-specific NULL equality.
schema_hash
property
¶
Hash row identity and biological meaning, excluding descriptions and display labels.
IncompatibleAltarApiError
¶
Bases: ManifestError
Raised when a plugin does not support the running Altar API.
IncompatibleResultError
¶
Bases: ManifestError
Raised before results are decoded with an incompatible plugin contract.
IneligibilityReason
¶
Bases: StrEnum
Why a model's :class:VariantEligibility excludes a variant.
ManifestError
¶
Bases: ValueError
Base class for invalid or incompatible plugin ABI data.
PluginManifest
dataclass
¶
PluginManifest(
plugin_id: str,
kind: PluginKind,
capabilities: frozenset[PluginCapability],
architecture: str | None,
plugin_version: str,
altar_api: ApiVersionRange,
configuration_schema_version: str,
result_schema: ResultSchema,
prioritization: PrioritizationPolicy,
runtimes: tuple[RuntimeRequirement, ...],
detail_schemas: tuple[DetailSchema, ...] = (),
variant_eligibility: VariantEligibility = VariantEligibility(),
manifest_schema_version: str = MANIFEST_SCHEMA_VERSION,
)
The immutable declaration shipped by one plugin release.
detail_schema
¶
detail_schema(name: str) -> DetailSchema
Return one declared detail schema by name or fail with the available names.
PluginRunIdentity
dataclass
¶
PluginRunIdentity(
plugin_id: str,
plugin_kind: PluginKind,
architecture: str | None,
plugin_version: str,
supported_altar_api: ApiVersionRange,
altar_api_version: str,
result_schema_hash: str,
runtimes: tuple[RuntimeIdentity, ...],
reducer: str | None = None,
configuration_hash: str | None = None,
output_selection_hash: str | None = None,
genome_build: str | None = None,
scientific_configuration_hash: str | None = None,
genome_hash: str | None = None,
)
Transport-safe exact identity attached to a run and its result set.
The optional scoring-context fields distinguish concrete model instances, not merely wrapper ABIs.
configuration_hash preserves the complete durable configuration for exact provenance, while
scientific_configuration_hash may omit typed operational credentials for cache compatibility.
genome_hash records the structured genome's scientific projection separately so an engine can verify
that a plan was built from the request it is about to execute. These fields are optional when inspecting
a plugin before configuring a concrete model, but persistence requires :meth:assert_scoring_context.
Prioritization, annotation dependencies, schema release labels, and the full manifest hash deliberately
stay out of this identity: changing downstream selection must not invalidate raw scores.
assert_altar_compatible
¶
Reject executing this recorded plan under an API its wrapper did not support.
assert_scoring_context
¶
Reject unconfigured identities that do not identify weights, outputs, and genome.
primary_result_compatibility
¶
primary_result_compatibility() -> ResultCompatibility
Return the scientific cache identity for the primary scalar score table.
scientific_dict
¶
Return the projection that keys and defines a generated score lineage.
A lineage identifies what can change a primary score: the plugin release and its exact runtimes,
the scientific configuration projection, the output selection, the genome build and the genome's
scientific hash, and the primary result schema. The exact configuration_hash stays out
because it also covers resource locations and every declared serialization detail, so a schema
migration that adds defaulted fields, or unchanged bytes served from a new location, would
otherwise orphan every cached score. The framework API versions stay out for the same reason.
Two runs whose projections are equal share one lineage ID, and a lineage store accepts either run
for that lineage (ScoreLineage.same_lineage), so each reuses the other's scores. Exact provenance
is not lost: the catalog keeps one deterministic representative identity, and each distinct exact
identity that registers to write scores under the lineage is recorded once in the store's lineage
provenance table, read back as a set with ScoreStore.read_lineage_provenance.
detail_result_compatibility
¶
detail_result_compatibility(
schema: DetailSchema,
) -> ResultCompatibility
Return the scientific cache identity for one named detail result table.
ResultCompatibility
dataclass
¶
ResultCompatibility(
role: ResultRole,
plugin_id: str,
plugin_kind: PluginKind,
architecture: str | None,
plugin_version: str,
runtimes: tuple[RuntimeIdentity, ...],
reducer: str | None,
configuration_hash: str,
output_selection_hash: str,
genome_build: str,
schema_name: str,
schema_hash: str,
)
Role-specific scientific identity used to decide whether persisted results may be reused.
Plugin releases and exact runtimes distinguish computations. Reusing a previous release requires an explicit score-lineage acceptance; version numbers alone never establish scientific equivalence. Framework compatibility and display metadata are not scientific cache gates. Prioritization is not part of score identity at all.
fingerprint
property
¶
Return the canonical role-specific compatibility fingerprint.
ResultField
dataclass
¶
ResultField(
name: str,
dtype: Literal["float", "bool", "int", "str"],
label: str,
meaning: str,
nullable: bool = True,
)
One ordered scalar result field, including its durable biological meaning.
ResultRole
¶
Bases: StrEnum
The persisted result role whose compatibility is being evaluated.
ResultSchema
dataclass
¶
ResultSchema(version: str, fields: tuple[ResultField, ...])
A versioned ordered schema with a biological-compatibility hash.
The complete schema, including display labels and its version, remains in the manifest for provenance.
schema_hash deliberately covers only fields that change how stored values are interpreted.
schema_hash
property
¶
Hash ordered storage and biological meaning, excluding presentation labels.
RuntimeIdentityError
¶
Bases: ManifestError
Raised when a run does not satisfy its manifest's runtime identity policy.
RuntimeIdentityPolicy
¶
Bases: StrEnum
How an actual runtime value is checked and recorded for a run.
RuntimeRequirement
dataclass
¶
RuntimeRequirement(
name: str,
kind: RuntimeKind,
policy: RuntimeIdentityPolicy,
declared_identity: str | None = None,
)
One named runtime role and the policy for its per-run identity.
VariantClass
¶
Bases: StrEnum
The allele-length class of one canonical biallelic variant.
Classification compares REF and ALT lengths exactly as written. It does not trim shared bases or
left-align, so both an anchored VCF insertion (A>AT) and a delins that lengthens the sequence
(A>GT) are insertions: a sequence model substitutes the whole ALT for the whole REF either way.
SNV
class-attribute
instance-attribute
¶
One reference base replaced by one different base.
VariantEligibility
dataclass
¶
VariantEligibility(
variant_classes: frozenset[VariantClass] = (
lambda: frozenset(VariantClass)
)(),
max_allele_length: int | None = None,
)
The variants a model can score, declared once in its manifest.
variant_classes lists the admitted :class:VariantClass values. max_allele_length, when set,
bounds the longer of REF and ALT in bases. The default admits every class at any length, which is
how a manifest that declares nothing behaves. Batch preparation excludes and reports ineligible
variants before cache lookup, and the scoring engine rejects result rows for them.
Eligibility filters inputs; it is not a score-changing setting. The score of an eligible variant does
not depend on it, so it is excluded from :class:PluginRunIdentity and score lineage.
is_unrestricted
property
¶
Whether every variant class at every allele length is eligible.
ineligibility
¶
ineligibility(
key: VariantKey,
) -> IneligibilityReason | None
Return why key is ineligible, or None when the model can score it.
Version
dataclass
¶
A deliberately small SemVer core used by the ABI (no prerelease syntax).
ConfigurationError
¶
Bases: ValueError
Base error for durable model configuration.
ConfigurationIdentityError
¶
Bases: ConfigurationError
A saved configuration belongs to a different plugin or schema.
ConfigurationManifest
dataclass
¶
ConfigurationManifest(
plugin_id: str,
plugin_version: str,
schema_id: str,
schema_version: str,
configuration_model: type[ModelConfiguration],
run_input_model: type[ModelRunInputs],
)
Plugin and schema identity attached to every saved configuration.
ConfigurationModel
¶
Bases: BaseModel
Strict, frozen, JSON-serializable base for Altar's durable public values.
Unknown fields are forbidden and defaults are validated. The wrapper methods are the supported public serialization/schema API; Pydantic remains an implementation detail behind them.
ConfigurationVersionError
¶
Bases: ConfigurationError
A saved configuration needs a migration that the plugin does not provide.
EmptyModelConfiguration
¶
LegacyArtifactError
¶
Bases: ConfigurationError
A legacy artifact mapping cannot be converted safely.
LegacyScoringConversion
dataclass
¶
LegacyScoringConversion(
configuration: ModelConfiguration,
run_inputs: ModelRunInputs,
output: OutputSelection,
runtime: RuntimeContext | None = None,
)
Explicit result of converting a supported legacy artifact mapping.
ModelConfiguration
¶
Bases: ConfigurationModel
Durable, architecture-specific model configuration.
to_dict() is the complete saved configuration. scientific_dict() is the cache-identity view:
fields typed to contain :class:SecretReference are operational and excluded recursively, while every
other durable value remains identity-bearing.
scientific_dict
¶
Return the scientific cache-identity projection of this configuration.
A field whose annotation contains SecretReference is wholly operational. This type-driven rule
applies even when its current value is None, so adding, removing, or rotating a credential does
not change scientific identity. Nested configuration models are traversed recursively.
ModelRunInputs
¶
Bases: ConfigurationModel
Base per-run biological inputs shared by scoring plugins.
genome_build is the readable assembly label used in paths and diagnostics. genome records the
optional reference resource for integrations that can identify one. Models that read local reference
sequence bytes should inherit :class:ReferenceGenomeRunInputs, which requires a content address.
NonDurableRuntimeError
¶
Bases: TypeError
A caller attempted to serialize runtime-only injected state.
OutputSelection
¶
Bases: ConfigurationModel
Requested named outputs and reducer, separate from model and biological inputs.
ReferenceGenomeRunInputs
¶
Bases: ModelRunInputs
Per-run inputs for models whose results depend on staged reference FASTA bytes.
ResourceReference
¶
Bases: ConfigurationModel
Durable identity for a model, weights, genome, or other resource.
digest is a content address and takes precedence over location metadata in scientific cache identity.
It addresses the resource bytes as narrowed by any subtype selectors; a selector covering multiple files
therefore needs a bundle or snapshot digest. Every additional subtype field is identity-bearing, so keep
operational location metadata in uri rather than adding fields such as region or endpoint. If no
digest is available, an immutable provider revision plus URI is the next-best identity. A bare URI remains
supported for legacy inputs, but identifies only the name, not the bytes later served there.
scientific_dict
¶
Return the strongest identity plus any subtype-specific selectors.
A digest identifies the selector-narrowed resource without its location. Revision-based identities retain the URI, and subclasses retain identity-bearing selectors such as a filename within a versioned repository.
RuntimeContext
¶
Explicitly non-durable injected clients/functions used only while executing a plan.
SavedConfiguration
¶
SecretReference
¶
Bases: ConfigurationModel
Durable pointer to a credential; it never contains the credential value itself.
SubjectSelection
¶
ContainerResultFiles
dataclass
¶
Collected native files for the primary table or one named detail table.
name is "primary" or a manifest-declared :class:DetailSchema name. Logical paths refer to
terminal task outputs before a deployment maps their storage URIs. Keeping this classification in the
portable plan prevents result codecs from guessing meaning from provider-specific paths or filenames.
ContainerScoringPlan
dataclass
¶
ContainerScoringPlan(
shards: list[ContainerTaskSpec],
plugin_identity: PluginRunIdentity,
summarize: ContainerTaskSpec | None = None,
ready_when: int = 1,
result_files: tuple[ContainerResultFiles, ...] = (),
)
A scoring run expressed as N container shards, plus an optional fan-in summarize shard.
This is the container form of ScoringPlan, used by architectures like ChromBPNet, ProCapNet, and
Cherimoya. Each shard is a ContainerTaskSpec that runs on any ExecutionBackend (local, Modal, or
Kubernetes), so the same plan scores on any of them. ready_when is the number of shards to wait for
before the summarize shard runs; it equals ModelPlugin.num_folds. result_files assigns collected
terminal logical paths to the primary table or manifest-declared named details. It may be omitted only
for the backward-compatible primary-only shape, where every terminal output is primary.
InlineScoringPlan
dataclass
¶
InlineScoringPlan(
plugin_identity: PluginRunIdentity,
executor: str,
payload: dict[str, JsonValue],
)
A durable inline execution descriptor, such as an AlphaGenome API scoring request.
The plan contains only JSON data and an executor discriminator. Runtime clients, secret resolvers, and
injected callables are supplied separately to ModelPlugin.execute_inline and never become plan state.
InterpretationRequest
dataclass
¶
InterpretationRequest(
run_id: str,
model_id: str,
configuration: ModelConfiguration,
run_inputs: ModelRunInputs,
resolver: PathResolver,
)
The inputs build_interpretation_plan needs to turn one (interpret-run, model) pair into an
InterpretationPlan. Like ScoringRequest, it is backend-agnostic: the binding names base-relative
logical paths and the backend PathResolver supplies the mount root. Interpretation outputs are
architecture-specific, so the binding owns its logical layout; core carries no interpretation layout.
run_id is the interpretation run's UUID, not a job id.
ModelPlugin
¶
Bases: ModelResults
One model architecture's full behavior: the results methods inherited from ModelResults, plus the
optional compute methods num_folds and the build_* plan builders. model_type remains the registry
key and model-family discriminator; manifest is the durable versioned identity. This is the class an
architecture author subclasses and registers.
model_type equals the model's stored architecture type (for example "CHROMBPNET"). Every other
plugin axis keys on a name discriminator; this one keeps model_type because the value is stored and
used elsewhere and cannot be renamed freely.
The compute methods are optional. By default they raise CapabilityNotSupported, so a plugin whose
scores are produced elsewhere can implement none of them.
Validate Altar API compatibility before this plugin can execute or interpret data.
num_folds
property
¶
Number of container shards one scoring run fans into (ChromBPNet uses 5; a single-shard or inline
model uses 1). This sets ContainerScoringPlan.ready_when and the number of shards the completion
fan-in waits for. Defaults to 1.
configuration_manifest
classmethod
¶
configuration_manifest() -> ConfigurationManifest
Return stable plugin/schema identity plus the two discoverable public models.
configuration_schema
classmethod
¶
Return this plugin's durable-configuration JSON Schema without loading a model SDK.
run_input_schema
classmethod
¶
Return this plugin's per-run biological-input JSON Schema.
parse_configuration
¶
parse_configuration(
value: Mapping[str, object] | ModelConfiguration,
) -> ModelConfiguration
Validate new-format durable configuration and reject unknown fields.
parse_run_inputs
¶
parse_run_inputs(
value: Mapping[str, object] | ModelRunInputs,
) -> ModelRunInputs
Validate new-format per-run inputs and reject unknown fields.
validate_output_selection
¶
validate_output_selection(
output: OutputSelection,
) -> OutputSelection
Reject output/reducer choices the plugin would otherwise silently ignore.
scoring_result_codec
¶
scoring_result_codec() -> ScoringResultCodec
Return the binding-owned decoder for collected container outputs.
The default accepts a headered TSV containing variant_id plus exactly the manifest's primary
result fields. Bindings override this when a runtime uses different source names or emits explicitly
ignored transport fields. Storage adapters never learn either runtime-file shape.
detail_result_codec
¶
detail_result_codec(name: str) -> DetailResultCodec
Return the binding-owned decoder for one collected container detail result.
The default accepts exact-schema TSV output. Bindings override this method when their runtime uses
native column names, while the engine remains responsible for selecting only the logical files that
the plan assigned to name.
save_configuration
¶
save_configuration(
configuration: ModelConfiguration,
) -> SavedConfiguration
Wrap validated configuration with the plugin/schema identity needed after a restart.
load_configuration
¶
load_configuration(
saved: SavedConfiguration | str,
) -> ModelConfiguration
Validate identity/version, migrate if needed, and rebuild typed configuration.
migrate_configuration
¶
migrate_configuration(
from_version: str, configuration: Mapping[str, object]
) -> Mapping[str, object]
Migration seam for older saved schema versions; plugins override when they add one.
convert_legacy_artifact
¶
convert_legacy_artifact(
artifact: Mapping[str, object], *, genome_build: str
) -> LegacyScoringConversion
Explicit compatibility seam; arbitrary old mappings are never accepted implicitly.
serialize_scoring_request
¶
serialize_scoring_request(request: ScoringRequest) -> str
Serialize a validated request with configuration and plugin identity.
deserialize_scoring_request
¶
deserialize_scoring_request(value: str) -> ScoringRequest
Rebuild a durable request after a process restart.
execute_inline
async
¶
execute_inline(
plan: InlineScoringPlan,
runtime: RuntimeContext | None = None,
) -> list[dict[str, Any]] | InlineScoringResult
Execute a durable inline descriptor using separately injected runtime state.
run_identity
¶
run_identity(
runtime_identities: Mapping[str, str],
*,
reducer: str | None = None,
) -> PluginRunIdentity
Snapshot plugin/output/runtime identity independently of downstream prioritization.
scoring_run_identity
¶
scoring_run_identity(
runtime_identities: Mapping[str, str],
*,
configuration: ModelConfiguration,
output: OutputSelection,
run_inputs: ModelRunInputs,
) -> PluginRunIdentity
Snapshot the exact model instance and scoring semantics for persisted results.
Durable model configuration (including weights/resources), the complete output selection, and genome
inputs affect result meaning and therefore the cache identity. The method derives both genome build
and structured genome identity from run_inputs so a binding cannot silently omit a supplied
genome resource. The exact hash retains resource locations; the scientific projection prefers
content digests and excludes typed credential references from compatibility.
Per-run subjects, job/model IDs, batch boundaries, path resolvers, and runtime clients deliberately
do not prevent incremental scoring into the same model result set.
build_scoring_plan
¶
build_scoring_plan(request: ScoringRequest) -> ScoringPlan
Turn one (model, job) pair into a ScoringPlan — the runtime image, argv, and I/O for each shard.
The model-specific knowledge (the subcommand, the flag names, the per-fold path math) lives here,
while the returned ContainerTaskSpec shards run on any ExecutionBackend. Paths come from
request.layout (logical paths) combined with request.resolver (the mount root), so the plan is
backend-agnostic: one plan runs unchanged locally or on Modal.
The default raises CapabilityNotSupported, since a results-only plugin has no plan. Container
architectures override this to return a ContainerScoringPlan; inline ones such as AlphaGenome
return an InlineScoringPlan.
build_interpretation_plan
¶
build_interpretation_plan(
request: InterpretationRequest,
) -> InterpretationPlan
Turn one (interpret-run, model) pair into a staged interpretation workflow.
The binding chooses the named stages and their dependencies. Model-specific commands and resources live in its backend-neutral container tasks; core does not assume folds, motifs, plots, or another particular interpretation shape.
The default raises CapabilityNotSupported; only container architectures with an interpretation
pipeline override it.
build_preprocessing_plan
¶
build_preprocessing_plan(
request: PreprocessingRequest,
) -> PreprocessingPlan
Turn one model into an optional staged preparation workflow.
Preparation derives reusable model artifacts before scoring; it is not model training. The binding chooses its stage graph and optional provenance outputs. Scoring remains free to consume trusted, externally managed artifacts without an Altar-produced provenance manifest.
The default raises CapabilityNotSupported; only container architectures with a preprocessing step
override it.
ModelResults
¶
Bases: ABC
The results half of a model architecture: the methods the score store uses to build one wide row per variant.
These are split out from the compute methods so a caller can depend on only this narrower contract. The
data-lake driver's ScoredModel.plugin is typed as ModelResults rather than the full plugin, so it
can only call these methods and never reach a build_*.
score_columns defines the column schema. prioritize_predicate is the rule that flags a variant as
prioritized, rendered to SQL or Python. to_model_score builds one model_scores JSONB entry. Both
score_columns and to_model_score default to the manifest's result schema, so prioritize_predicate
is the only results method a subclass must implement. prioritize is a plain-boolean convenience over
prioritize_predicate; note that core's materialize path does not call it — it evaluates the
predicate directly to keep the three-valued (Kleene) result (see prioritize below). A full
ModelPlugin adds the compute methods on top of this base.
validate_result_identity
¶
validate_result_identity(
identity: PluginRunIdentity,
) -> None
Fail before this results implementation decodes an incompatible primary score row.
Prioritization is applied after raw scores are read, so its exact expression remains provenance but is not part of primary-score compatibility. A changed policy can rematerialize the same raw scores. Runtime, configuration, output-selection, reducer, and genome compatibility belong to the concrete persisted result and are therefore checked by the score store, not against this wrapper's defaults.
score_columns
¶
score_columns() -> list[ScoreColumn]
Return the per-variant score columns, derived from the manifest's result schema.
The manifest's ordered ResultSchema is the single source for the column schema. A ScoreColumn is a
ResultField without its biological meaning, and ModelResultContract and the scoring engine both
require the two to match exactly, so the default is the only correct shape for a manifest-declared
model. Override it only when the columns come from somewhere other than manifest.result_schema.
prioritize_predicate
abstractmethod
¶
prioritize_predicate() -> Predicate
The prioritization rule as a Predicate. It renders to SQL, so the score store can filter with it,
or to Python for the reference store. Its columns may reference both score columns and annotation
columns.
to_model_score
¶
Build one model_scores JSONB entry from this model's per-variant score mapping.
The entry's keys are model_id, model_name, one key per score_columns() entry in order, and
prioritized. The default decodes each score value as its column's declared dtype. A missing, None,
or NaN value becomes None for every dtype: the default never turns an absent measurement into
False, 0, or an empty string, so a store can still tell "not measured" from a measured negative.
Decoding is strict, so a schema mismatch fails instead of producing a plausible wrong value. A
NumPy-style scalar (from a pandas frame) is unwrapped first. float accepts integers and floats but
not booleans. int accepts integers and whole floats, because a frame stores an integer column with
missing values as float, and rejects 3.7. bool accepts booleans and the integers 0 and 1, which
is how SQLite stores a flag. str accepts only strings. Any other value raises TypeError or
ValueError naming the column.
prioritized follows a different rule because it is the materialized per-model decision, not a score.
The materialize path passes the predicate's three-valued result, and a None (a required input was
absent) becomes False. That matches prioritize() and the warehouse export's
COALESCE(is_prioritized, FALSE), so both materialization paths produce the same entry. A present
value is decoded with the strict bool rule.
Override this only when a stored value needs a model-specific decoding. Changing how a stored value is decoded is a breaking result change (see the plugin ABI guide).
prioritize
¶
Evaluate prioritize_predicate against the merged score and annotation values as a plain boolean.
Score and annotation names must be disjoint. A collision raises SchemaCollisionError instead of
silently choosing one producer's value.
This is a convenience for callers who want a definite yes/no. It goes through to_python(), which
collapses the predicate's three-valued (Kleene) result — a missing/None column yields None — down
to False, matching a SQL boolean filter where only a definite True counts.
Core's materialize path does not call this. It evaluates the predicate with
prioritize_predicate().evaluate(...) (see data_lake/materialize.py), which returns the raw
True/False/None, then threads that value into to_model_score and aggregates the row-level flag
with an is True test. The net prioritized outcome matches prioritize() today — to_model_score
performs the same None→False collapse, and None is True is False — so the two do not disagree;
the difference is only where the collapse happens. Keeping the Kleene value until to_model_score
leaves a to_model_score implementation free to treat "the rule did not fire" (False) and "a
required input was absent" (None) differently. Prefer evaluate when that distinction matters;
prioritize is the shortcut when it does not.
PreprocessingRequest
dataclass
¶
PreprocessingRequest(
model_id: str,
configuration: ModelConfiguration,
run_inputs: ModelRunInputs,
resolver: PathResolver,
)
The inputs build_preprocessing_plan needs to turn one model into a PreprocessingPlan.
Like ScoringRequest, it is backend-agnostic: the binding names base-relative logical paths and the
backend PathResolver supplies the mount root. Derived artifacts are architecture-specific, so the binding
owns its logical layout; core carries no preprocessing layout. Staged file names come from the typed
configuration, with no runtime directory listing.
ScoreColumn
dataclass
¶
One per-variant score column a model plugin produces.
Declaring the column here sets four things that must otherwise be kept in sync by hand: the score-store schema (table columns / DDL), the write-time validator, the wide-TSV column names, and the frontend column definitions.
It lives in this shared leaf module — not in model/plugin.py — because it is the schema the score
store consumes (ScoreStore.ensure_schema / add_scores); keeping it here lets the data-lake axis
name it without importing back into the model axis. model.plugin re-exports it for architecture
authors.
ScoringRequest
dataclass
¶
ScoringRequest(
job_id: str,
model_id: str,
configuration: ModelConfiguration,
run_inputs: ModelRunInputs,
resolver: PathResolver,
output: OutputSelection,
scoring_batch: int | None = None,
scoring_batch_count: int | None = None,
layout: ScoringLayout = ScoringLayout(),
)
The inputs build_scoring_plan needs to turn one (model, job) pair into a ScoringPlan.
Paths are described independent of any backend, so a plugin can build the plan for any backend: they
come from a ScoringLayout (a base-relative logical layout) combined with the backend's PathResolver
(the mount root), never a concrete filesystem. configuration holds durable model/resource identity;
run_inputs holds per-run biology; output holds reducer/output selection. Runtime-injected clients
and callables are supplied separately and cannot enter this request.
genome_label
property
¶
Compatibility spelling for the typed run input's genome build.
DelimitedResultCodec
dataclass
¶
DelimitedResultCodec(
schema: ResultSchema,
source_names: Mapping[str, str] = dict(),
row_identity_fields: Mapping[str, str] = dict(),
allowed_extra_fields: tuple[str, ...] = (),
identity_field: str = "variant_id",
delimiter: str = "\t",
null_tokens: tuple[str, ...] = ("", "NA", "NaN", "nan"),
optional_extra_fields: tuple[str, ...] = (),
)
Strict decoder for a headered delimited primary-result file.
source_names maps each declared result name to its runtime-file spelling. row_identity_fields
maps transport columns such as chr/pos/ref/alt to scalar dtypes and preserves them for
storage adapters that need decomposed loci. allowed_extra_fields names required runtime columns that
are intentionally ignored; optional_extra_fields names ignored diagnostics accepted when present.
Every other undeclared column is rejected. This makes both preservation and projection explicit binding
decisions instead of silently dropping new runtime output.
decode
async
¶
decode(
storage: Storage,
output_uris: Sequence[str],
*,
batch_size: int = 10000,
) -> AsyncIterator[ResultBatch]
Stream every collected file in order and validate its complete input schema.
ScoringResultCodec
¶
Bases: Protocol
Decode collected runtime outputs into bounded, schema-shaped primary batches.
decode
¶
decode(
storage: Storage,
output_uris: Sequence[str],
*,
batch_size: int = 10000,
) -> AsyncIterator[ResultBatch]
Yield decoded batches without loading a complete output into memory.
ModelArtifactLayout
dataclass
¶
Paths to shared, model-owned inputs: the staged genome FASTA and files under models/{model_id}.
These live at the storage root under genomes/ and models/, not inside any job directory. Paths are
POSIX and relative to the storage root; the backend's PathResolver supplies the mount root. A binding
that needs architecture-specific model files (derived calibration data, an interpretation motif set, and
so on) names them with model_resource_file or subclasses this layout in its own package.
model_resource_file
¶
Return a fixed model-owned resource path, models/{model_id}/{filename}.
Bindings use this for content-addressed weight bundles, track catalogs, derived model artifacts, and similar resources. They supply a constant basename chosen from declared configuration rather than a provider path or URI suffix; rejecting path components keeps the logical layout portable and prevents traversal.
fold_model_file
¶
Path to one cross-validation ensemble member's weights, models/{model_id}/fold_{fold}_model{ext}.
ext includes the leading dot, e.g. .h5. Derive it from the binding's declared weight format, never
from a resource URI: a content-addressed URI need not carry a file extension.
ScoringLayout
dataclass
¶
Bases: ModelArtifactLayout
Model-neutral paths for one container scoring run, one method per file a scoring plan names.
Adds the scoring job's inputs and outputs on top of the shared genome and model-file paths from
ModelArtifactLayout. ScoringRequest.layout defaults to this class; the scoring engine and the portable
preparation path read the variant input and result paths from it.
ineligible_variants_file
¶
Return the report of candidates that batch preparation excluded as ineligible for this model.
fold_score_file
¶
Return one ensemble member's intermediate score file, reduced by the plan's summary task.
scoring_result_file
¶
Return the logical path of the primary result table the scoring engine collects.
detail_result_file
¶
detail_result_file(
job_id: str,
model_id: str,
name: str,
scoring_batch: int | None = None,
) -> str
Return the logical TSV path for one manifest-declared named detail result.
ModelWorkflowPlan
dataclass
¶
ModelWorkflowPlan(
operation: str,
stages: tuple[WorkflowStage, ...],
plugin_identity: PluginRunIdentity,
outputs: tuple[Transfer, ...] = (),
)
A backend-neutral DAG for one optional model operation.
operation is a stable plugin-facing discriminator such as "preparation" or
"interpretation". Each stage runs all of its tasks in parallel after every stage in needs has
completed. outputs names the terminal artifacts the caller wants copied back through the backend.
Intermediate products may remain on the backend's shared working volume.
WorkflowExecutionError
¶
Bases: RuntimeError
A model-workflow task failed on its execution backend.
WorkflowRunResult
dataclass
¶
WorkflowRunResult(
plan: ModelWorkflowPlan,
completed_stages: tuple[str, ...],
output_uris: tuple[str, ...],
)
The completed stages and collected terminal artifact locations for one workflow.
WorkflowStage
dataclass
¶
WorkflowStage(
name: str,
tasks: tuple[ContainerTaskSpec, ...],
needs: tuple[str, ...] = (),
)
One parallel task set whose declared prerequisite stages must finish first.
canonical_chromosome
¶
Return the canonical chr-prefixed spelling used in portable keys.
Primary chromosomes accept common aliases (1/chr1/CHR01 and
M/MT/chrM) in any case. Other reference contigs keep their exact
spelling after an optional case-insensitive chr prefix is normalized, so
CHRUn_KI270302v1 becomes chrUn_KI270302v1 but chrun_KI270302v1
stays as written. Non-ASCII text, whitespace inside the name, and characters
outside the VCF contig-name set (including :) raise
VariantIdentityError. This is a naming rule only; it does not claim that
contigs from different assemblies are interchangeable.
canonical_variant_id
¶
canonical_variant_id(
chromosome: str,
position: int,
reference_allele: str,
alternate_allele: str,
) -> str
Return the canonical text key for decomposed one-based locus fields.
canonical_json
¶
Return canonical UTF-8 JSON text: sorted object keys, compact separators, no NaN.
classify_variant
¶
classify_variant(key: VariantKey) -> VariantClass
Return the :class:VariantClass of key from its REF and ALT lengths.
Classification trusts that both alleles are literal bases. A placeholder allele such as - or * would
count as one base; rejecting those is :class:VariantKey's job, not this function's.
deterministic_hash
¶
Return a lowercase SHA-256 identity over :func:canonical_json.
validate_result_identity
¶
validate_result_identity(
manifest: PluginManifest, identity: PluginRunIdentity
) -> None
Reject a stored primary result before the current plugin assigns it new meaning.
This validates the binding-owned portion of primary-score compatibility. Concrete model configuration,
runtime, output, reducer, and genome are compared by stores through :class:ResultCompatibility.
Prioritization belongs to materialization rather than score identity.
require_resource_digests
¶
require_resource_digests(
value: ConfigurationModel, *, owner: str
) -> None
Raise ValueError unless every resource on value carries a SHA-256 digest.
Checks each top-level field whose value is a ResourceReference or a tuple or list of them; an unset
optional resource (None) is skipped. Nested configuration models and mappings are not traversed, so a
configuration that nests resources must also check the nested model. Missing resources are named in
field order, with an index for a sequence (weights[2]), and owner prefixes the message. Call it
from a model_validator(mode="after") on a configuration whose runtime verifies staged bytes, so
Pydantic reports the failure as a validation error before a plan can be built.
get_model_plugin
¶
get_model_plugin(model_type: str) -> ModelPlugin
Look up and instantiate the ModelPlugin for a model architecture, keyed by its model_type string
(the registry key, e.g. "CHROMBPNET").
A call site holds a model's model_type and asks for its plugin without naming a concrete class, which
is what lets the scoring, interpretation, and preprocessing flows work with any architecture. Adding an
architecture then only requires registering it. The registry stores the plugin class (an entry point
loads the class object); this instantiates it, and plugins take no constructor arguments. Raises
KeyError, listing the registered names, if model_type has no plugin.
get_model_plugin_registry
¶
get_model_plugin_registry() -> PluginRegistry[
type[ModelPlugin]
]
Return the shared registry of model-architecture plugins. It stores ModelPlugin subclasses, so
.get(model_type) returns a type[ModelPlugin].
resource_transfer
¶
resource_transfer(
resource: ResourceReference,
logical_path: str,
*,
locality_key: str | None = None,
) -> Transfer
Declare a configured resource as a container input staged at logical_path.
The transfer carries the resource's uri as its location and its digest as the content identity the
staging backend verifies, so relocating identical bytes changes neither the plan's meaning nor its cache
identity. revision and subtype selectors stay in the configuration's scientific identity; a storage
adapter reads only the location and the digest. locality_key is the optional co-scheduling hint
described on Transfer, typically the model ID for model-owned resources.
scoring_plan_from_dict
¶
Rebuild a scoring plan, rejecting unknown envelope fields.
scoring_plan_from_json
¶
Rebuild a scoring plan after a process restart.
scoring_plan_to_dict
¶
Serialize either scoring-plan shape to JSON-compatible values.
scoring_plan_to_json
¶
Serialize a durable scoring plan to JSON.
run_model_workflow
async
¶
run_model_workflow(
plan: ModelWorkflowPlan,
*,
backend: ExecutionBackend,
input_transfer: TransferMapper | None = None,
output_transfer: TransferMapper | None = None,
spec_mapper: SpecMapper | None = None,
poll_interval_s: float = 1.0,
max_ticks_per_wave: int = 3600,
) -> WorkflowRunResult
Execute all ready stages in dependency waves and collect declared terminal outputs.
workflow_plan_from_dict
¶
workflow_plan_from_dict(
data: Mapping[str, object],
) -> ModelWorkflowPlan
Rebuild and validate a workflow plan from durable values.
workflow_plan_from_json
¶
workflow_plan_from_json(value: str) -> ModelWorkflowPlan
Rebuild a durable model workflow plan from JSON.
workflow_plan_to_dict
¶
workflow_plan_to_dict(
plan: ModelWorkflowPlan,
) -> dict[str, Any]
Serialize a workflow plan to JSON-compatible values.
workflow_plan_to_json
¶
workflow_plan_to_json(plan: ModelWorkflowPlan) -> str
Serialize a durable model workflow plan to canonical JSON.