Enformer¶
Reference integration
Enformer is an independently installed model binding with an isolated Python 3.10 / TensorFlow 2.13 runtime. Altar does not make a remote, one-variant API call. The global variant set is partitioned into deterministic scoring batches, each batch becomes an independently retryable container task, and the runtime streams that bounded input through one loaded model.
Model-instance scope¶
An Enformer model instance is defined by:
- a content-addressed TensorFlow SavedModel archive;
- content-addressed official human target metadata;
- a strictly ordered selection of output indices;
- an immutable runtime image; and
- the content identity of the hg38 FASTA used for the run.
Resource locations are operational. Moving identical bytes between local storage, HTTPS, S3, or another supported transfer backend does not change scientific cache identity.
Result semantics¶
For each selected human output track, the runtime sums predictions across Enformer's 896 output bins and calculates the Basenji variant statistics linked from DeepMind's Enformer release:
SAD = alternate_activity - reference_activity
SADR = log2(1 + alternate_activity) - log2(1 + reference_activity)
The primary table contains one row per variant and identifies the track with the largest absolute SAD. Its
signed SAD and SADR remain in the row. The manifest-declared track_effects detail table preserves reference
activity, alternate activity, SAD, and SADR for every selected (variant, track) pair. SAR is a different
Basenji statistic—the sum of per-bin log ratios—and is not emitted under that name.
The initial prioritization predicate is abs(top_track_sad) >= 0.5. This is an explicit versioned policy,
not a claim of universal biological significance; callers should select biologically relevant tracks and
may rematerialize under a later policy without rerunning raw model inference.
Scale and execution¶
Outer scoring batches are Altar's unit of scheduling, retry, and checkpointing. Inside each task the runtime
uses a small tensor microbatch, writes rows incrementally, and never assembles the complete job or complete
detail matrix in memory. Primary and detail files are routed separately through the generic ScoreStore and
DetailStore axes.
The same plan runs through any execution backend that can satisfy its container and GPU resource request. The binding contains no Kubernetes, Modal, object-store, or model-registry client.
Current boundaries¶
- The published human model and this integration are hg38-only.
- The first runtime release accepts SNVs only. The manifest declares SNV-only variant eligibility, so batch preparation filters out and reports indels and multi-nucleotide variants, and the rest of a mixed cohort still scores. The runtime still rejects a non-SNV in a hand-staged batch. Indels stay out of scope until a versioned alignment policy is defined and validated.
- Track percentiles and other empirical calibration assets are not part of raw inference. They can be added later as separately versioned, content-addressed resources.
- A real deployment must prepare a tar archive containing exactly one Enformer SavedModel and register both
that archive and
targets_human.txtwith SHA-256 digests. The isolated runtime providesenformer-resources fetch-officialto fetch pinned upstream versions, validate all 5,313 metadata rows, build a deterministic archive, and emit a content-identity manifest.enformer-score smokeruns a full-shape inference against those exact bytes before they are admitted to durable model storage.
The current weight-free NVIDIA L40S validation
record exercises the official model and
metadata through both the full-shape smoke and complete scoring paths under the development semantics-2
lineage consolidated into published contract 0.0.1. It independently
checks the SAD/SADR formulas and max-absolute-SAD primary selection without redistributing model weights.
Runtime images are published under commit-addressed GHCR tags, and the image workflow reports the resulting
OCI digest. Model configurations use the digest form (ghcr.io/kundajelab/altar-enformer@sha256:...), never
the tag.
See the Enformer repository for the model implementation, published output definitions, and attribution.