Skip to content

Cherimoya

Reference integration

Cherimoya is a BPNet-style sequence model exposed through the independently installable altar-cherimoya binding. The lightweight binding builds work declarations; the PyTorch runtime and model dependencies remain in runtimes/cherimoya.

Inputs and execution

Each GPU task loads one staged fold resource, reads the reference genome and canonical variants, and writes per-fold scores. A CPU task validates the complete fold set and averages its outputs. The model runtime sees only mounted files; storage adapters may retrieve the weights from any supported location.

The runtime includes a provider-facing preparation command for the official CATv1 atlas. It downloads one experiment's five checkpoints from a pinned upstream revision, verifies the checkpoint container and fold identities, and emits a content-addressed manifest for durable registration:

cd runtimes/cherimoya
uv run cherimoya-resources fetch-official \
  --experiment-accession ENCSR000EOT \
  --output-dir /path/to/models/k562-dnase

ENCSR000EOT is the official K562 DNase reference used by Altar's smoke validation, not a global default. CATv1 has 1,518 experiment-specific GRCh38 models, and an experiment accession must be selected deliberately. The scoring runtime streams variants in bounded outer batches and uses a separate model inference batch, so memory use does not grow with the complete variant file. Fold summarization is likewise a row-aligned stream, so ensemble fan-in does not rematerialize the cohort.

Upstream compatibility

Altar loads checkpoints with the official Cherimoya.load API and executes them with Tangermeme's predict. Its reference and alternate windows are also parity-tested against Tangermeme 1.4.1's public substitution_effect, insertion_effect, and deletion_effect operations, including their default right-flank trimming convention. Altar retains its own reference-backed window builder because those operations accept already encoded tensors: they do not validate VCF reference alleles, retrieve genomic sequence, define one consistent variant-centered window, or stream a large canonical variant table.

Cherimoya's documented variant-effect example subtracts the alternate and reference count heads, which is the basis of counts_log2fc. Upstream exposes profile logits but does not prescribe a scalar profile-L1 variant score. profile_l1 is therefore an explicitly Altar-defined reduction over softmax-normalized profiles, not a claim of an upstream Cherimoya metric.

The artifact configuration declares:

  • one generic ResourceReference(uri=..., digest="sha256:...") per fold;
  • the input sequence window (the number of weight references is the fold count);
  • the binding's default immutable runtime image, or an explicit name@sha256:<digest> override; mutable tags are rejected.

Because a resource digest takes precedence over its URI in scientific identity, relocating identical weight bytes between local disk, HTTPS, GCS, S3, or a model registry does not invalidate compatible scores. The selected storage adapter resolves the URI and the execution backend verifies staged bytes before compute.

Score fields

Field Meaning
counts_log2fc Alternate-minus-reference single-track log(count + 1) effect, converted to log2 units.
profile_l1 L1 distance between the softmax-normalized reference and alternate profile distributions.

These are the fields declared by the current Cherimoya runtime. They are not renamed to ChromBPNet's logfc/jsd fields because the present computations and schema are not identical.

profile_l1 compares softmax-normalized profiles; comparing raw profile logits would be scientifically undefined because their additive offset is arbitrary. The runtime also rejects multi-track Cherimoya checkpoints; CATv1 checkpoints have the required profile (N, 1, 1000) and count (N, 1) outputs, while collapsing a multi-track model into these scalar fields would require a separately declared reducer and result schema.

Prioritization rule

The binding currently marks abs(counts_log2fc) >= 0.25 as prioritized. This is a conservative placeholder, not a calibrated biological threshold.

Reproducibility record

Record every checkpoint digest, fold order, input window, genome FASTA identity, runtime image digest, binding version, experiment accession, assay/biosample metadata, and threshold policy. URIs remain useful provenance but are not the checkpoint identity. CATv1 checkpoints are GRCh38-only; reject rather than score a different assembly.