Skip to content

ChromBPNet

Reference integration

ChromBPNet predicts base-resolution chromatin-accessibility profiles and total accessibility from DNA sequence. In Altar, one model instance represents one trained biological context such as a cell type.

Install the lightweight binding separately from the base framework:

pip install altar-chrombpnet

The install contributes the CHROMBPNET model-plugin registry entry. Direct imports come from altar_chrombpnet; the base altar wheel contains no ChromBPNet-specific implementation or entry point.

Inputs and execution

The binding consumes a reference genome, canonical variants, five fold models, and a matched peak set. Each raw resource must have a SHA-256 digest; the five fold digests must be distinct and ordered. Preprocessing derives the five peak distributions and safe interval index, while scoring runs one GPU task per fold and a CPU task that summarizes all five outputs. The separately locked Python 3.9/TensorFlow runtime lives in runtimes/chrombpnet.

The same plan can be submitted through local Docker, Modal, or Kubernetes. The model binding does not import those provider SDKs.

Every digest-bearing input a task reads is checked before use. The execution backend verifies each declared Transfer digest (local and Modal when staging, Kubernetes in a verify-inputs init container), and the runtime also checks the files named on its command line: fold weights and genome in scoring, all five weights, peaks, and genome in preparation, and the genome in interpretation. The first check of a file hashes it and writes a verification record beside it; later tasks accept the unchanged file from that record (see Hash once per volume). A 3 GB reference genome on network storage is therefore hashed once per volume rather than once per fold. The runtime logs whether each input was hashed or accepted from its record. The resource preparer and the smoke workflow always hash, because they are where bytes are first checked.

The runtime includes a reproducible reference preparer for ENCODE's K562 DNase ensemble. It verifies the two source archives, extracts the five bias-corrected models and matched all-input peak set, and writes a resource manifest suitable for model registration:

cd runtimes/chrombpnet
uv run chrombpnet-resources fetch-official --output-dir /path/to/models/k562-dnase

The checked-in official smoke workflow loads all five weights, verifies their digests and tensor contracts, and executes one forward pass per fold on a self-hosted GPU. This is a validation fixture, not a restriction to K562: other ChromBPNet model instances use the same five-weight-plus-peaks contract.

Model preprocessing streams the peak atlas in fixed 2,048-sequence batches and predicts only the count head. By default it predicts only the forward orientation, preserving varscore behavior. An explicit configuration option enables reverse-complement averaging for comparison with the dedicated ChromBPNet variant scorer. This keeps the official 196,077-peak K562 input within the normal-GPU pool instead of materializing the full one-hot atlas or reserving an entire accelerator.

Derived artifacts remain at the conventional models/{model_id} paths. Preparation also emits an optional provenance manifest whose identity covers the ordered weight digests, peaks digest, genome digest, strand policy, and preprocessing-semantics version. The manifest separately records the exact producer image and every output's SHA-256 digest and byte size, and is written atomically after all six outputs.

The normal scoring plan does not require that manifest or an Altar-owned custody layout. Existing trusted artifacts can be used in place without copying or attestation. Standalone runtime callers may opt into strict verification by supplying both the manifest and expected preparation identity. In that mode, fold scoring verifies its selected peak distribution and summary verifies the interval index before reading either file.

On NVIDIA L40S validation hardware, the forward-only K562 preprocessing run completed in about 25 minutes while sampled usage stayed near 4.4 GB host RSS and 17.9 GB GPU memory. Registration therefore needs a 24 GB-class accelerator with the current five-model preprocessing implementation; Altar's normal Modal GPU pool maps to an A10G. Five independent cold fold runs over 64 variants each completed in 10–13 seconds per fold. Reverse-complement averaging performs both orientations and is expected to take roughly twice the model inference work; it does not yet have a full-atlas timing baseline. These are reproducibility measurements, not scheduler guarantees; they include model startup but exclude queue and image-pull time.

Resource formats and staged paths

Every resource is staged at a binding-owned path under models/{model_id} whose name comes from declared configuration, never from the resource URI. A content-addressed URI therefore needs no file extension.

Field Values Staged file
weights + weights_format h5 (default; Keras HDF5) fold_{fold}_model.h5
peaks + peaks_compression gzip (default) or none peaks.bed.gz or peaks.bed
motifs (optional) TF-MoDISco HDF5 motif set motifs.h5

The formats only choose staged file names for bytes whose digests already identify them, so they do not enter the score-cache projection. Configuration schema 0.0.1 inferred both from URI suffixes; loading a saved 0.0.1 configuration records the same choice explicitly once (h5, and gzip when the peaks URI or archive member ends in .gz).

Interpretation

Interpretation runs one contribution-score task per fold, averages them, calls motif hits for each allele with finemo, and renders plots. It requires a configured, content-addressed motifs resource; the plan stages it with its digest beside the model's weights instead of reading an unidentified motifs.h5 from the storage root. Fold weights, the genome, and the motif set are declared as digest-carrying transfers exactly as in scoring, so the backend verifies them exactly as it verifies scoring inputs, verification records included. The interpretation runtime's command line names no weight digest, so, unlike fold scoring, it does not re-check the weights itself; the backend check is the guarantee. The interpretation commands are unchanged apart from the motif path, so they run on the pinned default runtime image.

The interpretation plan identity covers the complete configuration, genome, both runtime images, and the motif set. The motif set does not change any score, so it is excluded from the scoring projection: adding or replacing it keeps cached scores valid, while it changes the interpretation identity.

Score fields

Field Meaning
logfc Predicted log fold-change between alternate and reference alleles. Sign records effect direction.
jsd Jensen–Shannon divergence between predicted allele profiles.
active_allele_quantile Quantile of the more active allele relative to the model's peak activity distribution.
in_peak Whether the variant is in the model's accessibility peak set.

Scores are meaningful only in the context of the trained model instance and its assay/cell type. Do not compare values from unrelated model instances without confirming compatible training and calibration. Genome-invalid inputs remain present with all four fields null and are never passed to the model. Their raw result rows retain error_reason and error_message diagnostics; those operational columns are not public score fields. The runtime treats a variant as invalid when ChromBPNet's 2114 bp input window doesn't fit on its contig, or when it is on chrM. Altar's generic validation accepts both, so they reach the runtime and come back with null scores.

Prioritization rule

The reference predicate requires:

  • abs(logfc) >= 0.25;
  • active_allele_quantile >= 0.05; and
  • a promoter annotation, membership in a peak, or a positive predicted accessibility gain outside a peak.

The rule reads region_type, which the manifest declares as a dependency on the org.kundajelab.altar.annotation.regions contract (version 1.0.0). Materialization fails when no configured annotation source supplies region_type. A variant without a region row leaves the promoter branch unknown. altar.variants publishes the contract as REGION_ANNOTATION_CONTRACT. RegionAnnotationSource implements it, and a host table loaded from the annotate command declares it:

from altar.sources import BigQueryAnnotationSource
from altar.variants import REGION_ANNOTATION_CONTRACT

regions = BigQueryAnnotationSource.connect(
    "my-project.annotations.variant_regions",
    REGION_ANNOTATION_CONTRACT.columns,
    name="regions",
    genome_default="GRCh38",
    identity=REGION_ANNOTATION_CONTRACT.identity(
        release="ensembl-116.ccre-v4.genes-gencode-47.logic-1", genome_build="hg38"
    ),
)

If Altar also writes this region table, add identity_scoped=True to it and to every other source that reads it, so a release change re-annotates the table. Migrate an existing table first, as described in Migrate an existing BigQuery lake. See Region annotation source and Use your own annotation tables.

The thresholds are an analysis heuristic, not a clinical classifier.

Reproducibility record

The configuration records the five fold artifacts and checksums, their declared weight format, the peak set checksum and compression, the optional motif set checksum, immutable runtime image, reverse-complement policy, and indel-JSD policy. Also record the genome FASTA digest, cell type/assay, generated peak distributions, binding version, and prioritization rule. Existing configurations without explicit scoring policies migrate to the varscore-compatible average_reverse=false and adjust_indel_jsd=false defaults. Pattern-only registrations cannot be migrated safely because one digest cannot establish the identities of five different files; prepare and register them again.

Provenance is not a prerequisite for use. Trusted artifacts managed by a downstream host can be scored directly; historical score rows remain valid and are not rewritten. When preparation provenance is available, it can be retained as metadata and optionally checked by the runtime. A legacy store with no Altar identity does not need a metadata import before its rows can be used.

A trusted host holding an older pickle-backed peaks.dnatree may convert its own known artifact to JSON Lines without running ChromBPNet. Altar still never deserializes an untrusted pickle. This serialization migration is separate from the earlier coordinate bug: an interval index known to contain the old BED-end error must not be described as corrected semantics until that structural error is repaired.

Peak preprocessing converts between one-based variant positions and zero-based BED peak windows. Peak trees generated by an earlier runtime used an inclusive representation for an exclusive BED end. Re-run preprocessing with the current runtime before comparing or publishing in_peak; the numeric logfc, jsd, and active_allele_quantile definitions are unchanged.

Scoring canonicalizes chromosome aliases in variant and peak inputs before computing in_peak. Regenerate scores produced from unprefixed contigs such as 1; earlier runtimes silently reported in_peak=false for those rows. Fold scoring also disables output reuse by default because fold paths do not carry complete scientific identity. Opt-in reuse is reserved for independently verified standalone runs.

The varscore compatibility mode remains the default: forward-only prediction and direct JSD for every allele length. The dedicated ChromBPNet variant scorer instead defaults to two-strand inference and indel-aware JSD. Altar exposes these as independent opt-ins rather than silently choosing one scientific interpretation. Reverse-complement logits are returned to forward coordinates before profile averaging, and logcounts encode counts averaged in linear space. Indel adjustment removes the predicted positional shift attributable only to the allele-length difference before JSD. Both choices are part of scientific cache identity.

Enabling reverse-complement averaging requires regenerating numeric scores and every fold's peak distribution. Enabling only indel adjustment changes indel JSD and downstream IPS, but not peak distributions or SNV scores. The prediction-free peaks.intervals.jsonl does not change.

Reference parity and deliberate differences

Altar uses kundajelab/varscore at 3497f6c as the default behavioral reference. “Compatible” means the same default model orientation and numeric score definitions, not preservation of known coordinate bugs or unsafe operational behavior.

Concern varscore Altar Scientific or operational impact
Model orientation Forward only Forward only by default; reverse-complement averaging is opt-in Opting in changes counts, profiles, all scores, and peak calibration, and roughly doubles inference work.
JSD for indels Direct comparison of allele profiles Direct comparison by default; allele-length realignment is opt-in Opting in changes indel JSD and IPS only. SNVs and peak calibration are unchanged.
Variant-to-peak coordinates Queries a zero-based peak tree with the one-based variant position Converts the variant position to zero-based before lookup Deliberate correctness fix; in_peak can differ at peak boundaries. Numeric model scores are unchanged.
BED interval end Stores the BED-exclusive end as an inclusive tree endpoint Converts the exclusive end to the tree's inclusive representation Deliberate correctness fix; removes one extra base from in_peak membership. Numeric model scores are unchanged.
Chromosome aliases Uses input contig spelling directly Canonicalizes aliases such as 1 and chr1 Deliberate correctness fix; prevents false in_peak=false results.
Existing-output reuse Enabled by default Disabled by default Operational integrity choice; avoids reusing scores whose full scientific identity was not verified.
Peak index artifact Pickle-backed peaks.dnatree Versioned JSON-lines intervals Security/portability change only; the in-memory interval-query structure is retained.
Invalid row in a batch Assertion aborts the complete batch Scores valid neighbors and retains the invalid row with null scores plus raw diagnostics Operational scale hardening; model inputs and valid-row score formulas are unchanged.
Derived-artifact provenance Peak distributions and interval index live under a model directory without a parent manifest Conventional paths remain supported; preparation emits an optional manifest and runtime verification is opt-in Adds reproducibility metadata without requiring Altar custody or changing score formulas.
Scale and validation Whole-input-oriented reference implementation Bounded streaming, content digests, tensor-contract validation, and fail-closed fold reduction Operational hardening intended not to change the score formulas.

The optional two-strand and indel policies reproduce the relevant behavior from kundajelab/variant-scorer at 0e1e341. They are not promoted to Altar defaults without a separate scientific decision and benchmark.