Run a real scoring tutorial¶
Altar's primary tutorial is a four-part Marimo notebook series. It starts with real local model inference, adds a hosted model, joins a real annotation dataset, and then evaluates an explicit client policy over the stored model scores.
The notebooks do not manufacture variants, genomes, scores, API responses, or annotation rows. Their first editable cell declares every path, credential reference, model identity, and analysis parameter you must supply. Expensive work runs only after you press a labeled run button.
Prepare the checkout¶
Altar requires Python 3.12 or newer. From a source checkout:
This installs the framework, the Cherimoya, AlphaGenome, and SpliceAI bindings, and the pinned Marimo version used to validate the notebooks.
Prepare real inputs¶
Tutorial 1 requires Docker with the NVIDIA runtime, an hg38 FASTA with its SHA-256 digest, a real cohort in Altar's canonical headerless TSV format, an immutable Cherimoya runtime image, and one official CATv1 model ensemble. Prepare the five content-addressed CATv1 folds with:
uv run --project runtimes/cherimoya cherimoya-resources fetch-official \
--experiment-accession ENCSR000EOT \
--output-dir /absolute/path/to/cherimoya-resources
ENCSR000EOT is a real K562 DNase experiment used in Altar's validated Cherimoya run. Select a different
CATv1 experiment when the biological context requires it, and keep the accession, model ID, model name, and
resource manifest aligned in the notebook configuration cell.
Later notebooks require:
- an AlphaGenome API key, provided through the environment rather than stored in notebook source;
- an explicit AlphaGenome scorer and ontology selection appropriate for the analysis; and
- Illumina's licensed hg38 SpliceAI data converted to Parquet with
altar-spliceai-build.
See the notebook README for the complete input and licensing checklist.
Start tutorial 1¶
Edit the configuration cell, resolve every displayed prerequisite, and press Run or load real Cherimoya scores. The notebook verifies the FASTA digest, stages the real inputs, executes the five CATv1 folds on the local GPU, persists their summarized results in SQLite, and displays the binding's default priority flag.
The configured work directory is the handoff between notebooks. Existing scientifically compatible scores are loaded rather than recomputed; newly added cohort variants are the only variants sent for scoring.
Continue through the series¶
01_score_with_cherimoya.pyruns one real local model and its default predicate.02_add_alphagenome.pycalls the real hosted API, stores lossless per-track details, and materializes two models together.03_add_spliceai.pyadds real precomputed evidence and reports model and annotation-source provenance separately.04_custom_prioritization.pycompares binding-owned defaults with user-supplied per-model and cross-model predicates.
The fourth notebook keeps custom analysis decisions separate from Altar's materialized default-policy provenance. A custom cross-model rule is not silently presented as though it came from either binding.
Interpret the result¶
Each materialized variant retains separate model score entries, an overall default-policy decision, the IDs
of models whose rules fired, the names of prioritizing annotation sources, and exact plugin/run identities.
A prioritized value means a declared rule matched. It is not a pathogenicity, causality, clinical, or
statistical conclusion. Read what prioritized means before interpreting it.