End-to-end examples¶
The recommended examples are four progressive Marimo notebooks under examples/tutorials/. They use only
Altar's supported public façades and installed binding interfaces, and every scientific row comes from real
model execution or a real reference dataset.
Progressive real-data series¶
| Notebook | Adds | External requirement |
|---|---|---|
01_score_with_cherimoya.py |
Local CATv1 Cherimoya scores and the binding-default priority flag | NVIDIA GPU, Docker, hg38 FASTA, cohort TSV, official CATv1 resources |
02_add_alphagenome.py |
Hosted AlphaGenome scores and track-level details | AlphaGenome API key and explicit scorer/tissue scope |
03_add_spliceai.py |
Gene-level SpliceAI evidence and source priority provenance | Licensed Illumina data prepared as Parquet |
04_custom_prioritization.py |
Explicit per-model and cross-model client policies | User-justified thresholds |
Start with the real scoring quickstart, or open the first notebook directly:
uv sync --all-packages --group tutorials
uv run --group tutorials marimo edit examples/tutorials/01_score_with_cherimoya.py
All four notebooks share one configured work directory and SQLite score store. The sequence demonstrates the M+N design directly: a second model does not change the first binding, and adding an annotation source does not change either model result set.
Reproducibility and cost controls¶
The notebook files are plain Python and pass marimo check --strict in CI. Public CI does not call model
services or execute GPU work. Each notebook validates prerequisites and exposes a deliberate run button;
compatible stored scores are reused, and the scoring notebooks submit only variants missing from the model
cache.
Model execution still has real costs: AlphaGenome consumes hosted API quota, Cherimoya consumes local GPU time, and reference datasets may have license restrictions. Review the configuration cell before every run.
Lower-level client and contract programs¶
The remaining directories under examples/ exercise narrower public-interface compositions used by CI:
scoring/exposes plan construction and backend orchestration as command-line programs;annotation_sources/checks that scientific source behavior is stable across storage backends;variant_gene_links/checks canonical relation ingestion and query behavior; andanalysis/covers evidence-presence and missing-evidence branches.
Some of those programs use explicitly labeled fixtures or smoke tasks to test a contract. They are not the scientific tutorial, and their fixture outputs must not be interpreted as model predictions. Use the Marimo series whenever the goal is to learn or demonstrate an actual analysis.