Skip to content

Sei

Experimental integration

The Sei binding is not yet scientifically validated against the released official weights. Use it for integration testing and evaluation, not as a validated production scorer. A pinned default image and an Altar validation record will follow real-weight CPU/GPU parity testing.

Sei is an independently installed model binding with an isolated PyTorch runtime. The binding contains only portable configuration, schemas, and an execution plan; importing it does not import PyTorch or load the approximately 3.4 GB official checkpoint.

Install the binding alongside Altar. It is not on PyPI; install it from a source checkout, which also installs Altar:

uv pip install -e bindings/sei

Scientific contract

An Altar Sei model instance identifies four official resources by SHA-256 digest:

  • sei.pth, the 4,096 bp sequence model predicting 21,907 chromatin profiles;
  • projvec_targets.npy, the published sequence-class projection vectors;
  • histone_inds.npy, the chromatin-profile indices used for nucleosome-occupancy adjustment; and
  • seqclass.names, the ordered catalog of 40 sequence classes.

For each SNV, the runtime extracts the same even-length window used by Selene: the variant occupies zero-based index 2,047. Reference and alternate alleles are each predicted on the forward sequence and reverse complement, and their profile predictions are averaged. The runtime then applies FunctionLab's published sc_hnorm_varianteffect procedure:

  1. sum all selected histone-profile predictions separately for reference and alternate;
  2. scale each allele's histone profiles to the mean of those two totals;
  3. project the adjusted 21,907-profile vectors onto the published class vectors; and
  4. report adjusted alternate minus adjusted reference for the first 40 sequence classes.

The pairwise adjustment must happen before projection. A runtime that projects first, or returns ordinary alt - ref, is not numerically equivalent.

The primary score table retains the class with largest absolute effect, including its signed effect. The sequence_class_effects named detail contains one row for every (variant, class) pair, so the compact primary choice does not discard any published class score. The 21,907 underlying profile predictions are transient model intermediates in this first contract and are not persisted as scalar scores.

Scale and execution

The binding emits one independently retryable container task for each staged batch. The runtime loads the checkpoint once, streams the staged variant file, and uses bounded tensor microbatches; it does not launch a model process per variant.

prepare_scoring_batches cache-filters and densely repacks Sei batches. Pass the DetailStore that holds sequence_class_effects; a variant is a cache hit only when its primary row and its class-effect rows are both stored (see bindings with named details).

The current binding accepts hg19 and hg38, matching the two reference-genome workflows in the official Sei repository. Like the official workflow, it does not score a variant without a complete 4,096 bp contig window; Altar records an explicit invalid primary row and continues with valid neighbors instead of silently padding or failing the complete batch. Reference-allele mismatches are handled the same way. This is deliberately stricter than Selene's fallback of replacing a mismatching FASTA base with the VCF reference while marking ref_match false. Unknown N bases inside an otherwise valid window use Selene's uniform 0.25-per-channel encoding and are reported through contains_unk; other IUPAC window bases and unsupported allele symbols produce an invalid row. Repeated canonical variant identifiers are deduplicated before inference. It initially supports SNVs only. The manifest declares SNV-only variant eligibility, so batch preparation filters out and reports indels and multi-nucleotide variants before they reach the runtime. The runtime still writes an invalid row for a non-SNV in a batch staged some other way, and the engine rejects that row rather than storing it, which fails the batch. A BigQuery deployment must therefore pass the manifest's eligibility to prepare_unscored_table() or export_unscored() (see variant eligibility). Upstream Selene also defines indel centering and truncation, but Altar will add that as an explicit versioned policy rather than silently treating indels like substitutions.

The default predicate marks max_abs_class_effect >= 1 as prioritized. Sei does not publish that value as a universal biological cutoff; it is an operational default and should be calibrated for the cohort and analysis.

Upstream parity and Chorus

The numerical authority is FunctionLab's Sei framework and its sc_hnorm_varianteffect implementation. Chorus was inspected as an independent integration and helped expose a useful failure mode: its changelog records that earlier Sei scores omitted the nucleosome adjustment. Altar uses Chorus as corroboration and implementation evidence, not as the scientific standard or execution architecture. In particular, Altar keeps its dense cache-miss batching and backend-neutral plans rather than adopting Chorus's on-demand oracle process model.

Licensing and resources

The official Sei software and model are licensed for academic and research use only. The separately licensed runtime retains that upstream notice; the MIT-licensed binding does not redistribute weights or model code. Prepare content-addressed runtime inputs from the official Zenodo model record with sei-resources package, then register the resulting file digests in SeiConfiguration.