Precomputed evidence¶
These bindings retrieve released evidence and do not execute model inference during an Altar analysis.
Each annotation binding publishes its columns as an AnnotationContract and declares an
AnnotationSourceIdentity from identity(). Another source, such as a host's warehouse table, can implement
the same contract; see Use your own annotation tables.
| Binding | Contract constant | source_id |
Identity release |
|---|---|---|---|
| GPN-Star | GPNSTAR_ANNOTATIONS |
org.kundajelab.altar.annotation.gpnstar |
<score_set>@<source_release_revision> |
| SpliceAI | SPLICEAI_ANNOTATIONS |
org.kundajelab.altar.annotation.spliceai |
<source_release>-masked or <source_release>-raw |
| AlphaMissense | ALPHAMISSENSE_ANNOTATIONS |
org.kundajelab.altar.annotation.alphamissense |
<source_release> |
| gnomAD | GNOMAD_ANNOTATIONS |
org.kundajelab.altar.annotation.gnomad |
<version>-<dataset> |
| ClinVar | CLINVAR_ANNOTATIONS |
org.kundajelab.altar.annotation.clinvar |
<YYYYMMDD> release date |
All five contracts are at version 1.0.0. The per-row provenance columns below remain part of the GPN-Star, SpliceAI, and AlphaMissense contracts. The gnomAD and ClinVar contracts have none: their release is only in the identity.
GPN-Star¶
altar-gpnstar retrieves allele-specific scores from a pinned GPN-Star release. It supports only biallelic
SNVs covered by the selected assembly and whole-genome alignment.
| Field | Interpretation |
|---|---|
gpnstar_llr_calibrated |
Mutation-rate-calibrated log-likelihood ratio of alternate versus reference; more negative indicates stronger evolutionary constraint or predicted effect. |
gpnstar_abs_llr_calibrated |
Independently calibrated absolute-LLR statistic; it is not abs(gpnstar_llr_calibrated). |
The source also emits checkpoint, alignment, species set, genome build, release revision, and source-asset checksum. It intentionally defines no default prioritization threshold.
SpliceAI¶
altar-spliceai normalizes the precomputed Illumina release. A variant may have several gene records. Scalar
fields come from one coherent record with the greatest delta score, while spliceai_gene_scores preserves
every record as repeated structured detail.
SpliceAIRelease declares the dataset's genome build (hg19 or hg38), whether it holds Illumina's masked
or raw scores (masked), and the upstream release. Masked files set a delta score to 0 when it reports a gain
of an existing splice site or a loss of a non-site, so the two are different quantities. The flag is required
because the score files do not record it. The source refuses to open for a different analysis build, and
every annotation carries spliceai_genome_build, spliceai_masked, and spliceai_source_release.
The default rule is spliceai_ds_max >= 0.5; callers may choose the provider's lower or higher cutoffs for a
different recall/precision tradeoff. Code, models, and score data have non-commercial licensing constraints.
The binding does not redistribute them.
AlphaMissense¶
altar-alphamissense aggregates transcript-level records. am_pathogenicity and am_class come from the
most pathogenic matching transcript. am_transcript_scores is a repeated struct that preserves every matching
transcript in descending score order, with its UniProt accession, transcript, protein variant, pathogenicity,
and class.
AlphaMissenseRelease declares the dataset's genome build (hg19 or hg38) and the exact upstream release.
The source refuses to open for a different analysis build. It also rejects any backend record whose native
genome value is another build. Every annotation carries am_genome_build and am_source_release.
The default rule selects likely_pathogenic and ambiguous classifications. It remains upstream predictive
evidence, not an Altar clinical classification.
gnomAD¶
altar-gnomad looks up gnomAD v4 population allele frequencies. GnomADRelease names gnomAD's exact release
version as it appears in the file names, such as 4.1.1 for the v4.1.1 files, and the dataset
(genomes, exomes, or joint). Its canonical release label, <version>-<dataset>, is the identity's
release. Only hg38 is accepted. gnomAD v2.1.1 uses other field names and ancestry groups, so its values
would not mean the same thing under these columns.
| Field | Interpretation |
|---|---|
gnomad_af, gnomad_ac, gnomad_an, gnomad_nhomalt |
Allele frequency, allele count, allele number, and homozygote count. gnomad_af is null when AN is 0. |
gnomad_af_grpmax, gnomad_grpmax |
gnomAD's grpmax frequency and the group it comes from. grpmax leaves out asj, fin, and remaining, plus ami and mid in genomes and ami in joint; exomes have no ami. |
gnomad_faf95_max |
The greatest filtering allele frequency (Poisson 95% CI) over groups. |
gnomad_af_<group> |
Frequency in afr, ami, amr, asj, eas, fin, mid, nfe, sas, and remaining. |
gnomad_af_max_ancestry |
The greatest non-null frequency over those ten groups in every dataset, as one scalar for filters. |
gnomad_filters |
The site's VCF FILTER value, PASS or the filters it failed. |
A variant gnomAD has no row for is omitted, not given a frequency of zero: an unknown frequency is not
evidence of rarity. The source has no default prioritization rule, because frequency is a filter. A rarity
filter that keeps variants gnomAD lacks is
Or(IsNull(Col("gnomad_af_max_ancestry")), Le(Col("gnomad_af_max_ancestry"), Lit(1e-4))), which renders the
same rule to Python and SQL. An AC0 site is stored with gnomad_af 0.0, so such a filter counts it as
observed and rare; add Eq(Col("gnomad_filters"), Lit("PASS")) or Gt(Col("gnomad_ac"), Lit(0)) if an
observed, passing frequency is what you mean.
Normalize multi-allelic cohorts before ingestion
gnomAD's alleles are split, trimmed, and left-aligned. Altar's VCF ingestion splits ALTs but keeps the
record's REF and does not trim or left-align. A cohort record chr1 100 CA C,TA yields chr1:100:CA:TA,
while gnomAD holds that allele as the SNV chr1:100:C:T. The variant finds no gnomAD row, and a filter in
which null passes then keeps it as rare. Run bcftools norm -m -any -f <reference> on any cohort with
multi-allelic records, SNVs included, or with indels.
altar-gnomad-build streams gnomAD's per-chromosome sites VCFs into Parquet keyed by canonical
chr:pos:ref:alt, reading the dataset's INFO names (the joint dataset's carry a _joint suffix). It writes
only into an empty output directory, rejects a variant key that occurs twice across its inputs, deletes what
it wrote if any input fails, and checks inputs named like gnomAD's files against the release. Each Parquet
file records the version, dataset, genome build, and builder version, and GnomADSource over the Parquet
backend refuses files built for another release. Any AllelicRecordBackend can serve the table. Most cohort
variants are in gnomAD, so for BigQuery materialization use bigquery_source, which reads a loaded copy of the
table through a relation= the store fuses into its query and declares the same identity, rather than
staged_annotation_source, which is for sparse sources.
gnomAD's primary data is released under CC0; its terms ask users to cite the release and not to attempt to re-identify participants. The binding downloads and redistributes no data. See the binding README for the column contract, build options, and filtering examples.
ClinVar¶
altar-clinvar looks up ClinVar's aggregate germline classifications. ClinVarRelease names the release by
the YYYYMMDD date in ClinVar's file names (clinvar_20260928.vcf.gz is 20260928), which is the identity's
release, and the genome_build: hg38 for ClinVar's GRCh38 files or hg19 for its GRCh37 files. ClinVar
builds both from the same records, so the columns mean the same thing on either assembly. Releases before
20240127 are rejected: until then CLNSIG mixed germline and somatic submissions, and the oncogenicity and
somatic clinical impact classifications had no fields of their own.
| Field | Interpretation |
|---|---|
clinvar_variation_id, clinvar_allele_id |
ClinVar's Variation ID and Allele ID. |
clinvar_classification |
The aggregate germline classification (CLNSIG) in ClinVar's VCF spelling, such as Pathogenic/Likely_pathogenic. Null for a record with no germline classification. |
clinvar_pathogenic |
Whether any /- or \|-separated term of the classification is Pathogenic or Likely_pathogenic. |
clinvar_low_penetrance |
Whether any term is Pathogenic,_low_penetrance or Likely_pathogenic,_low_penetrance. A low-penetrance classification alone is not clinvar_pathogenic. |
clinvar_review_status, clinvar_review_stars |
The germline review status and ClinVar's 0–4 gold stars for it; stars are null for a status the binding does not know. |
clinvar_conditions, clinvar_conflicting_classifications, clinvar_molecular_consequences |
Repeated: condition names, each conflicting submission's classification and count, and consequence names. |
clinvar_variant_type, clinvar_origin, clinvar_origins |
The variant type, the allele-origin bitmask, and the origins it decodes to. |
clinvar_genes |
Repeated struct: the genes ClinVar reports for the variant, each {symbol, gene_id} with its NCBI Gene ID. |
clinvar_oncogenicity, clinvar_somatic_impact |
The oncogenicity and somatic clinical impact classifications, each with a review-status column. |
A variant ClinVar has no record for is omitted: the absence of a submission is not evidence that a variant is
benign. By default the source has no prioritization rule, because ClinVar is curated evidence rather than a
model prediction and is often used as a filter or a display column. ClinVarSource(...,
prioritize_pathogenic=True) returns PATHOGENIC_PREDICATE,
And(Eq(Col("clinvar_pathogenic"), Lit(True)), Ge(Col("clinvar_review_stars"), Lit(1))), which leaves out
classifications submitted without assertion criteria. A host that also wants low-penetrance classifications
adds clinvar_low_penetrance to its own predicate.
altar-clinvar-build streams ClinVar's VCF into Parquet keyed by canonical chr:pos:ref:alt, with the same
safeguards as the gnomAD builder: an empty output directory, unique keys across inputs, removal of everything
it wrote on failure, and the release recorded in Parquet metadata, which ClinVarSource checks. It also checks
each header's ##fileDate and ##reference against the release. It writes 1 as chr1 and MT as chrM,
and skips and counts records with no ALT, a symbolic ALT, or an N base, and records on unplaced or
alternate contigs. An hg19 build also skips MT, because ClinVar's GRCh37 MT is the rCRS while UCSC hg19's
chrM is a different sequence. As with gnomAD, run bcftools norm -m -any -f <reference> on a cohort with multi-allelic
records or indels before ingestion, or such variants miss ClinVar's keys. For BigQuery, bigquery_source
reads a loaded copy of the table through a fused relation=.
NCBI places no restrictions on the use or distribution of ClinVar data and asks for attribution to ClinVar as a data source. ClinVar's information is not intended for direct diagnostic use without review by a genetics professional. The binding downloads and redistributes no data. See the binding README for the column contract, build options, and the full data terms.
Open Targets E2G¶
altar-opentargets-e2g normalizes the Open Targets enhancer_to_gene dataset into
VariantGeneLinkSource. Each relation preserves its regulatory element, target gene, biosample context,
scores, distances, dataset release, genome build, and source provenance.
OpenTargetsE2GImporter streams normalized records into any VariantGeneLinkStore; OpenTargetsE2GSource
queries the published evidence through the same storage-neutral filters and pagination as other relation
sources. The binding does not define PostgreSQL-, Parquet-, or service-specific query adapters. Operators
remain responsible for reading the upstream Parquet release and choosing a stable generation identifier for
that exact extraction. Import configuration also records whether the Open Targets pipeline's declared score
threshold was applied, preserving the distinction between a released filtered table and raw rE2G scores.
ENCODE-rE2G and scE2G atlases¶
altar-e2g-atlas ingests released ENCODE DCC rE2G BED3+ files and the pinned scE2G v1.3 multiome prediction
schema into version-retaining, indexed element-to-gene evidence. This is lookup of already generated atlas
rows, not execution of either scientific pipeline.
The binding preserves all meaningful source fields in typed rE2G/scE2G records, records the input SHA-256 and
threshold semantics, reports invalid rows, and keeps older accessions/releases. Both formats normalize into
VariantGeneEvidenceRecord; E2GAtlasSource then reads them from any VariantGeneLinkStore. Altar's generic
in-memory and SQLite stores run the same conformance suite and supply interval, gene, context, method/source,
release, score-bound, generation-visibility, and seek-pagination behavior. See the
binding README for supported schemas,
coordinates, thresholds, licensing, acquisition, and known uncertainties.
Asset responsibility¶
Operators acquire upstream datasets, review their licenses, choose storage, and preserve release manifests. Bindings fail on build or asset mismatches rather than silently lifting coordinates or substituting releases.