Skip to content

Precomputed evidence

These bindings retrieve released evidence and do not execute model inference during an Altar analysis.

Each annotation binding publishes its columns as an AnnotationContract and declares an AnnotationSourceIdentity from identity(). Another source, such as a host's warehouse table, can implement the same contract; see Use your own annotation tables.

Binding Contract constant source_id Identity release
GPN-Star GPNSTAR_ANNOTATIONS org.kundajelab.altar.annotation.gpnstar <score_set>@<source_release_revision>
SpliceAI SPLICEAI_ANNOTATIONS org.kundajelab.altar.annotation.spliceai <source_release>-masked or <source_release>-raw
AlphaMissense ALPHAMISSENSE_ANNOTATIONS org.kundajelab.altar.annotation.alphamissense <source_release>
gnomAD GNOMAD_ANNOTATIONS org.kundajelab.altar.annotation.gnomad <version>-<dataset>
ClinVar CLINVAR_ANNOTATIONS org.kundajelab.altar.annotation.clinvar <YYYYMMDD> release date

All five contracts are at version 1.0.0. The per-row provenance columns below remain part of the GPN-Star, SpliceAI, and AlphaMissense contracts. The gnomAD and ClinVar contracts have none: their release is only in the identity.

GPN-Star

altar-gpnstar retrieves allele-specific scores from a pinned GPN-Star release. It supports only biallelic SNVs covered by the selected assembly and whole-genome alignment.

Field Interpretation
gpnstar_llr_calibrated Mutation-rate-calibrated log-likelihood ratio of alternate versus reference; more negative indicates stronger evolutionary constraint or predicted effect.
gpnstar_abs_llr_calibrated Independently calibrated absolute-LLR statistic; it is not abs(gpnstar_llr_calibrated).

The source also emits checkpoint, alignment, species set, genome build, release revision, and source-asset checksum. It intentionally defines no default prioritization threshold.

SpliceAI

altar-spliceai normalizes the precomputed Illumina release. A variant may have several gene records. Scalar fields come from one coherent record with the greatest delta score, while spliceai_gene_scores preserves every record as repeated structured detail.

SpliceAIRelease declares the dataset's genome build (hg19 or hg38), whether it holds Illumina's masked or raw scores (masked), and the upstream release. Masked files set a delta score to 0 when it reports a gain of an existing splice site or a loss of a non-site, so the two are different quantities. The flag is required because the score files do not record it. The source refuses to open for a different analysis build, and every annotation carries spliceai_genome_build, spliceai_masked, and spliceai_source_release.

The default rule is spliceai_ds_max >= 0.5; callers may choose the provider's lower or higher cutoffs for a different recall/precision tradeoff. Code, models, and score data have non-commercial licensing constraints. The binding does not redistribute them.

AlphaMissense

altar-alphamissense aggregates transcript-level records. am_pathogenicity and am_class come from the most pathogenic matching transcript. am_transcript_scores is a repeated struct that preserves every matching transcript in descending score order, with its UniProt accession, transcript, protein variant, pathogenicity, and class.

AlphaMissenseRelease declares the dataset's genome build (hg19 or hg38) and the exact upstream release. The source refuses to open for a different analysis build. It also rejects any backend record whose native genome value is another build. Every annotation carries am_genome_build and am_source_release.

The default rule selects likely_pathogenic and ambiguous classifications. It remains upstream predictive evidence, not an Altar clinical classification.

gnomAD

altar-gnomad looks up gnomAD v4 population allele frequencies. GnomADRelease names gnomAD's exact release version as it appears in the file names, such as 4.1.1 for the v4.1.1 files, and the dataset (genomes, exomes, or joint). Its canonical release label, <version>-<dataset>, is the identity's release. Only hg38 is accepted. gnomAD v2.1.1 uses other field names and ancestry groups, so its values would not mean the same thing under these columns.

Field Interpretation
gnomad_af, gnomad_ac, gnomad_an, gnomad_nhomalt Allele frequency, allele count, allele number, and homozygote count. gnomad_af is null when AN is 0.
gnomad_af_grpmax, gnomad_grpmax gnomAD's grpmax frequency and the group it comes from. grpmax leaves out asj, fin, and remaining, plus ami and mid in genomes and ami in joint; exomes have no ami.
gnomad_faf95_max The greatest filtering allele frequency (Poisson 95% CI) over groups.
gnomad_af_<group> Frequency in afr, ami, amr, asj, eas, fin, mid, nfe, sas, and remaining.
gnomad_af_max_ancestry The greatest non-null frequency over those ten groups in every dataset, as one scalar for filters.
gnomad_filters The site's VCF FILTER value, PASS or the filters it failed.

A variant gnomAD has no row for is omitted, not given a frequency of zero: an unknown frequency is not evidence of rarity. The source has no default prioritization rule, because frequency is a filter. A rarity filter that keeps variants gnomAD lacks is Or(IsNull(Col("gnomad_af_max_ancestry")), Le(Col("gnomad_af_max_ancestry"), Lit(1e-4))), which renders the same rule to Python and SQL. An AC0 site is stored with gnomad_af 0.0, so such a filter counts it as observed and rare; add Eq(Col("gnomad_filters"), Lit("PASS")) or Gt(Col("gnomad_ac"), Lit(0)) if an observed, passing frequency is what you mean.

Normalize multi-allelic cohorts before ingestion

gnomAD's alleles are split, trimmed, and left-aligned. Altar's VCF ingestion splits ALTs but keeps the record's REF and does not trim or left-align. A cohort record chr1 100 CA C,TA yields chr1:100:CA:TA, while gnomAD holds that allele as the SNV chr1:100:C:T. The variant finds no gnomAD row, and a filter in which null passes then keeps it as rare. Run bcftools norm -m -any -f <reference> on any cohort with multi-allelic records, SNVs included, or with indels.

altar-gnomad-build streams gnomAD's per-chromosome sites VCFs into Parquet keyed by canonical chr:pos:ref:alt, reading the dataset's INFO names (the joint dataset's carry a _joint suffix). It writes only into an empty output directory, rejects a variant key that occurs twice across its inputs, deletes what it wrote if any input fails, and checks inputs named like gnomAD's files against the release. Each Parquet file records the version, dataset, genome build, and builder version, and GnomADSource over the Parquet backend refuses files built for another release. Any AllelicRecordBackend can serve the table. Most cohort variants are in gnomAD, so for BigQuery materialization use bigquery_source, which reads a loaded copy of the table through a relation= the store fuses into its query and declares the same identity, rather than staged_annotation_source, which is for sparse sources.

gnomAD's primary data is released under CC0; its terms ask users to cite the release and not to attempt to re-identify participants. The binding downloads and redistributes no data. See the binding README for the column contract, build options, and filtering examples.

ClinVar

altar-clinvar looks up ClinVar's aggregate germline classifications. ClinVarRelease names the release by the YYYYMMDD date in ClinVar's file names (clinvar_20260928.vcf.gz is 20260928), which is the identity's release, and the genome_build: hg38 for ClinVar's GRCh38 files or hg19 for its GRCh37 files. ClinVar builds both from the same records, so the columns mean the same thing on either assembly. Releases before 20240127 are rejected: until then CLNSIG mixed germline and somatic submissions, and the oncogenicity and somatic clinical impact classifications had no fields of their own.

Field Interpretation
clinvar_variation_id, clinvar_allele_id ClinVar's Variation ID and Allele ID.
clinvar_classification The aggregate germline classification (CLNSIG) in ClinVar's VCF spelling, such as Pathogenic/Likely_pathogenic. Null for a record with no germline classification.
clinvar_pathogenic Whether any /- or \|-separated term of the classification is Pathogenic or Likely_pathogenic.
clinvar_low_penetrance Whether any term is Pathogenic,_low_penetrance or Likely_pathogenic,_low_penetrance. A low-penetrance classification alone is not clinvar_pathogenic.
clinvar_review_status, clinvar_review_stars The germline review status and ClinVar's 0–4 gold stars for it; stars are null for a status the binding does not know.
clinvar_conditions, clinvar_conflicting_classifications, clinvar_molecular_consequences Repeated: condition names, each conflicting submission's classification and count, and consequence names.
clinvar_variant_type, clinvar_origin, clinvar_origins The variant type, the allele-origin bitmask, and the origins it decodes to.
clinvar_genes Repeated struct: the genes ClinVar reports for the variant, each {symbol, gene_id} with its NCBI Gene ID.
clinvar_oncogenicity, clinvar_somatic_impact The oncogenicity and somatic clinical impact classifications, each with a review-status column.

A variant ClinVar has no record for is omitted: the absence of a submission is not evidence that a variant is benign. By default the source has no prioritization rule, because ClinVar is curated evidence rather than a model prediction and is often used as a filter or a display column. ClinVarSource(..., prioritize_pathogenic=True) returns PATHOGENIC_PREDICATE, And(Eq(Col("clinvar_pathogenic"), Lit(True)), Ge(Col("clinvar_review_stars"), Lit(1))), which leaves out classifications submitted without assertion criteria. A host that also wants low-penetrance classifications adds clinvar_low_penetrance to its own predicate.

altar-clinvar-build streams ClinVar's VCF into Parquet keyed by canonical chr:pos:ref:alt, with the same safeguards as the gnomAD builder: an empty output directory, unique keys across inputs, removal of everything it wrote on failure, and the release recorded in Parquet metadata, which ClinVarSource checks. It also checks each header's ##fileDate and ##reference against the release. It writes 1 as chr1 and MT as chrM, and skips and counts records with no ALT, a symbolic ALT, or an N base, and records on unplaced or alternate contigs. An hg19 build also skips MT, because ClinVar's GRCh37 MT is the rCRS while UCSC hg19's chrM is a different sequence. As with gnomAD, run bcftools norm -m -any -f <reference> on a cohort with multi-allelic records or indels before ingestion, or such variants miss ClinVar's keys. For BigQuery, bigquery_source reads a loaded copy of the table through a fused relation=.

NCBI places no restrictions on the use or distribution of ClinVar data and asks for attribution to ClinVar as a data source. ClinVar's information is not intended for direct diagnostic use without review by a genetics professional. The binding downloads and redistributes no data. See the binding README for the column contract, build options, and the full data terms.

Open Targets E2G

altar-opentargets-e2g normalizes the Open Targets enhancer_to_gene dataset into VariantGeneLinkSource. Each relation preserves its regulatory element, target gene, biosample context, scores, distances, dataset release, genome build, and source provenance.

OpenTargetsE2GImporter streams normalized records into any VariantGeneLinkStore; OpenTargetsE2GSource queries the published evidence through the same storage-neutral filters and pagination as other relation sources. The binding does not define PostgreSQL-, Parquet-, or service-specific query adapters. Operators remain responsible for reading the upstream Parquet release and choosing a stable generation identifier for that exact extraction. Import configuration also records whether the Open Targets pipeline's declared score threshold was applied, preserving the distinction between a released filtered table and raw rE2G scores.

ENCODE-rE2G and scE2G atlases

altar-e2g-atlas ingests released ENCODE DCC rE2G BED3+ files and the pinned scE2G v1.3 multiome prediction schema into version-retaining, indexed element-to-gene evidence. This is lookup of already generated atlas rows, not execution of either scientific pipeline.

The binding preserves all meaningful source fields in typed rE2G/scE2G records, records the input SHA-256 and threshold semantics, reports invalid rows, and keeps older accessions/releases. Both formats normalize into VariantGeneEvidenceRecord; E2GAtlasSource then reads them from any VariantGeneLinkStore. Altar's generic in-memory and SQLite stores run the same conformance suite and supply interval, gene, context, method/source, release, score-bound, generation-visibility, and seek-pagination behavior. See the binding README for supported schemas, coordinates, thresholds, licensing, acquisition, and known uncertainties.

Asset responsibility

Operators acquire upstream datasets, review their licenses, choose storage, and preserve release manifests. Bindings fail on build or asset mismatches rather than silently lifting coordinates or substituting releases.