Skip to content

Annotation data files

Region classification, cCRE overlap, and allele-frequency annotation read three generated lookup files. The altar wheel does not include them. Build them once, or run annotation in the variants runtime image, which includes them.

File Used by Source Size
region_annotations.parquet region_annotations, region_filter, the preprocessing pipeline's region step, annotate Ensembl 116 GRCh38 GFF3 about 32M intervals
ccres.dnatree ccre_overlap, annotate SCREEN Registry v4 GRCh38 cCRE BED about 2.3M elements
variants.pkl.gz get_ot_variant, annotate Open Targets 25.03 variant shards (about 3.1 GB) about 320 MB

Nearest-gene lookup reads gene_df.tsv, which ships in the wheel: the gene records of GENCODE release 47 (Ensembl 113, GRCh38). A data directory may hold its own gene_df.tsv, which then replaces the shipped one for annotation that reads that directory.

region_provenance.json records where the region files came from; see Provenance.

Where Altar looks

Altar resolves the directory that holds these files in this order:

  1. The data_dir argument of the annotation functions, or the --data_dir option of altar.variants.annotation.annotate and altar.variants.preprocessing.region_filter, or --data-dir of altar.variants.preprocessing.pipeline.
  2. The ALTAR_VARIANTS_DATA_DIR environment variable.
  3. The data directory of the installed altar.variants package. The runtime image installs the files here. In an editable source checkout this is altar/altar/variants/data.

A missing file raises AnnotationDataNotFoundError, a FileNotFoundError. The message names the file, the directory searched, and the command that builds the file.

Build the files

Set the data directory first. The download scripts and build commands all default to it.

export ALTAR_VARIANTS_DATA_DIR="$HOME/altar-variants-data"

The download scripts are in altar/altar/variants/scripts/ of the repository and in altar/variants/scripts/ of the installed package. The examples run them from the repository root. With ALTAR_VARIANTS_DATA_DIR set, each one writes under $ALTAR_VARIANTS_DATA_DIR/raw from any directory. A missing-file error prints the installed scripts' full paths.

# Region annotations (Ensembl 116 GFF3)
bash altar/altar/variants/scripts/download_ensembl_gff.sh
python -m altar.variants.scripts.construct_region_annotations \
    -i "$ALTAR_VARIANTS_DATA_DIR/raw/Homo_sapiens.GRCh38.116.gff3.gz"

# cCRE index (SCREEN Registry v4)
bash altar/altar/variants/scripts/download_ccres.sh
python -m altar.variants.scripts.construct_ccre_dnatree

# Open Targets allele frequencies (needs rsync and about 7 GB of memory)
bash altar/altar/variants/scripts/download_variants.sh
python -m altar.variants.scripts.construct_variants_df

After the build, $ALTAR_VARIANTS_DATA_DIR/raw is no longer needed.

construct_region_annotations reads the Ensembl release from the GFF3 file name (or --ensembl-release), and construct_ccre_dnatree records SCREEN release v4 (or --release). Each records the input it read, with its SHA-256, in region_provenance.json. Ensembl 116 and SCREEN v4 are pinned: a file labelled with one of these releases must have the pinned SHA-256, and the command refuses before it builds anything otherwise.

The runtime image checks the SHA-256 of the downloaded cCRE BED, of ccres.dnatree, and of the decompressed variants.pkl.gz pickle. construct_ccre_dnatree prints the digest of the tree it wrote. To produce the same allele-frequency table as the image, follow runtimes/variants/README.md in the repository.

Provenance

region_provenance.json, in the same directory, records the inputs the region values come from, the region logic version that built and reads them, and the SHA-256 of each derived file:

{
  "format": 2,
  "genome_build": "hg38",
  "logic_version": 1,
  "inputs": {
    "ensembl_gff3": {"source": "Ensembl GRCh38 GFF3", "release": "116", "sha256": "…"},
    "ccre_bed": {"source": "SCREEN cCRE Registry GRCh38", "release": "v4", "sha256": "…"},
    "gene_table": {"source": "GENCODE GRCh38 gene annotation", "release": "gencode-47", "sha256": "…"}
  },
  "derived": {
    "region_annotations.parquet": {"sha256": "…", "logic_version": 1},
    "ccres.dnatree": {"sha256": "…", "logic_version": 1}
  }
}

The build commands above write each entry as they build the file, and the variants runtime image writes the manifest for the files it bakes in. The gene-table input is the table annotation reads: the data directory's own gene_df.tsv, or the shipped one. A gene table other than the shipped GENCODE 47 table is recorded with release unrecorded and its digest.

To record files built by this Altar version from inputs you still have, run:

python -m altar.variants.annotation.provenance \
    --ensembl-release 116 --ensembl-gff "$ALTAR_VARIANTS_DATA_DIR/raw/Homo_sapiens.GRCh38.116.gff3.gz" \
    --ccre-release v4 --ccre-bed "$ALTAR_VARIANTS_DATA_DIR/raw/GRCh38-cCREs.bed"

The command takes --data-dir, hashes the inputs and files, writes the manifest, and prints the release it yields. A tree with the runtime image's pinned v4 digest is itself evidence of the v4 BED, so --ccre-bed may be left out for it. The command never records a release from a label alone, refuses a pinned release whose input has another digest, and refuses to give an input whose bytes are already recorded a different release. A manifest from an earlier Altar, or from another region logic version, is refused; rebuild the files.

The region-annotation source identity is derived from the manifest. Its release names the inputs and the logic version, for example ensembl-116.ccre-v4.genes-gencode-47.logic-1, and adds a short digest of the input digests when an input is not a pinned release. It does not depend on the derived files' bytes. Annotation checks those files against the manifest before it loads them, so a manifest that no longer describes its directory is an error, not a mislabelled release. See Region annotation source.

gene_df.tsv entered the repository without the script that built it. Its rows are the gene records of GENCODE 47's gencode.v47.annotation.gtf, in file order: the gene name, chromosome, start and end ordered 5' to 3' (a minus-strand gene has start greater than end), strand, gene type, and the unversioned gene ID. Its 78,724 genes and its biotype totals equal the GENCODE 47 statistics, and its chr1 rows equal that GTF's chr1 gene records field for field. Its SHA-256 is 52807cc8f3e74645d1a2ebc27b95081fc031a56d766aa3f31e3114c5315cc370.

Trust

The provenance checks catch drift and accidents. They do not protect against someone who can write the data directory, who can rewrite the manifest as well.

ccres.dnatree and variants.pkl.gz are Python pickles, and loading a pickle can run arbitrary code. Point ALTAR_VARIANTS_DATA_DIR only at files you built or that come from a source you trust.