Annotation data files¶
Region classification, cCRE overlap, and allele-frequency annotation read three generated lookup files. The
altar wheel does not include them. Build them once, or run annotation in the variants runtime image,
which includes them.
| File | Used by | Source | Size |
|---|---|---|---|
region_annotations.parquet |
region_annotations, region_filter, the preprocessing pipeline's region step, annotate |
Ensembl 116 GRCh38 GFF3 | about 32M intervals |
ccres.dnatree |
ccre_overlap, annotate |
SCREEN Registry v4 GRCh38 cCRE BED | about 2.3M elements |
variants.pkl.gz |
get_ot_variant, annotate |
Open Targets 25.03 variant shards (about 3.1 GB) | about 320 MB |
Nearest-gene lookup reads gene_df.tsv, which ships in the wheel: the gene records of GENCODE release 47
(Ensembl 113, GRCh38). A data directory may hold its own gene_df.tsv, which then replaces the shipped one for
annotation that reads that directory.
region_provenance.json records where the region files came from; see Provenance.
Where Altar looks¶
Altar resolves the directory that holds these files in this order:
- The
data_dirargument of the annotation functions, or the--data_diroption ofaltar.variants.annotation.annotateandaltar.variants.preprocessing.region_filter, or--data-dirofaltar.variants.preprocessing.pipeline. - The
ALTAR_VARIANTS_DATA_DIRenvironment variable. - The
datadirectory of the installedaltar.variantspackage. The runtime image installs the files here. In an editable source checkout this isaltar/altar/variants/data.
A missing file raises AnnotationDataNotFoundError, a FileNotFoundError. The message names the file, the
directory searched, and the command that builds the file.
Build the files¶
Set the data directory first. The download scripts and build commands all default to it.
The download scripts are in altar/altar/variants/scripts/ of the repository and in altar/variants/scripts/
of the installed package. The examples run them from the repository root. With ALTAR_VARIANTS_DATA_DIR set,
each one writes under $ALTAR_VARIANTS_DATA_DIR/raw from any directory. A missing-file error prints the
installed scripts' full paths.
# Region annotations (Ensembl 116 GFF3)
bash altar/altar/variants/scripts/download_ensembl_gff.sh
python -m altar.variants.scripts.construct_region_annotations \
-i "$ALTAR_VARIANTS_DATA_DIR/raw/Homo_sapiens.GRCh38.116.gff3.gz"
# cCRE index (SCREEN Registry v4)
bash altar/altar/variants/scripts/download_ccres.sh
python -m altar.variants.scripts.construct_ccre_dnatree
# Open Targets allele frequencies (needs rsync and about 7 GB of memory)
bash altar/altar/variants/scripts/download_variants.sh
python -m altar.variants.scripts.construct_variants_df
After the build, $ALTAR_VARIANTS_DATA_DIR/raw is no longer needed.
construct_region_annotations reads the Ensembl release from the GFF3 file name (or --ensembl-release), and
construct_ccre_dnatree records SCREEN release v4 (or --release). Each records the input it read, with its
SHA-256, in region_provenance.json. Ensembl 116 and SCREEN v4 are pinned: a file labelled with one of
these releases must have the pinned SHA-256, and the command refuses before it builds anything otherwise.
The runtime image checks the SHA-256 of the downloaded cCRE BED, of ccres.dnatree, and of the decompressed
variants.pkl.gz pickle. construct_ccre_dnatree prints the digest of the tree it wrote. To produce the same allele-frequency table as the image, follow runtimes/variants/README.md in the
repository.
Provenance¶
region_provenance.json, in the same directory, records the inputs the region values come from, the region
logic version that built and reads them, and the SHA-256 of each derived file:
{
"format": 2,
"genome_build": "hg38",
"logic_version": 1,
"inputs": {
"ensembl_gff3": {"source": "Ensembl GRCh38 GFF3", "release": "116", "sha256": "…"},
"ccre_bed": {"source": "SCREEN cCRE Registry GRCh38", "release": "v4", "sha256": "…"},
"gene_table": {"source": "GENCODE GRCh38 gene annotation", "release": "gencode-47", "sha256": "…"}
},
"derived": {
"region_annotations.parquet": {"sha256": "…", "logic_version": 1},
"ccres.dnatree": {"sha256": "…", "logic_version": 1}
}
}
The build commands above write each entry as they build the file, and the variants runtime image writes the
manifest for the files it bakes in. The gene-table input is the table annotation reads: the data directory's own
gene_df.tsv, or the shipped one. A gene table other than the shipped GENCODE 47 table is recorded with release
unrecorded and its digest.
To record files built by this Altar version from inputs you still have, run:
python -m altar.variants.annotation.provenance \
--ensembl-release 116 --ensembl-gff "$ALTAR_VARIANTS_DATA_DIR/raw/Homo_sapiens.GRCh38.116.gff3.gz" \
--ccre-release v4 --ccre-bed "$ALTAR_VARIANTS_DATA_DIR/raw/GRCh38-cCREs.bed"
The command takes --data-dir, hashes the inputs and files, writes the manifest, and prints the release it
yields. A tree with the runtime image's pinned v4 digest is itself evidence of the v4 BED, so --ccre-bed may be
left out for it. The command never records a release from a label alone, refuses a pinned release whose input
has another digest, and refuses to give an input whose bytes are already recorded a different release. A
manifest from an earlier Altar, or from another region logic version, is refused; rebuild the files.
The region-annotation source identity is derived from the manifest. Its release names the inputs and the logic
version, for example ensembl-116.ccre-v4.genes-gencode-47.logic-1, and adds a short digest of the input digests
when an input is not a pinned release. It does not depend on the derived files' bytes. Annotation checks those
files against the manifest before it loads them, so a manifest that no longer describes its directory is an
error, not a mislabelled release. See Region annotation source.
gene_df.tsv entered the repository without the script that built it. Its rows are the gene records of
GENCODE 47's gencode.v47.annotation.gtf, in file order: the gene name, chromosome, start and end ordered 5' to
3' (a minus-strand gene has start greater than end), strand, gene type, and the unversioned gene ID. Its 78,724
genes and its biotype totals equal the GENCODE 47 statistics, and its chr1 rows equal that GTF's chr1 gene records
field for field. Its SHA-256 is 52807cc8f3e74645d1a2ebc27b95081fc031a56d766aa3f31e3114c5315cc370.
Trust¶
The provenance checks catch drift and accidents. They do not protect against someone who can write the data directory, who can rewrite the manifest as well.
ccres.dnatree and variants.pkl.gz are Python pickles, and loading a pickle can run arbitrary code. Point
ALTAR_VARIANTS_DATA_DIR only at files you built or that come from a source you trust.