ENCODE GRAMMAR: Model tracks

Description

This track shares genome-wide resources from ENCODE GRAMMAR (ENCODE Genomic Regulatory Atlas of sequence Models, Motifs, Annotations and Rules) a collection of 3,888 experiment-specific model sets across the ENCODE data compendium, spanning several layers of gene regulation, including the binding of regulatory proteins to DNA, chromatin accessibility, transcription initiation, and regulatory activity measured using high-throughput reporter assays. For each experiment, we also release model-predicted, de-noised biochemical activity profiles at single-base resolution; model-estimated contributions of individual DNA bases to the activity of each regulatory element in each cellular context; recurring predictive DNA patterns (motifs) learned by the models; the genomic locations of predictive motif instances, into 43,237 browser-ready files on the hg38 genome assembly from BPNet, ChromBPNet, ProCapNet, ReporterNet models.

The BPNet family contains four related Deep Learning Models of regulatory DNA sequence (sometimes called sequence-to-function models). Each model uses a related BPNet-style convolutional neural network (CNN) architecture to train on, predict, and explain a distinct regulatory functional assay:

As part of the ENCODE Project, we trained 2,339 BPNet models on TF ChIP-seq experiments across 788 transcription factors and 175 biosamples; 1,512 ChromBPNet models trained on DNase-seq and ATAC-seq experiments across 407 biosamples; 6 ProCapNet models trained on PRO-cap experiments across 6 biosamples; and 5 ReporterNet models trained on lentiMPRA, ATAC-STARR, WG-STARR, and aggregated MPRA datasets across 4 biosamples.

We provide browser-ready tracks of the models' outputs for an easy interactive session. For each model, we provide the following types of output tracks:

The model outputs were generated by processing each model through a uniform model training, prediction, and interpretation workflow:

All other non-browser-related files are available at the ENCODE Portal.

For more information about the process workflow, see the ENCODE 4 Flagship paper.

Display conventions and configuration

Click a model collection to view available datasets. Subtracks can be further filtered by experiment accession, model accession, assay, target, biosample, tissue, organ, system, data type, strand, score type, quality-control flag, and file accession. Normalized signal tracks are recommended for visualization and comparison across datasets. For stranded assays, plus and minus strand signal tracks are provided separately.

Data access and use

The ENCODE GRAMMAR resource on the UCSC Genome Browser can be explored interactively with the Table Browser or the Data Integrator. For automated download and analysis, each subtrack points directly to an ENCODE bigWig or bigBed file. The data may also be explored interactively using the UCSC REST API. The original data files are available from the ENCODE portal. Clicking any accession in the track's configuration table links directly to the corresponding ENCODE experiment, model, or file details page.

External data users may freely download, analyze and publish results based on any ENCODE data without restrictions. This applies to all datasets, regardless of type or size, and includes no grace period for ENCODE data producers, either as individual members or as part of the Consortium. Researchers using unpublished ENCODE data are encouraged to contact the data producers to discuss possible coordinated publications; however, this is optional. The Consortium will continue to publish the results of its own analysis efforts in independent publications.

Credits

Data were generated by the ENCODE Consortium through ENCODE production and analysis laboratories. Model training, interpretation, browser-ready packaging, and trackhub integration were developed as part of ENCODE deep learning model efforts, including work from the Kundaje lab.

References

Avsec Z et al. Base-resolution models of transcription-factor binding reveal soft motif syntax. Nat Genet. 2021;53:354-366.

Pampari A et al. ChromBPNet: bias factorized, base-resolution Deep Learning Models of chromatin accessibility reveal cis-regulatory sequence syntax, transcription factor footprints and regulatory variants. bioRxiv. 2025.

Cochran K et al. Dissecting the cis-regulatory syntax of transcription initiation with deep learning. bioRxiv. 2024.

Contact

For questions, contact: Chang M. Yun (chang.m.yun@stanford.edu), Vivekanandan Ramalingam (vir@stanford.edu), Vivian Hecht (vhecht@stanford.edu), Anshul Kundaje (akundaje@stanford.edu).