Skip to content

Latest commit

 

History

67 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

DANSy_Applications

Here, we provide example applications of our Domain Architecture Network Syntax (DANSy) to different applications on either the whole human proteome, post-translational modification (PTM) systems, or fusion genes. This results of these applications are summarized in our bioRxiv paper and here we provide the code to produce the results and figures in the manuscript.

How to cite: Please cite our bioRxiv paper

DANSy Overview

Overview of the general workflow

Getting started

For this work, create a local copy of this repository by copying the following into your terminal:

git clone https://github.com/NaegleLab/DANSy_Applications

Then create a virtual environment with all the dependencies of this repo using:

conda create env -f dansy_apps.yml

Activate the environment using conda activate dansy_apps for specific scripts or select the dansy_apps kernel for Jupyter notebooks.

Proteome Reference File

DANSy relies on reference files generated by CoDIAC. We have provided the reference files for the analysis conducted in our manuscript in the data folder. This version of the reference file was generated on July 15, 2026 (UniProt version v2026_02 and InterPro v109), and will be the default file used across all applications.

If you wish to generate a new reference file to use for analysis, take the following steps.

  • First download the SwissProt MetaData file from Gencode into your local copy of this repo.

  • Then, modify the whole_proteome_reference.py file by changing the reference file suffix variable to the current date and verify the gencode file name is correct. Activate your dansy environment and run the following code. (Note: This can take up to 2 hours after a fresh install as it will also establish a biomart sqlite database for use in other notebooks.)

    conda activate dansy_apps
    python scripts/whole_proteome_reference.py
    conda deactivate
    

To use the new build change the instances for the dansy.import_reference_files() commands to look for your reference file instead by specifying the exact directory and file suffix used import_proteome_files(ref_file_dir='data/path/to/reference_file',ref_file_suffix=new_fetch_date_suffix).

Additional Datasets

For more specific analysis, the following datasets should be downloaded from their source.

Analysis Dataset Source/Code
Fusion Gene Analysis ChimerSeq Excel File from ChimerDB
PTM Systems Provided in the Multispecies Reference Files imported from UniProt using the reference proteome fetching script and reference file generating script
Cancer Cell Line Encyclopedia (CCLE) Fusion and CRISPR Screen Data From the DepMap project. We used the 26Q1 files. We specifically downloaded the 1) OmicsFusionFilteredSupplementary.csv 2) Models.csv and 3) CRISPRGeneEffect.csv files.

What is provided and how to run the code.

We have provided several Jupyter notebooks that serve as examples of how to use DANSy. Below are short summaries of their applications and the results, which are discussed in our manuscript.

For custom uses of DANSy, please visit our DANSy repository, which provides more general use cases and information on how to get started with the DANSy package for new datasets.

Complete human proteome analysis

Analyzing the domain architectures across the human proteome and broadly characterizing the resulting network and information encoded by different versions related to n-gram length. Here, we find limiting n-grams to 10-grams in length will recapitulate network characteristics of the complete proteome and that specific domains such as the protein kinase, zinc finger C2H2, and EGF-like domains are n-grams that frequently lie along the shortest path and are connected to the most other n-gram nodes in the network.

Associated notebooks:

  • human_proteome The broad characterization of the human proteome to explore DANSy in general.
  • human_model_comp The comparison of different n-gram lengths used to construct DANSy models. Note: this is one of the most computationally demanding notebooks.
  • human_proteome_correlation_analysis Additional correlative analysis of network centrality measurements to protein counts and a comparison to randomized null models.

Additional Files:

  • entropyCalc Utility functions to quickly calculate some entropy metrics.

PTM System Analysis

Focused on reversible post-translational modification systems (e.g. phosphorylation, methylation, acetylation) that operate under a reader-writer-eraser paradigm. Characterizing broad properties of how individual components combine in domain architectures and identifying general grammatical rules where eraser domains do not require additional reader domains to modify their activity. Meanwhile, reader domains will more frequently create domain combinations with writer domains, but rarely with eraser domains.

Associated notebooks:

PTM System Evolution

Additional analysis on the phosphorylation systems compares the network characteristics of domains associated with phosphotyrosine (pTyr) and phosphoserine/threonine (pSer/Thr) systems during the evolutionary period from yeast to humans. The pTyr system is evolutionarily younger and rapidly expanded during the transition to metazoans. Our analysis suggests that during this transition, species were sampling several configurations of the network before converging to a similar set of grammatical rules observed in other PTM systems.

Associated notebooks:

Fusion Gene Analysis

Here, we analyze how fusion genes may provide an avenue to explore new grammatical structures of domain architectures. We studied fusion genes reported in TCGA and CCLE and their predicted chimera protein domain architectures. We find that the domain architectures of gene fusions rarely generate novel domain architectures to suggest they largely respect existing domain combination rules of the natural proteome. Using the CCLE CRISPR screen data, we do not observe specific domains to be highly enriched in fusions whose partner genes create dependency effects in the cell line models. However, we did observe that kinase domains were the most frequently involved in gene fusions. Thus, we further explored kinase fusions to understand if the flexibility in generating domain combinations in then natural proteome reflected its tendency to form fusions.

Associated notebooks:

  • tcga_fusion_gene Baseline comparison of fusions from ChimerDB on TCGA patients samples
  • ccle_fusions Expansion of fusion analysis to compliment TCGA results with CCLE based analysis.
  • ccle_fusion_contribution_analysis Analyzing how specific domains are enriched in novel domain architectures and contributions to fusions whose partner genes suggest dependency effects.
  • kinase_fusion_analysis Specific breakdown of differences between tyrosine kinase fusions and serine/threonine fusions and how they recapture syntax rules established in the proteome.

Addtional files to enable fusion analysis:

  • fusionClasses A python class to create individual and collections of fusion genes and hold the results together.
  • fusionGeneAnalysis Utilities to establish the fusion domain architectures from genome coordinates.
  • domain_buffers Establishing the file for the domain amino acid buffers used for annotating domain architectures in fusion genes.
  • fpr_stability Analysis to evaluate the stability of FPR values used in domain enrichment analysis.
  • fusion_null_check Generates multiple collections of randomized fusion gene fragment pairs and builds domain architectures for the predicted chimera proteins to act as background distributions to determine FPR values. Note: This notebook can take several hours (>8hrs) to generate all the values because it is creating >=100 collections of 15000 gene fragments

Supplementary notebooks

The Jupyter notebooks provided above relate to specific topics, but there are additional notebooks which are self-contained vignettes to complement the results. Short descriptions and links to them are provided below:

About

Example applications of the DANSy linguistic network analysis

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages