scEnrichment¶
In this tutorial, we will provide a quick guideline for applying the scEnrichment feature of discotoolkit on DEGs (Differentially Expressed Genes). Follow these steps:
- First, we load the DEGs from Example 1, available on the DISCO website.
- Finally, we convert the retrieved data into a Pandas DataFrame, which will serve as the input for the
dt.CELLiD_enrichmentfunction. By running this function, we can obtain the desired enrichment analysis.
In [1]:
Copied!
# for google colab
# pip install discotoolkit
# import package
import discotoolkit as dt
import scanpy as sc
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
# adding ignore warning to clean the code
import warnings
# Ignore all warnings
warnings.filterwarnings('ignore')
%load_ext autoreload
%autoreload 2
sns.set_theme(rc={'figure.dpi': 300})
# for google colab
# pip install discotoolkit
# import package
import discotoolkit as dt
import scanpy as sc
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
# adding ignore warning to clean the code
import warnings
# Ignore all warnings
warnings.filterwarnings('ignore')
%load_ext autoreload
%autoreload 2
sns.set_theme(rc={'figure.dpi': 300})
The user can input either a gene list or a gene list with log fold change as the input to the dt.CELLiD_enrichment function.
Example of the DEGs:
In [2]:
Copied!
# testing DEG from DEGs of acinar cell(Peng, Junya et al.)
# Reference is in DISCO website
test_genes = {"gene":["PRSS1", "CTRB2", "CELA3A", "CTRB1", "REG1B", "CLPS", "CPB1", "CPA1", "PLA2G1B", "REG3A", "CTRC"],
# "fc":[5.236605052, 5.179753462, 4.724315075, 4.702706704, 4.65145949, 4.513887613, 4.351968886, 4.311988687, 4.272185339, 4.253882194, 4.208933992]
}
deg_df = pd.DataFrame(test_genes)
deg_df.head()
# testing DEG from DEGs of acinar cell(Peng, Junya et al.)
# Reference is in DISCO website
test_genes = {"gene":["PRSS1", "CTRB2", "CELA3A", "CTRB1", "REG1B", "CLPS", "CPB1", "CPA1", "PLA2G1B", "REG3A", "CTRC"],
# "fc":[5.236605052, 5.179753462, 4.724315075, 4.702706704, 4.65145949, 4.513887613, 4.351968886, 4.311988687, 4.272185339, 4.253882194, 4.208933992]
}
deg_df = pd.DataFrame(test_genes)
deg_df.head()
Out[2]:
| gene | |
|---|---|
| 0 | PRSS1 |
| 1 | CTRB2 |
| 2 | CELA3A |
| 3 | CTRB1 |
| 4 | REG1B |
Now that we have our data ready, we can proceed to run the function.
In [3]:
Copied!
# get enrichment analysis result using dt.CELLiD_enrichment
results = dt.CELLiD_enrichment(deg_df)
# get enrichment analysis result using dt.CELLiD_enrichment
results = dt.CELLiD_enrichment(deg_df)
INFO:root:Downloading gene set reference...
INFO:root:Comparing the ranked gene list to reference gene sets...
[Parallel(n_jobs=1)]: Done 40 tasks | elapsed: 0.3s
[Parallel(n_jobs=1)]: Done 161 tasks | elapsed: 0.9s
[Parallel(n_jobs=1)]: Done 364 tasks | elapsed: 1.6s
[Parallel(n_jobs=1)]: Done 647 tasks | elapsed: 2.9s
[Parallel(n_jobs=1)]: Done 1012 tasks | elapsed: 4.3s
[Parallel(n_jobs=1)]: Done 1253 out of 1253 | elapsed: 5.5s finished
In [4]:
Copied!
# get the top 10 results
results.head(10)
# get the top 10 results
results.head(10)
Out[4]:
| pval | or | name | gene | background | overlap | geneset | |
|---|---|---|---|---|---|---|---|
| 9 | 0.0 | 700.394 | Marker of Acinar cell in pancreas_cell | CTRC,CELA3A,REG1B,REG3A,CPB1,CLPS,CPA1,PRSS1,P... | 4228 | 11 | 81 |
| 3 | 0.0 | 609.570 | Marker of Acinar cell in PDAC_pancreas | CTRC,CELA3A,CPB1,CLPS,CPA1,PRSS1,PLA2G1B,CTRB2... | 4105 | 11 | 89 |
| 22 | 0.0 | 317.643 | Marker of Paneth cell in intestine | REG1B,REG3A | 4493 | 2 | 43 |
| 13 | 0.0 | 278.933 | Marker of Acinar cell in type_1_diabetes_pancreas | CTRC,CELA3A,REG1B,CPB1,CLPS,CPA1,PRSS1,PLA2G1B... | 1104 | 11 | 55 |
| 0 | 0.0 | 159.589 | Marker of Acinar cell in type_2_diabetes_pancreas | CTRC,CELA3A,REG1B,REG3A,CPB1,CLPS,CPA1,PRSS1,P... | 1543 | 11 | 117 |
| 21 | 0.0 | 154.783 | type 1 diabetes vs control for T cell in type_... | REG1B,PRSS1,REG3A,CTRB2,CLPS,CTRB1,CELA3A,PLA2... | 1104 | 9 | 31 |
| 20 | 0.0 | 117.889 | type 1 diabetes vs control for Perivascular ce... | REG3A,REG1B,PRSS1,CELA3A,CTRB2,CLPS,PLA2G1B,CP... | 1104 | 9 | 38 |
| 16 | 0.0 | 109.330 | type 1 diabetes vs control for Delta cell in t... | REG3A,REG1B,PRSS1,CLPS,CELA3A,PLA2G1B,CTRB1,CTRB2 | 1104 | 8 | 29 |
| 14 | 0.0 | 89.915 | type 1 diabetes vs control for Alpha cell in t... | REG3A,REG1B,PRSS1,CTRB2,CELA3A,CTRB1,CLPS,PLA2... | 1104 | 9 | 47 |
| 18 | 0.0 | 39.280 | type 1 diabetes vs control for Fibroblast in t... | REG1B,REG3A,CTRB2,PRSS1,CTRB1,PLA2G1B,CPB1,CEL... | 1104 | 10 | 143 |
We can visualize the results using a horizontal barplot with the help of the seaborn library.
In [5]:
Copied!
# setting the theme to whitegrid
sns.set_theme(style="whitegrid")
# Initialize the matplotlib figure
f, ax = plt.subplots(figsize=(6, 15))
# plot only the top 20 results
sns.barplot(x="or", y="name", data=results.head(20), color="b")
# adding axis title
ax.set(ylabel="Gene sets",
xlabel="Odds Ratio")
# make the plot look cleaner and nicer
sns.despine(left=True, bottom=True)
# setting the theme to whitegrid
sns.set_theme(style="whitegrid")
# Initialize the matplotlib figure
f, ax = plt.subplots(figsize=(6, 15))
# plot only the top 20 results
sns.barplot(x="or", y="name", data=results.head(20), color="b")
# adding axis title
ax.set(ylabel="Gene sets",
xlabel="Odds Ratio")
# make the plot look cleaner and nicer
sns.despine(left=True, bottom=True)