DISCOtoolkit¶
DISCOtoolkit is the Python package for the DISCO database: filter and download single-cell data, annotate your own data with CELLiD, and test gene sets with scEnrichment. This documentation describes version 1.3.0, and is rebuilt, with every tutorial re-run against the live database, whenever the package changes.
Install¶
Python 3.9 or newer.
The dependencies (scanpy, pandas, numpy and others) are installed with it. For the development
version: pip install "git+https://github.com/JinmiaoChenLab/DISCOtoolkit_py.git".
Start here¶
| Page | What it shows |
|---|---|
| Quickstart | From a query to a UMAP: filter, download, cluster, plot. |
| Download data | Filters, cell type confidence and downloading in depth. |
| Cell type annotation | Annotate your own clusters against the DISCO reference. |
| Enrichment | Test a gene list against DISCO differential-expression gene sets. |
| Gene search | A gene's expression across every annotated cell type. |
| API reference | Every function and its arguments. |
Each tutorial can be opened in Google Colab, and the notebook downloaded, from its page.
Which server¶
The toolkit is built for and tested against DISCO V1, which it uses by default. DISCO V2 is not supported; it has its own R package, DISCOtoolkit.
import discotoolkit as dt
dt.get_server() # DISCO V1
dt.set_server("https://my.mirror/disco_v3_api/") # a mirror or test copy of DISCO V1
Citation¶
If you use DISCO in your work, please cite: Li M. et al., DISCO: a database of Deeply Integrated human Single-Cell Omics data, Nucleic Acids Research 2022; and Li M. et al., Rediscovering publicly available single-cell data with the DISCO platform, Nucleic Acids Research 2025.