Skip to content

API reference

This page is generated from the docstrings in the package, so it always describes the code it was built from (version 1.3.0).

import discotoolkit as dt

Choose a server

set_server

set_server(server: str) -> str

Choose which DISCO server the toolkit uses for the rest of the session.

Parameters:

Name Type Description Default
server str

"v1" (DISCO v1, the default), or the full URL of the API root of a mirror or test copy of DISCO v1, such as "http://127.0.0.1:8889/disco_v3_api/".

required

Returns:

Name Type Description
str str

the API root now in use.

get_server

get_server() -> str

The API root the toolkit is currently using.

Find and download data

Filter

Which DISCO samples and cells to look for.

Every argument is optional; a filter with none of them matches everything. Give a string or a list of strings. The list of values each field accepts is list_metadata_item(field).

Parameters:

Name Type Description Default
sample_id str or list

sample identifier, e.g. "ERX2757110".

None
project_id str or list

project (a study or dataset), e.g. "GSE147520".

None
tissue str or list

e.g. "lung", "bladder".

None
disease str or list

e.g. "COVID-19".

None
platform str or list

sequencing platform, e.g. "10x3'".

None
sample_type str or list

e.g. "control".

None
cell_type str or list

keep samples that contain these cell types, and download only those cells.

None
cell_type_confidence str

how sure DISCO's annotation must be: "high", "medium" or "all". Defaults to "medium".

'medium'
include_cell_type_children bool

also match the more specific cell types under cell_type. Defaults to True.

True
min_cell_per_sample int

drop samples with fewer matching cells than this. Defaults to 100.

100

FilterData

The result of a filter: the matching samples, with their counts.

Returned by filter_disco_metadata and passed to download_disco_data.

Attributes:

Name Type Description
sample_metadata DataFrame

one row per matching sample.

cell_type_metadata DataFrame

the cell types found in each sample.

sample_count int

number of matching samples.

cell_count int

number of matching cells.

filter Filter

the filter that produced this result.

filter_disco_metadata

filter_disco_metadata(
    filter: Filter = Filter(),
) -> FilterData

Find the DISCO samples and cells that match a filter.

Parameters:

Name Type Description Default
filter Filter

what to look for. With no argument, everything matches.

Filter()

Returns:

Name Type Description
FilterData FilterData

the matching samples and counts, to pass to download_disco_data.

list_metadata_item

list_metadata_item(field: str) -> list

List element inside the metadata columns

Parameters:

Name Type Description Default
field str

metadata columns or field from the disco database

required

Returns:

Name Type Description
list list

the unique elements of the metadata column, as a reference for filtering

list_all_columns

list_all_columns() -> list

list all the columns found in the metadata of the disco database

Returns:

Name Type Description
list list

the names of the metadata columns, as a list of strings

find_celltype

find_celltype(
    term: str = "", cell_ontology: dict = None
) -> list

find the celltype within the disco dataset

Parameters:

Name Type Description Default
term String

term refer to string of the cell type

''
cell_ontology dict

cell_ontology can be provided by the user in the format of dictionary datatype in python. Defaults to None.

None

Returns:

Name Type Description
list list

the cell types that match the term

get_celltype_children

get_celltype_children(
    cell_type: Union[str, list], cell_ontology: dict = None
) -> list

get the children of the input celltype from the user

Parameters:

Name Type Description Default
cell_type Union[str, list]

the input can be either string or list of string

required
cell_ontology dict

cell_ontology can be provided by the user in the format of dictionary datatype in python. Defaults to None.

None

Returns:

Name Type Description
list list

the children of the defined cell type, as a list of strings

download_disco_data

download_disco_data(
    metadata, output_dir: str = "DISCOtmp"
) -> None

Download the samples of a filter result, one AnnData .h5ad file per sample.

Each file holds only the cells that matched the filter, with DISCO's annotation in obs["cell_type"].

Parameters:

Name Type Description Default
metadata FilterData

the result of filter_disco_metadata.

required
output_dir str

directory to save the files in; created if missing. Defaults to "DISCOtmp".

'DISCOtmp'

Returns:

Name Type Description
None None

the files are written to output_dir.

Cell type annotation (CELLiD)

get_atlas

get_atlas(ref_data=None, ref_path=None) -> list

get the all the atlas string from the DISCO website and return to the user

Returns:

Name Type Description
list list

the atlas names, as a list of strings

CELLiD_cluster

CELLiD_cluster(
    rna,
    ref_data: DataFrame = None,
    ref_deg: DataFrame = None,
    atlas: str = None,
    n_predict: int = 1,
    ref_path: str = None,
    ncores: int = 10,
) -> pd.DataFrame

Cell type annotation using reference data and compute the correlation between the user cell gene expression as compare to the reference data. The celltype with highest correlation will be concluded as the celltype

Parameters:

Name Type Description Default
rna Pandas DataFrame | Numpy array

user define dataframe. Need to transpose so that the index is the genes

required
ref_data Pandas DataFrame

Reference dataframe used to compute for the cell type annotation. Defaults to None.

None
ref_deg Pandas DataFrame

reference DEG database. Defaults to None.

None
atlas String

String of atlas that the user want to use as the reference. Defaults to None.

None
n_predict Integer

number of predicted celltype. Defaults to 1.

1
ref_path string

path string to the reference data. Defaults to None.

None
ncores Integer

number of CPU cores used to run the data. Defaults to 10.

10

Returns:

Type Description
DataFrame

Pandas DataFrame: return the Pandas DataFrame along with the correlation score.

Enrichment (scEnrichment)

CELLiD_enrichment

CELLiD_enrichment(
    input: DataFrame,
    reference: DataFrame = None,
    ref_path: str = None,
    ncores: int = 1,
) -> pd.DataFrame

Function to generate enrichment analysis based on the reference gene sets and following the DISCO pipeline.

Parameters:

Name Type Description Default
input Pandas DataFrame

User defined Dataframe in the format of (gene, fc). gene refer to gene name and fc refer to log fold change.

required
reference Pandas DataFrame

Reference datasets from DISCO. Recommend to put as None as the function will automatically retrieve the dataset from the server. Defaults to None.

None
ref_path String

Path to the reference dataset or reading the file if it is existed. Defaults to None.

None
ncores Integer

Number of CPU cores to run the function. Defaults to 1.

1

Returns:

Type Description
DataFrame

Pandas DataFrame: return the significant gene sets that is over-represented in a large set of genes.

gene_search(
    gene: str,
    atlas: Union[str, list] = None,
    figsize: tuple = None,
    dpi: int = 300,
) -> None

Function to search for the gene expression level the same as the input gene search bar in DISCO website.

Parameters:

Name Type Description Default
gene String

name of the gene in capital letter. e.g. LYVE1.

required
atlas String or List of String, Optional

User defined atlas for visualisation. Default to None to search for all Atlases.

None
figsize (tuple, Optional)

Size of the generated figure in tuple. Default to None.

None
dpi (int, Optional)

DPI resolution for the figure. Default to 300.

300

Returns:

Name Type Description
None None

This function does not return anything beside plotting