API reference¶
This page is generated from the docstrings in the package, so it always describes the code it was built from (version 1.3.0).
Choose a server¶
set_server ¶
Choose which DISCO server the toolkit uses for the rest of the session.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
server
|
str
|
"v1" (DISCO v1, the default), or the full URL of the API root of a mirror or test copy of DISCO v1, such as "http://127.0.0.1:8889/disco_v3_api/". |
required |
Returns:
| Name | Type | Description |
|---|---|---|
str |
str
|
the API root now in use. |
Find and download data¶
Filter ¶
Which DISCO samples and cells to look for.
Every argument is optional; a filter with none of them matches everything. Give a string or a
list of strings. The list of values each field accepts is list_metadata_item(field).
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
sample_id
|
str or list
|
sample identifier, e.g. "ERX2757110". |
None
|
project_id
|
str or list
|
project (a study or dataset), e.g. "GSE147520". |
None
|
tissue
|
str or list
|
e.g. "lung", "bladder". |
None
|
disease
|
str or list
|
e.g. "COVID-19". |
None
|
platform
|
str or list
|
sequencing platform, e.g. "10x3'". |
None
|
sample_type
|
str or list
|
e.g. "control". |
None
|
cell_type
|
str or list
|
keep samples that contain these cell types, and download only those cells. |
None
|
cell_type_confidence
|
str
|
how sure DISCO's annotation must be: "high", "medium" or "all". Defaults to "medium". |
'medium'
|
include_cell_type_children
|
bool
|
also match the more specific cell types under
|
True
|
min_cell_per_sample
|
int
|
drop samples with fewer matching cells than this. Defaults to 100. |
100
|
FilterData ¶
The result of a filter: the matching samples, with their counts.
Returned by filter_disco_metadata and passed to download_disco_data.
Attributes:
| Name | Type | Description |
|---|---|---|
sample_metadata |
DataFrame
|
one row per matching sample. |
cell_type_metadata |
DataFrame
|
the cell types found in each sample. |
sample_count |
int
|
number of matching samples. |
cell_count |
int
|
number of matching cells. |
filter |
Filter
|
the filter that produced this result. |
filter_disco_metadata ¶
Find the DISCO samples and cells that match a filter.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
filter
|
Filter
|
what to look for. With no argument, everything matches. |
Filter()
|
Returns:
| Name | Type | Description |
|---|---|---|
FilterData |
FilterData
|
the matching samples and counts, to pass to |
list_metadata_item ¶
List element inside the metadata columns
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
field
|
str
|
metadata columns or field from the disco database |
required |
Returns:
| Name | Type | Description |
|---|---|---|
list |
list
|
the unique elements of the metadata column, as a reference for filtering |
list_all_columns ¶
list all the columns found in the metadata of the disco database
Returns:
| Name | Type | Description |
|---|---|---|
list |
list
|
the names of the metadata columns, as a list of strings |
find_celltype ¶
find the celltype within the disco dataset
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
term
|
String
|
term refer to string of the cell type |
''
|
cell_ontology
|
dict
|
cell_ontology can be provided by the user in the format of dictionary datatype in python. Defaults to None. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
list |
list
|
the cell types that match the term |
get_celltype_children ¶
get the children of the input celltype from the user
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
cell_type
|
Union[str, list]
|
the input can be either string or list of string |
required |
cell_ontology
|
dict
|
cell_ontology can be provided by the user in the format of dictionary datatype in python. Defaults to None. |
None
|
Returns:
| Name | Type | Description |
|---|---|---|
list |
list
|
the children of the defined cell type, as a list of strings |
download_disco_data ¶
Download the samples of a filter result, one AnnData .h5ad file per sample.
Each file holds only the cells that matched the filter, with DISCO's annotation in
obs["cell_type"].
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
metadata
|
FilterData
|
the result of |
required |
output_dir
|
str
|
directory to save the files in; created if missing. Defaults to "DISCOtmp". |
'DISCOtmp'
|
Returns:
| Name | Type | Description |
|---|---|---|
None |
None
|
the files are written to |
Cell type annotation (CELLiD)¶
get_atlas ¶
get the all the atlas string from the DISCO website and return to the user
Returns:
| Name | Type | Description |
|---|---|---|
list |
list
|
the atlas names, as a list of strings |
CELLiD_cluster ¶
CELLiD_cluster(
rna,
ref_data: DataFrame = None,
ref_deg: DataFrame = None,
atlas: str = None,
n_predict: int = 1,
ref_path: str = None,
ncores: int = 10,
) -> pd.DataFrame
Cell type annotation using reference data and compute the correlation between the user cell gene expression as compare to the reference data. The celltype with highest correlation will be concluded as the celltype
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
rna
|
Pandas DataFrame | Numpy array
|
user define dataframe. Need to transpose so that the index is the genes |
required |
ref_data
|
Pandas DataFrame
|
Reference dataframe used to compute for the cell type annotation. Defaults to None. |
None
|
ref_deg
|
Pandas DataFrame
|
reference DEG database. Defaults to None. |
None
|
atlas
|
String
|
String of atlas that the user want to use as the reference. Defaults to None. |
None
|
n_predict
|
Integer
|
number of predicted celltype. Defaults to 1. |
1
|
ref_path
|
string
|
path string to the reference data. Defaults to None. |
None
|
ncores
|
Integer
|
number of CPU cores used to run the data. Defaults to 10. |
10
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Pandas DataFrame: return the Pandas DataFrame along with the correlation score. |
Enrichment (scEnrichment)¶
CELLiD_enrichment ¶
CELLiD_enrichment(
input: DataFrame,
reference: DataFrame = None,
ref_path: str = None,
ncores: int = 1,
) -> pd.DataFrame
Function to generate enrichment analysis based on the reference gene sets and following the DISCO pipeline.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
input
|
Pandas DataFrame
|
User defined Dataframe in the format of |
required |
reference
|
Pandas DataFrame
|
Reference datasets from DISCO. Recommend to put as None as the function will automatically retrieve the dataset from the server. Defaults to None. |
None
|
ref_path
|
String
|
Path to the reference dataset or reading the file if it is existed. Defaults to None. |
None
|
ncores
|
Integer
|
Number of CPU cores to run the function. Defaults to 1. |
1
|
Returns:
| Type | Description |
|---|---|
DataFrame
|
Pandas DataFrame: return the significant gene sets that is over-represented in a large set of genes. |
Gene search¶
gene_search ¶
gene_search(
gene: str,
atlas: Union[str, list] = None,
figsize: tuple = None,
dpi: int = 300,
) -> None
Function to search for the gene expression level the same as the input gene search bar in DISCO website.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
gene
|
String
|
name of the gene in capital letter. e.g. LYVE1. |
required |
atlas
|
String or List of String, Optional
|
User defined atlas for visualisation. Default to None to search for all Atlases. |
None
|
figsize
|
(tuple, Optional)
|
Size of the generated figure in tuple. Default to None. |
None
|
dpi
|
(int, Optional)
|
DPI resolution for the figure. Default to 300. |
300
|
Returns:
| Name | Type | Description |
|---|---|---|
None |
None
|
This function does not return anything beside plotting |