Identify PBMC populations with scRNA-seq#
A cluster number tells us which cells have similar RNA profiles. To give a cluster a biological name, we need evidence from its genes. Here we will inspect a prepared blood-cell analysis, compare a small marker panel, and assign broad cell types.
Start with the quick start if you want to run the standard pipeline on your own counts. This page uses saved results so we can concentrate on interpreting them.
Open the analysis and look at the clusters#
# Open count stores and run Scarf analyses.
import scarf
# Keep routine execution messages out of the results.
scarf.configure_output(level="WARNING", progress=False)
# Download the prepared example store.
dataset = scarf.cytebase.connect("scarf_docs").download_dataset(
"tenx_5K_pbmc_rnaseq",
destination="scarf_datasets",
zarr=True,
)
# Open the count store for this analysis.
ds = scarf.DataStore(f"{dataset}/data.zarr")
# Open the saved analysis and its exact results.
run = ds.pipeline.open(label="docs_default")
# Inspect the opened store's cells and features.
ds
DataStore has 3948 (5025) cells with 1 assays: RNA
Cell metadata:
'I', 'ids', 'names', 'RNA_G2M_score', 'RNA_S_score',
'RNA_UMAP1', 'RNA_UMAP2', 'RNA_cell_cycle_phase', 'RNA_clusters', 'RNA_doublet_score',
'RNA_leiden_cluster', 'RNA_nCounts', 'RNA_nFeatures', 'RNA_paris_cluster', 'RNA_percentMito',
'RNA_percentRibo'
RNA assay has 33538 features and following metadata:
'I', 'ids', 'names', 'dropOuts', 'feature_type',
'genome', 'nCells'
The prepared run contains the selected cells, clusters, and UMAP coordinates. An earlier release
wrote its marker tables, which get_markers refuses; rank markers for its clusters with
ds.run_marker_search(run["clusters"], features=run["feature_universe"]), as Annotate cell types and states
does. Its name is docs_default, but it used dataset-specific filtering, 500 variable genes, and
15 PCs. Those settings preserve the example we will interpret; they are not the current pipeline
defaults.
# Locate the numbered clusters in the prepared PBMC UMAP.
ds.plots.embedding(run=run, color_by="clusters")
Look for the main groups and the smaller populations. Nearby cells have similar profiles in this view, but the size of a gap between clusters does not establish how different their cell types are. We will use marker expression to investigate that.
Compare markers before naming the groups#
The panel below includes several genes for each broad lineage. In a dot plot, a larger dot means more cells express the gene; its colour shows the mean expression in the cluster.
# Choose several markers for each broad lineage.
marker_panel = {
"Monocyte": ["LST1", "S100A8", "FCGR3A"],
"B cell": ["MS4A1", "CD79A"],
"T cell": ["CD3D", "IL7R"],
"NK cell": ["NKG7", "GNLY"],
"pDC-like": ["GZMB", "JCHAIN"],
}
# Compare marker expression and detection across the selected groups.
ds.plots.dotplot(features=marker_panel, groups=run["clusters"])
MS4A1 and CD79A support B-cell identities; CD3D and IL7R support T cells. NKG7 and GNLY help identify cytotoxic populations, so compare them with the T-cell markers before calling a group NK cells. LST1, S100A8, and FCGR3A help distinguish the monocyte populations.
GZMB together with JCHAIN suggests a small pDC-like population. This is a provisional label: the Annotate cell types and states tutorial checks IL3RA and LILRA4 and examines competing interpretations.
For a gene with much lower expression than the others, it can help to scale each gene separately. The next plot changes only the colour scale: values are now relative within each gene, so colours cannot be used to compare absolute expression between different genes.
# Compare marker expression and detection across the selected groups.
ds.plots.dotplot(features=marker_panel, groups=run['clusters'], standardize='feature')
Give the clusters broad names#
The marker evidence supports the following broad labels for this prepared result. Several clusters share a label because they belong to the same lineage. These cluster numbers are specific to this analysis and must not be copied to another dataset.
# Assign broad names supported by the marker evidence.
cell_type_by_cluster = {
"1": "CD14 monocytes",
"2": "monocytes",
"3": "B cells",
"4": "T cells",
"5": "T cells",
"6": "NK cells",
"7": "T cells",
"8": "T cells",
"9": "B cells",
"10": "pDC-like cells",
}
# Review the names assigned to the prepared clusters.
cell_type_by_cluster
{'1': 'CD14 monocytes',
'2': 'monocytes',
'3': 'B cells',
'4': 'T cells',
'5': 'T cells',
'6': 'NK cells',
'7': 'T cells',
'8': 'T cells',
'9': 'B cells',
'10': 'pDC-like cells'}
Save the names in the cell table. The run contains only the analyzed cells, while the cell table contains every cell, so fill the other rows with an explicit label before inserting the column.
# Work with numeric arrays and cell masks.
import numpy as np
# Locate the analyzed cells within the full cell table.
analysis_cells = run.cells.fetch_all("I").astype(bool)
# Read cluster labels in the selected cells' order.
cluster_values = run.cells.fetch("clusters").astype(str)
# Give cells outside the analysis an explicit label.
cell_types = np.full(ds.cells.N, "Not analyzed", dtype=object)
# Fill analyzed rows with their cluster's chosen cell-type label.
cell_types[analysis_cells] = [cell_type_by_cluster[value] for value in cluster_values]
# Save the calculated values in the cell table.
ds.cells.insert("pbmc_cell_type", cell_types, overwrite=True)
# Count analyzed cells assigned to each cell type.
labels, counts = np.unique(cell_types[analysis_cells], return_counts=True)
# Show the number of analyzed cells assigned to each cell type.
{str(label): int(count) for label, count in zip(labels, counts, strict=True)}
{'B cells': 495,
'CD14 monocytes': 716,
'NK cells': 319,
'T cells': 2378,
'monocytes': 25,
'pDC-like cells': 15}
This writes our labels to pbmc_cell_type; rerunning the cell replaces that column. The saved
clusters remain available. Use their UMAP to display the new names:
# Show the assigned broad cell types on the saved UMAP.
ds.plots.embedding(layout=run["umap"], color_by="pbmc_cell_type")
Build on this analysis#
Continue with Annotate cell types and states to distinguish finer cell types and states using supporting and negative markers. Before interpreting your own data, review Quality control across assays and check that the retained cells make sense for your study.
When the first result raises a specific question, use Choosing informative features, Choosing dimensionality reductions, or Clustering and cluster evidence to compare one choice at a time. Building neighbourhood graphs step by step explains the individual computation steps, and Provenance and reuse shows how to reopen and compare saved analyses.
Marker tests here compare cells, not biological replicates. For differences between conditions, use a suitable study design and the Pseudobulk and differential expression (DE) primer or Compare biological conditions across donors guide.