Research Protocol for Single-cell Omics Analysis
Materials Required
Background
Single-cell omics analysis measures molecular features in individual cells to resolve cellular heterogeneity, rare populations, cell states, developmental trajectories, and disease-specific cell programs that are obscured in bulk assays[1][2].
Single-cell RNA-seq is the most established modality and typically requires cell-level quality control, normalization, feature selection, dimensionality reduction, clustering, marker-gene identification, cell-type annotation, and differential analysis[3][4].
Single-cell ATAC-seq measures chromatin accessibility at single-cell resolution and requires modality-specific QC, peak or bin quantification, dimensionality reduction, motif analysis, and integration with transcriptomic data when regulatory interpretation is needed[5][6].
Single-cell multi-omics can jointly or computationally integrate transcriptomic, epigenomic, protein, spatial, or lineage information, but unresolved problems include batch effects, sparse data, doublets, cell-type annotation uncertainty, donor-level replication, and distinguishing correlation from mechanism[7][8][9].
MCE has not independently verified the accuracy of these methods. They are for reference only.
Project Analysis
• Generate or import raw sequencing data, demultiplex samples when needed, quantify reads or fragments into cell-by-feature matrices, attach metadata, and preserve raw files and software versions for reproducibility[3][10].
• Perform sample-level and cell-level QC, remove low-quality cells, likely doublets, empty droplets, and outlier libraries, then normalize each modality with methods appropriate to the data type[3][4][5][10].
• Conduct dimensionality reduction, clustering, marker detection, reference-assisted annotation, and manual biological review, then verify that clusters are represented across biological replicates rather than driven by batch or sample identity[4][11].
• Analyze condition effects using sample-aware differential expression, differential detection, differential accessibility, or differential abundance methods within annotated cell types or states[12][13][14].
• Integrate multi-omics or spatial data only after modality-specific QC, then validate major discoveries with independent datasets, protein-level assays, spatial localization, or functional perturbation[7][8][9][15].
Phased Objectives
Objective 1: Establish the single-cell dataset and quality-control framework.
• Research approach: generate or import single-cell omics data and apply predefined sample-level and cell-level QC.• Experimental model: fresh tissue, frozen nuclei, cultured cells, organoids, animal tissue, or clinical specimens.
• Experimental groups: control, disease or treatment, biological replicates, technical replicates when available, and excluded low-quality cells.
• Key techniques: single-cell RNA-seq, single-nucleus RNA-seq, scATAC-seq, CITE-seq, multiome sequencing, quality-control visualization, and metadata curation.
• Detection indices: number of cells, detected genes or fragments per cell, mitochondrial transcript fraction, library complexity, doublet score, fragment enrichment, transcription start-site enrichment for scATAC-seq, and batch structure.
• Expected results: high-quality cells or nuclei should separate from low-quality droplets, debris, and doublets.
• Interpretation: downstream analysis should use only QC-passed cells with documented exclusion rules[3][4][5][10].
Objective 2: Identify cell populations and biological states.
• Research approach: normalize data, reduce dimensionality, cluster cells, and annotate cell identities using markers and references.• Experimental model: QC-passed single-cell matrix.
• Experimental groups: unsupervised clusters, reference-labeled populations, manually curated cell types, and uncertain clusters.
• Key techniques: normalization, highly variable feature selection, PCA or latent semantic indexing, UMAP/t-SNE visualization, graph clustering, marker-gene analysis, reference mapping, and cell-type annotation.
• Detection indices: cluster stability, marker specificity, known lineage markers, annotation confidence, batch mixing, and replicate representation.
• Expected results: major expected cell types and biologically meaningful subclusters should be detected.
• Interpretation: clusters should not be interpreted as new cell types unless supported by markers, replication, and biological context[3][4][11].
Objective 3: Detect condition-associated cellular and molecular changes.
• Research approach: compare cell proportions, gene expression, chromatin accessibility, protein abundance, or pathway scores between biological conditions.• Experimental model: annotated single-cell dataset with biological replicates.
• Experimental groups: control versus disease or treatment within each cell type.
• Key techniques: pseudobulk differential expression, differential abundance testing, differential accessibility testing, pathway enrichment, gene-set scoring, and sample-aware statistical modeling.
• Detection indices: log fold change, adjusted P value, donor-level consistency, cell-type-specific effect size, differential cell abundance, pathway enrichment, and chromatin motif activity.
• Expected results: disease or treatment should alter specific cell states, cell proportions, or molecular programs.
• Interpretation: donor-aware or sample-aware analysis is required to avoid mistaking many cells from one donor for true biological replication[12][13][14].
Objective 4: Integrate modalities and validate prioritized findings.
• Research approach: integrate single-cell transcriptomic, epigenomic, protein, spatial, or lineage information and validate key findings with independent assays.• Experimental model: matched or comparable single-cell multi-omics datasets, spatial tissue sections, sorted cells, or validation cohort.
• Experimental groups: discovery dataset, integrated dataset, independent validation dataset, and experimental validation group.
• Key techniques: scRNA-seq/scATAC-seq integration, CITE-seq integration, spatial mapping, motif-to-gene linking, trajectory inference, RNA velocity when justified, immunofluorescence, flow cytometry, RT-qPCR, Western blot, or functional perturbation.
• Detection indices: cross-modality concordance, marker replication, spatial localization, trajectory consistency, protein validation, and functional readout.
• Expected results: high-priority cell states or markers should replicate across modalities or independent methods.
• Interpretation: integrated single-cell findings support candidate mechanisms but require orthogonal validation before causal claims[7][8][9][15].
Critical Points
Objective 1
• Produce a clean single-cell object with reproducible QC decisions, retained cells, excluded cells, and interpretable sample-level metrics; failure at this step suggests poor sample quality or inappropriate preprocessing[3][10].Objective 2
• Identify known cell classes and biologically interpretable subpopulations; clusters lacking markers, replicate support, or annotation confidence should be treated as uncertain[4][11].Objective 3
• Identify cell-type-specific molecular or compositional changes associated with the phenotype; results that disappear under donor-aware analysis should not be considered robust[12][13][14].Objective 4
• Prioritize mechanisms supported by transcriptomic, epigenomic, protein, spatial, or functional evidence; lack of validation suggests that the signal may reflect batch effects, annotation error, or dataset-specific noise[7][8][9][15].Troubleshooting
1: the dissociation or nuclei isolation can bias cell recovery and alter apparent cell proportions.
Alternative: compare cell recovery with expected the biology and validate key populations by flow cytometry, immunohistochemistry, or spatial methods[3][15].2: doublets can create artificial hybrid clusters.
Alternative: use computational doublet detection, inspect co-expression of incompatible lineage markers, and remove suspected doublets before annotation[10][16].3: batch correction can remove true biology or fail to remove technical effects.
Alternative: inspect batch structure before and after integration and preserve biological contrasts during correction[4][7][8].4: cell-level differential testing can inflate significance when biological replication is ignored.
Alternative: use pseudobulk or mixed-model approaches that respect sample or donor structure[12][13][14].5: marker-based annotation can be biased or ambiguous.
Alternative: combine canonical markers, reference mapping, cluster-specific marker analysis, and orthogonal validation for key cell types[11][15].References:
- [1]. Tang F, et al. mRNA-Seq whole-transcriptome analysis of a single cell. Nat Methods. 2009;6(5):377-382. [Content Brief]
- [2]. Kolodziejczyk AA, et al. The technology and biology of single-cell RNA sequencing. Mol Cell. 2015;58(4):610-620. [Content Brief]
- [3]. Luecken MD, et al. Current best practices in single-cell RNA-seq analysis: a tutorial. Mol Syst Biol. 2019;15(6):e8746. [Content Brief]
- [4]. Stuart T, et al. Comprehensive integration of single-cell data. Cell. 2019;177(7):1888-1902.e21. [Content Brief]
- [5]. Buenrostro JD, et al. Single-cell chromatin accessibility reveals principles of regulatory variation. Nature. 2015;523(7561):486-490. [Content Brief]
- [6]. Chen H, et al. Assessment of computational methods for the analysis of single-cell ATAC-seq data. Genome Biol. 2019;20(1):241. [Content Brief]
- [7]. Colomé-Tatché M, et al. Statistical single cell multi-omics integration. Curr Opin Syst Biol. 2018;7:54-59.
- [8]. Argelaguet R, et al. MOFA+: a statistical framework for comprehensive integration of multi-modal single-cell data. Genome Biol. 2020;21(1):111. [Content Brief]
- [9]. Xiao C, et al. Benchmarking multi-omics integration algorithms across single-cell RNA and ATAC data. Brief Bioinform. 2024;25(3):bbae095. [Content Brief]
- [10]. McGinnis CS, et al. DoubletFinder: doublet detection in single-cell RNA sequencing data using artificial nearest neighbors. Cell Syst. 2019;8(4):329-337.e4. [Content Brief]
- [11]. Diaz A, et al. SCell: integrated analysis of single-cell RNA-seq data. Bioinformatics. 2016;32(14):2219-2220. [Content Brief]
- [12]. Crowell HL, et al. Muscat detects subpopulation-specific state transitions from multi-sample multi-condition single-cell transcriptomics data. Nat Commun. 2020;11(1):6077. [Content Brief]
- [13]. Zimmerman KD, et al. A practical solution to pseudoreplication bias in single-cell studies. Nat Commun. 2021;12(1):738. [Content Brief]
- [14]. Gilis J, et al. Differential detection workflows for multi-sample single-cell RNA-seq data. BMC Genomics. 2025;26(1):420. [Content Brief]
- [15]. Ståhl PL, et al. Visualization and analysis of gene expression in tissue sections by spatial transcriptomics. Science. 2016;353(6294):78-82. [Content Brief]
- [16]. DePasquale EAK, et al. DoubletDecon: deconvoluting doublets from single-cell RNA-sequencing data. Cell Rep. 2019;29(6):1718-1727.e8. [Content Brief]