Research Protocol for Single-cell Omics Analysis

Materials Required

/

Background

Single-cell omics analysis measures molecular features in individual cells to resolve cellular heterogeneity, rare populations, cell states, developmental trajectories, and disease-specific cell programs that are obscured in bulk assays[1][2].

Single-cell RNA-seq is the most established modality and typically requires cell-level quality control, normalization, feature selection, dimensionality reduction, clustering, marker-gene identification, cell-type annotation, and differential analysis[3][4].

Single-cell ATAC-seq measures chromatin accessibility at single-cell resolution and requires modality-specific QC, peak or bin quantification, dimensionality reduction, motif analysis, and integration with transcriptomic data when regulatory interpretation is needed[5][6].

Single-cell multi-omics can jointly or computationally integrate transcriptomic, epigenomic, protein, spatial, or lineage information, but unresolved problems include batch effects, sparse data, doublets, cell-type annotation uncertainty, donor-level replication, and distinguishing correlation from mechanism[7][8][9].

MCE has not independently verified the accuracy of these methods. They are for reference only.

Project Analysis

• Begin by defining the biological question, specimen source, experimental groups, replicate structure, dissociation or nuclei-isolation strategy, sequencing modality, and primary endpoints before data generation or reanalysis[3][4].

• Generate or import raw sequencing data, demultiplex samples when needed, quantify reads or fragments into cell-by-feature matrices, attach metadata, and preserve raw files and software versions for reproducibility[3][10].

• Perform sample-level and cell-level QC, remove low-quality cells, likely doublets, empty droplets, and outlier libraries, then normalize each modality with methods appropriate to the data type[3][4][5][10].

• Conduct dimensionality reduction, clustering, marker detection, reference-assisted annotation, and manual biological review, then verify that clusters are represented across biological replicates rather than driven by batch or sample identity[4][11].

• Analyze condition effects using sample-aware differential expression, differential detection, differential accessibility, or differential abundance methods within annotated cell types or states[12][13][14].

• Integrate multi-omics or spatial data only after modality-specific QC, then validate major discoveries with independent datasets, protein-level assays, spatial localization, or functional perturbation[7][8][9][15].

Phased Objectives

Objective 1: Establish the single-cell dataset and quality-control framework.

• Research approach: generate or import single-cell omics data and apply predefined sample-level and cell-level QC.
• Experimental model: fresh tissue, frozen nuclei, cultured cells, organoids, animal tissue, or clinical specimens.
• Experimental groups: control, disease or treatment, biological replicates, technical replicates when available, and excluded low-quality cells.
• Key techniques: single-cell RNA-seq, single-nucleus RNA-seq, scATAC-seq, CITE-seq, multiome sequencing, quality-control visualization, and metadata curation.
• Detection indices: number of cells, detected genes or fragments per cell, mitochondrial transcript fraction, library complexity, doublet score, fragment enrichment, transcription start-site enrichment for scATAC-seq, and batch structure.
• Expected results: high-quality cells or nuclei should separate from low-quality droplets, debris, and doublets.
• Interpretation: downstream analysis should use only QC-passed cells with documented exclusion rules[3][4][5][10].

Objective 2: Identify cell populations and biological states.

• Research approach: normalize data, reduce dimensionality, cluster cells, and annotate cell identities using markers and references.
• Experimental model: QC-passed single-cell matrix.
• Experimental groups: unsupervised clusters, reference-labeled populations, manually curated cell types, and uncertain clusters.
• Key techniques: normalization, highly variable feature selection, PCA or latent semantic indexing, UMAP/t-SNE visualization, graph clustering, marker-gene analysis, reference mapping, and cell-type annotation.
• Detection indices: cluster stability, marker specificity, known lineage markers, annotation confidence, batch mixing, and replicate representation.
• Expected results: major expected cell types and biologically meaningful subclusters should be detected.
• Interpretation: clusters should not be interpreted as new cell types unless supported by markers, replication, and biological context[3][4][11].

Objective 3: Detect condition-associated cellular and molecular changes.

• Research approach: compare cell proportions, gene expression, chromatin accessibility, protein abundance, or pathway scores between biological conditions.
• Experimental model: annotated single-cell dataset with biological replicates.
• Experimental groups: control versus disease or treatment within each cell type.
• Key techniques: pseudobulk differential expression, differential abundance testing, differential accessibility testing, pathway enrichment, gene-set scoring, and sample-aware statistical modeling.
• Detection indices: log fold change, adjusted P value, donor-level consistency, cell-type-specific effect size, differential cell abundance, pathway enrichment, and chromatin motif activity.
• Expected results: disease or treatment should alter specific cell states, cell proportions, or molecular programs.
• Interpretation: donor-aware or sample-aware analysis is required to avoid mistaking many cells from one donor for true biological replication[12][13][14].

Objective 4: Integrate modalities and validate prioritized findings.

• Research approach: integrate single-cell transcriptomic, epigenomic, protein, spatial, or lineage information and validate key findings with independent assays.
• Experimental model: matched or comparable single-cell multi-omics datasets, spatial tissue sections, sorted cells, or validation cohort.
• Experimental groups: discovery dataset, integrated dataset, independent validation dataset, and experimental validation group.
• Key techniques: scRNA-seq/scATAC-seq integration, CITE-seq integration, spatial mapping, motif-to-gene linking, trajectory inference, RNA velocity when justified, immunofluorescence, flow cytometry, RT-qPCR, Western blot, or functional perturbation.
• Detection indices: cross-modality concordance, marker replication, spatial localization, trajectory consistency, protein validation, and functional readout.
• Expected results: high-priority cell states or markers should replicate across modalities or independent methods.
• Interpretation: integrated single-cell findings support candidate mechanisms but require orthogonal validation before causal claims[7][8][9][15].

Critical Points

Objective 1

• Produce a clean single-cell object with reproducible QC decisions, retained cells, excluded cells, and interpretable sample-level metrics; failure at this step suggests poor sample quality or inappropriate preprocessing[3][10].

Objective 2

• Identify known cell classes and biologically interpretable subpopulations; clusters lacking markers, replicate support, or annotation confidence should be treated as uncertain[4][11].

Objective 3

• Identify cell-type-specific molecular or compositional changes associated with the phenotype; results that disappear under donor-aware analysis should not be considered robust[12][13][14].

Objective 4

• Prioritize mechanisms supported by transcriptomic, epigenomic, protein, spatial, or functional evidence; lack of validation suggests that the signal may reflect batch effects, annotation error, or dataset-specific noise[7][8][9][15].

Troubleshooting

1: the dissociation or nuclei isolation can bias cell recovery and alter apparent cell proportions.

Alternative: compare cell recovery with expected the biology and validate key populations by flow cytometry, immunohistochemistry, or spatial methods[3][15].

2: doublets can create artificial hybrid clusters.

Alternative: use computational doublet detection, inspect co-expression of incompatible lineage markers, and remove suspected doublets before annotation[10][16].

3: batch correction can remove true biology or fail to remove technical effects.

Alternative: inspect batch structure before and after integration and preserve biological contrasts during correction[4][7][8].

4: cell-level differential testing can inflate significance when biological replication is ignored.

Alternative: use pseudobulk or mixed-model approaches that respect sample or donor structure[12][13][14].

5: marker-based annotation can be biased or ambiguous.

Alternative: combine canonical markers, reference mapping, cluster-specific marker analysis, and orthogonal validation for key cell types[11][15].

参考文献: