UTMB logo LiverHomo

Methods

Acquisition and Preprocessing of scRNAseq Data

Human liver scRNA-seq datasets were systematically curated from public repositories, including: (i) the NCBI Gene Expression Omnibus (GEO); (ii) the Sequence Read Archive (SRA); and (iii) the Tabula Sapiens consortium. Studies were selected based on relevance to human liver tissues across healthy and disease states, with inclusion criteria requiring a 10x Genomics-compatible platform and metadata including species (human only), disease classification, cell-type annotations, and experimental preparation protocols to ensure atlas quality and consistency. Raw sequencing data were processed using a standardized pipeline. FASTQ files were aligned to the human reference genome (GRCh38) with CellRanger (v7.1.0), generating Gene expression count matrices for further analysis. This unified preprocessing workflow ensured comparability across all datasets.

Quality Control and Integration

Gene expression matrices were analyzed using Scanpy from the scverse framework. Standard quality control procedures were applied to remove low-quality cells and potential doublets. Data were normalized, log-transformed, and highly variable genes were selected for downstream analysis. Batch effects across datasets were corrected using Harmony, followed by dimensionality-reduced embeddings using principal component analysis and uniform manifold approximation and projection (UMAP). Cell clustering was conducted using the Leiden algorithm.

Cell-Type Annotation

Cell-type annotation was initially performed with CellTypist to get major cell groups, which was validated with DEG analysis and cross referenced with markers from literature and CellMarker database. More detailed annotation of T cells was performed using starCAT.

Protein Interaction Network

Cell-to-cell interactions were inferred using CellphoneDB. Consensus CCI networks were constructed for each disease condition by retaining interactions consistently observed across iterations.

Disease-associated DEG genes were identified for each condition relative to healthy controls using the memento-de method, while cell-type-specific marker genes were identified by contrasting the expression profiles of each cell type with those of all other cell types in the healthy cohort. The subset of all upregulated genes were provided to PrePPI to create protein protein interaction networks. For each cell type, two complementary PPI subnetworks were constructed by intersecting the prior network with the gene sets defined above: (i) a baseline subnetwork comprising interactions among cell-type–specific marker genes, representing constitutive cellular wiring, and (ii) disease-specific subnetworks comprising interactions among genes upregulated in each disease condition, capturing pathology-driven rewiring. Interactions were retained only among genes with detectable expression in the corresponding cell-type and were further prioritized by PrePPI-AF LR score and structural compatibility.

Annotation and Enrichment Analysis

The reviewed dataset from UniProt was utilized to annotate gene products. Finally, Gene Ontology (GO) enrichment analysis was performed to identify biological processes associated with the network.

Curated Datasets

Dataset Publication year Organization Platform Cells
GSE1361032019University of EdinburghIllumina HiSeq 400057,677
GSE1496142021State Key Laboratory of Proteomics, Beijing Proteome Research Center, Beijing Institute of Radiation MedicineIllumina NovaSeq 600068,985
GSE1515302021National Cancer InstituteIllumina NextSeq 500; Illumina HiSeq 4000; Illumina NovaSeq 600048,973
GSE1566252021Genome Institute Of SingaporeIllumina HiSeq 250044,410
GSE1599772021M3 Forschungszentrum für Malignom, Metabolom und MikrobiomIllumina NovaSeq 600080,056
GSE1694462021Weizmann Institute of ScienceIllumina NovaSeq 60009,210
GSE1719002021Istituto Clinico HumanitasNextSeq 55028,531
GSE1899032022National Cancer InstituteIllumina NovaSeq 6000109,767
GSE1927422022VIB-University of GhentIllumina HiSeq 4000; Illumina NovaSeq 6000187,162
GSE2363822024Shandong UniversityHiSeq X Ten25,668
GSE2428892024Wenzhou Medical UniversityIllumina NovaSeq 600042,443
GSE2439812024University of TorontoIllumina HiSeq 2500; Illumina NovaSeq 6000275,106
PRJNA8368682022Zhongshan HospitalIllumina NovaSeq 600093,724
PRJNA9329372023University of Hong KongIllumina NovaSeq 600038,409
PRJNA9472702023The First Hospital of Jilin UniversityIllumina NovaSeq 600039,765
Liver_TSP1_Nov1220242022Tabula Sapiens Consortium10x Genomics Chromium; Illumina NovaSeq 60007,248

Data Download

Download the LiverHomo metadata dataset.

Download adata_reprocessed_gzip.h5ad