Methods
Human liver scRNA-seq datasets were systematically curated from public repositories, including: (i) the NCBI Gene Expression Omnibus (GEO); (ii) the Sequence Read Archive (SRA); and (iii) the Tabula Sapiens consortium. Studies were selected based on relevance to human liver tissues across healthy and disease states, with inclusion criteria requiring a 10x Genomics-compatible platform and metadata including species (human only), disease classification, cell-type annotations, and experimental preparation protocols to ensure atlas quality and consistency. Raw sequencing data were processed using a standardized pipeline. FASTQ files were aligned to the human reference genome (GRCh38) with CellRanger (v7.1.0), generating Gene expression count matrices for further analysis. This unified preprocessing workflow ensured comparability across all datasets.
Quality Control and IntegrationGene expression matrices were analyzed using Scanpy from the scverse framework. Standard quality control procedures were applied to remove low-quality cells and potential doublets. Data were normalized, log-transformed, and highly variable genes were selected for downstream analysis. Batch effects across datasets were corrected using Harmony, followed by dimensionality-reduced embeddings using principal component analysis and uniform manifold approximation and projection (UMAP). Cell clustering was conducted using the Leiden algorithm.
Cell-Type AnnotationCell-type annotation was initially performed with CellTypist to get major cell groups, which was validated with DEG analysis and cross referenced with markers from literature and CellMarker database. More detailed annotation of T cells was performed using starCAT.
Protein Interaction NetworkCell-to-cell interactions were inferred using CellphoneDB. Consensus CCI networks were constructed for each disease condition by retaining interactions consistently observed across iterations.
Disease-associated DEG genes were identified for each condition relative to healthy controls using the memento-de method, while cell-type-specific marker genes were identified by contrasting the expression profiles of each cell type with those of all other cell types in the healthy cohort. The subset of all upregulated genes were provided to PrePPI to create protein protein interaction networks. For each cell type, two complementary PPI subnetworks were constructed by intersecting the prior network with the gene sets defined above: (i) a baseline subnetwork comprising interactions among cell-type–specific marker genes, representing constitutive cellular wiring, and (ii) disease-specific subnetworks comprising interactions among genes upregulated in each disease condition, capturing pathology-driven rewiring. Interactions were retained only among genes with detectable expression in the corresponding cell-type and were further prioritized by PrePPI-AF LR score and structural compatibility.
Annotation and Enrichment AnalysisThe reviewed dataset from UniProt was utilized to annotate gene products. Finally, Gene Ontology (GO) enrichment analysis was performed to identify biological processes associated with the network.
Curated Datasets
| Dataset | Publication year | Organization | Platform | Cells |
|---|---|---|---|---|
| GSE136103 | 2019 | University of Edinburgh | Illumina HiSeq 4000 | 57,677 |
| GSE149614 | 2021 | State Key Laboratory of Proteomics, Beijing Proteome Research Center, Beijing Institute of Radiation Medicine | Illumina NovaSeq 6000 | 68,985 |
| GSE151530 | 2021 | National Cancer Institute | Illumina NextSeq 500; Illumina HiSeq 4000; Illumina NovaSeq 6000 | 48,973 |
| GSE156625 | 2021 | Genome Institute Of Singapore | Illumina HiSeq 2500 | 44,410 |
| GSE159977 | 2021 | M3 Forschungszentrum für Malignom, Metabolom und Mikrobiom | Illumina NovaSeq 6000 | 80,056 |
| GSE169446 | 2021 | Weizmann Institute of Science | Illumina NovaSeq 6000 | 9,210 |
| GSE171900 | 2021 | Istituto Clinico Humanitas | NextSeq 550 | 28,531 |
| GSE189903 | 2022 | National Cancer Institute | Illumina NovaSeq 6000 | 109,767 |
| GSE192742 | 2022 | VIB-University of Ghent | Illumina HiSeq 4000; Illumina NovaSeq 6000 | 187,162 |
| GSE236382 | 2024 | Shandong University | HiSeq X Ten | 25,668 |
| GSE242889 | 2024 | Wenzhou Medical University | Illumina NovaSeq 6000 | 42,443 |
| GSE243981 | 2024 | University of Toronto | Illumina HiSeq 2500; Illumina NovaSeq 6000 | 275,106 |
| PRJNA836868 | 2022 | Zhongshan Hospital | Illumina NovaSeq 6000 | 93,724 |
| PRJNA932937 | 2023 | University of Hong Kong | Illumina NovaSeq 6000 | 38,409 |
| PRJNA947270 | 2023 | The First Hospital of Jilin University | Illumina NovaSeq 6000 | 39,765 |
| Liver_TSP1_Nov122024 | 2022 | Tabula Sapiens Consortium | 10x Genomics Chromium; Illumina NovaSeq 6000 | 7,248 |
Data Download
Download the LiverHomo metadata dataset.
Download adata_reprocessed_gzip.h5ad