서브메뉴
검색
Advancing Statistical Rigor in Single-Cell and Spatial Omics Analysis Through In Silico Control Data
Advancing Statistical Rigor in Single-Cell and Spatial Omics Analysis Through In Silico Control Data
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103635
- ISBN
- 9798315799320
- DDC
- 310
- 저자명
- Yan, Guanao.
- 서명/저자
- Advancing Statistical Rigor in Single-Cell and Spatial Omics Analysis Through In Silico Control Data
- 발행사항
- [Sl] : University of California, Los Angeles, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 198 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
- 주기사항
- Advisor: Li, Jingyi.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Los Angeles, 2025.
- 초록/해제
- 요약Over the past decade, single-cell and spatial transcriptomics technologies have transformed our ability to study cellular diversity and tissue organization. These advances have led to the rapid development of computational methods for analyzing high-dimensional omics data. However, benchmarking these methods and ensuring their statistical rigor remain challenging, largely due to the absence of realistic synthetic data with ground truths and the conceptual ambiguity in defining key biological features such as spatially variable genes (SVGs). This dissertation addresses these gaps through two simulation frameworks and a comprehensive review that improve the statistical rigor and interpretability of tool development and evaluation.My first project introduces scReadSim, a simulator designed to generate realistic synthetic data for single-cell RNA sequencing (scRNA-seq) and chromatin accessibility profiling (scATAC-seq). It produces simulated sequencing reads in standard formats by mimicking the characteristics of real datasets, while allowing users to specify key ground truths, such as transcript abundance for scRNA-seq and cell-type-specific open chromatin regions for scATAC-seq. scReadSim supports flexible simulation settings, including varying cell numbers and sequencing depths, and enables systematic benchmarking of preprocessing tools. Using scReadSim, we show that UMI-tools achieves higher accuracy in transcript quantification for scRNA-seq, while HMMRATAC and MACS3 perform best in peak calling for scATAC-seq.My second project presents scIsoSim, a simulator that generates single-cell RNA sequencing data with known isoform structures and their corresponding expression levels. In gene expression, a single gene can give rise to multiple isoforms-different versions of RNA transcripts-through a biological process called alternative splicing, where segments of RNA are included or excluded in various combinations. scIsoSim supports widely used experimental protocols, including Smart-seq2 and 10x Genomics 3' and 5' platforms, and captures realistic splicing patterns observed in real datasets. This tool enables systematic evaluation of computational methods for quantifying isoform expression and detecting alternative splicing events. Benchmarking results show that bulk RNA-seq tools, such as Salmon, perform accurately on Smart-seq2 data with high computational efficiency. In contrast, Scasa-the only existing tool for 10x 3' data-shows limited accuracy due to sparse data. Among splicing analysis tools, brie demonstrates better overall accuracy than outrigger but is less effective in detecting cell-specific splicing events.My third project is a review of 34 state-of-the-art SVG detection methods for spatial transcriptomics data. The review introduces a new categorization framework that defines SVGs as overall, cell-type-specific, or spatial-domain-marker genes, based on their spatial expression patterns and analytic objectives. It summarizes the underlying assumptions and statistical hypothesis tests used by each method, and discusses trade-offs between power and specificity. The review also identifies limitations in existing benchmarks, such as inappropriate method comparisons and oversimplified simulation designs, and calls for category-specific benchmarking using well-annotated datasets and realistic simulators.
- 일반주제명
- Statistics
- 일반주제명
- Cellular biology
- 일반주제명
- Biostatistics
- 일반주제명
- Genetics
- 키워드
- In silico
- 키워드
- Simulator
- 키워드
- Spatial omics
- 기타저자
- University of California, Los Angeles Statistics 0891
- 기본자료저록
- Dissertations Abstracts International. 86-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358039
■00520260202103635
■006m o d
■007cr#unu||||||||
■020 ▼a9798315799320
■035 ▼a(MiAaPQ)AAI32047458
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aYan, Guanao.
■24510▼aAdvancing Statistical Rigor in Single-Cell and Spatial Omics Analysis Through In Silico Control Data
■260 ▼a[Sl]▼bUniversity of California, Los Angeles▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a198 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: B.
■500 ▼aAdvisor: Li, Jingyi.
■5021 ▼aThesis (Ph.D.)--University of California, Los Angeles, 2025.
■520 ▼aOver the past decade, single-cell and spatial transcriptomics technologies have transformed our ability to study cellular diversity and tissue organization. These advances have led to the rapid development of computational methods for analyzing high-dimensional omics data. However, benchmarking these methods and ensuring their statistical rigor remain challenging, largely due to the absence of realistic synthetic data with ground truths and the conceptual ambiguity in defining key biological features such as spatially variable genes (SVGs). This dissertation addresses these gaps through two simulation frameworks and a comprehensive review that improve the statistical rigor and interpretability of tool development and evaluation.My first project introduces scReadSim, a simulator designed to generate realistic synthetic data for single-cell RNA sequencing (scRNA-seq) and chromatin accessibility profiling (scATAC-seq). It produces simulated sequencing reads in standard formats by mimicking the characteristics of real datasets, while allowing users to specify key ground truths, such as transcript abundance for scRNA-seq and cell-type-specific open chromatin regions for scATAC-seq. scReadSim supports flexible simulation settings, including varying cell numbers and sequencing depths, and enables systematic benchmarking of preprocessing tools. Using scReadSim, we show that UMI-tools achieves higher accuracy in transcript quantification for scRNA-seq, while HMMRATAC and MACS3 perform best in peak calling for scATAC-seq.My second project presents scIsoSim, a simulator that generates single-cell RNA sequencing data with known isoform structures and their corresponding expression levels. In gene expression, a single gene can give rise to multiple isoforms-different versions of RNA transcripts-through a biological process called alternative splicing, where segments of RNA are included or excluded in various combinations. scIsoSim supports widely used experimental protocols, including Smart-seq2 and 10x Genomics 3' and 5' platforms, and captures realistic splicing patterns observed in real datasets. This tool enables systematic evaluation of computational methods for quantifying isoform expression and detecting alternative splicing events. Benchmarking results show that bulk RNA-seq tools, such as Salmon, perform accurately on Smart-seq2 data with high computational efficiency. In contrast, Scasa-the only existing tool for 10x 3' data-shows limited accuracy due to sparse data. Among splicing analysis tools, brie demonstrates better overall accuracy than outrigger but is less effective in detecting cell-specific splicing events.My third project is a review of 34 state-of-the-art SVG detection methods for spatial transcriptomics data. The review introduces a new categorization framework that defines SVGs as overall, cell-type-specific, or spatial-domain-marker genes, based on their spatial expression patterns and analytic objectives. It summarizes the underlying assumptions and statistical hypothesis tests used by each method, and discusses trade-offs between power and specificity. The review also identifies limitations in existing benchmarks, such as inappropriate method comparisons and oversimplified simulation designs, and calls for category-specific benchmarking using well-annotated datasets and realistic simulators.
■590 ▼aSchool code: 0031.
■650 4▼aStatistics
■650 4▼aCellular biology
■650 4▼aBiostatistics
■650 4▼aGenetics
■653 ▼aIn silico
■653 ▼aSimulator
■653 ▼aSingle cell omics
■653 ▼aSpatial omics
■653 ▼aSpatially variable genes
■690 ▼a0463
■690 ▼a0379
■690 ▼a0369
■690 ▼a0308
■71020▼aUniversity of California, Los Angeles▼bStatistics 0891.
■7730 ▼tDissertations Abstracts International▼g86-12B.
■790 ▼a0031
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358039▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


