서브메뉴
검색
Advances in Synthetic Data Generation and Causal Inference
Advances in Synthetic Data Generation and Causal Inference
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103034
- ISBN
- 9798286444434
- DDC
- 310
- 서명/저자
- Advances in Synthetic Data Generation and Causal Inference
- 발행사항
- [Sl] : Yale University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 149 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: A.
- 주기사항
- Advisor: Sekhon, Jasjeet;Forastiere, Laura.
- 학위논문주기
- Thesis (Ph.D.)--Yale University, 2025.
- 초록/해제
- 요약Synthetic data generation has become essential for addressing data scarcity, privacy concerns, and generalization problems when applying machine learning in various domains. Hence, data is the main ingredient of any machine learning model. If synthetic data is too similar to real data, the risk of privacy breaches is significantly higher; if it is too different, its practical applicability is undermined. This dissertation proposes two approaches to enhance synthetic data generation. Chapter 2 develops SC-GOAT, a framework that integrates a supervised component tailored to the specific downstream task and employs a meta-learning approach to learn the optimal mixture distribution of existing synthetic distributions. Thus, synthetic data are made more useful for specific real-world applications. Chapter 3 introduces a framework designed to enhance existing clinical models, Private Synthetic Hypercube Augmentation (PriSHA). We use generative models to produce synthetic data as a means to augment these models while adhering to strict privacy standards. This approach has the potential to improve model performance without compromising patient confidentiality. To our knowledge, our framework is the first synthetic data augmentation framework that merges privacy-preserving tabular data and real data from multiple sources.Causal inference is central to distinguishing causation from correlation and thus facilitating informed decision-making in many fields, from economics to epidemiology and artificial intelligence. This dissertation makes two contributions to the literature on causal inference. Chapter 4 introduces CLOUD-CG, a clustering method for longitudinal data that uses temporal-directed acyclic graphs (T-DAG) to identify clusters with similar causal structures. While preserving individual-level heterogeneity, CLOUD-CG provides interpretable insights into time-dependent causal representation to evaluate financial stability in emerging economies. Chapter 5 introduces the causal machine learning model in sports analytics through its application to age-curve modeling. The Age-Conditioned Treatment Effect (ACTE) is presented to investigate the causal impact of interventions such as rest days on the performance of athletes at different stages of their careers. Using ACTE in a meta-learning framework, this work provides a load management strategy based on granular game-level data in professional sports.This dissertation advances the data science pipeline, developing synthetic data generation methods to improve data quality and availability, and causal inference frameworks to learn causal relationships. Approaching the shortcomings in both domains reinforces the reliability of the decision-making process in diverse fields.
- 일반주제명
- Statistics
- 일반주제명
- Computer science
- 일반주제명
- Information science
- 키워드
- Causal discovery
- 키워드
- Causal inference
- 키워드
- Data clustering
- 키워드
- Synthetic data
- 기타저자
- Yale University Statistics and Data Science
- 기본자료저록
- Dissertations Abstracts International. 86-12A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017356782
■00520260202103034
■006m o d
■007cr#unu||||||||
■020 ▼a9798286444434
■035 ▼a(MiAaPQ)AAI31846141
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aNakamura-Sakai, Shinpei.
■24510▼aAdvances in Synthetic Data Generation and Causal Inference
■260 ▼a[Sl]▼bYale University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a149 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: A.
■500 ▼aAdvisor: Sekhon, Jasjeet;Forastiere, Laura.
■5021 ▼aThesis (Ph.D.)--Yale University, 2025.
■520 ▼aSynthetic data generation has become essential for addressing data scarcity, privacy concerns, and generalization problems when applying machine learning in various domains. Hence, data is the main ingredient of any machine learning model. If synthetic data is too similar to real data, the risk of privacy breaches is significantly higher; if it is too different, its practical applicability is undermined. This dissertation proposes two approaches to enhance synthetic data generation. Chapter 2 develops SC-GOAT, a framework that integrates a supervised component tailored to the specific downstream task and employs a meta-learning approach to learn the optimal mixture distribution of existing synthetic distributions. Thus, synthetic data are made more useful for specific real-world applications. Chapter 3 introduces a framework designed to enhance existing clinical models, Private Synthetic Hypercube Augmentation (PriSHA). We use generative models to produce synthetic data as a means to augment these models while adhering to strict privacy standards. This approach has the potential to improve model performance without compromising patient confidentiality. To our knowledge, our framework is the first synthetic data augmentation framework that merges privacy-preserving tabular data and real data from multiple sources.Causal inference is central to distinguishing causation from correlation and thus facilitating informed decision-making in many fields, from economics to epidemiology and artificial intelligence. This dissertation makes two contributions to the literature on causal inference. Chapter 4 introduces CLOUD-CG, a clustering method for longitudinal data that uses temporal-directed acyclic graphs (T-DAG) to identify clusters with similar causal structures. While preserving individual-level heterogeneity, CLOUD-CG provides interpretable insights into time-dependent causal representation to evaluate financial stability in emerging economies. Chapter 5 introduces the causal machine learning model in sports analytics through its application to age-curve modeling. The Age-Conditioned Treatment Effect (ACTE) is presented to investigate the causal impact of interventions such as rest days on the performance of athletes at different stages of their careers. Using ACTE in a meta-learning framework, this work provides a load management strategy based on granular game-level data in professional sports.This dissertation advances the data science pipeline, developing synthetic data generation methods to improve data quality and availability, and causal inference frameworks to learn causal relationships. Approaching the shortcomings in both domains reinforces the reliability of the decision-making process in diverse fields.
■590 ▼aSchool code: 0265.
■650 4▼aStatistics
■650 4▼aComputer science
■650 4▼aInformation science
■653 ▼aCausal discovery
■653 ▼aCausal inference
■653 ▼aData clustering
■653 ▼aDifferential privacy
■653 ▼aLongitudinal data analysis
■653 ▼aSynthetic data
■690 ▼a0463
■690 ▼a0984
■690 ▼a0800
■690 ▼a0723
■71020▼aYale University▼bStatistics and Data Science.
■7730 ▼tDissertations Abstracts International▼g86-12A.
■790 ▼a0265
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17356782▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


