서브메뉴
검색
Novel Experimental Design Techniques for Data Science
Novel Experimental Design Techniques for Data Science
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105521
- ISBN
- 9798263341213
- DDC
- 600
- 저자명
- Huang, Chaofan.
- 서명/저자
- Novel Experimental Design Techniques for Data Science
- 발행사항
- [Sl] : Georgia Institute of Technology, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 215 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
- 주기사항
- Advisor: Joseph, Roshan V.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2024.
- 초록/해제
- 요약Experimental design [1, 2, 3] is a fundamental area of statistics, with many fascinating techniques developed in recent years. Many of these methodologies possess broader applications beyond experimental design-related problems but are not yet explored. This thesis presents novel applications of experimental design techniques for data science problems, with a focus on sampling, optimization, and machine learning.Chapter 1 and Chapter 2 integrate principles from optimal space-filling design into the weighted resampling methods, an integral part of sequential sampling, survey sampling, etc. The most straightforward approach is to draw sample independently with replacement according to their weights. However, this may result in clustered resamples that provide duplicated information and fail to cover certain regions of the sample space. Chapter 1 introduces a novel deterministic weighted sampling scheme known as the Importance Support Points (ISP) resampling. ISP resampling selects optimal resamples that not only best represent the weighted samples in terms of energy distance but also ensures space-fillingness. We incorporate ISP resampling into sequential sampling methods, and demonstrate its empirical improvement over the existing weighted resampling techniques. However, the quadratic complexity of ISP computation restricts its practicality in large data settings. To address this shortcoming, Chapter 2 presents Weighted Twinning, a nearest-neighbor based heuristic algorithm that is orders of magnitude faster for computing ISP, making it applicable to a broader class of problems.Chapter 3 explores another deterministic resampling method based on the minimum energy design (MinED). MinED resampling also aims to find the set of space-filling resamples that are representative for the target distribution. The effectiveness of MinED resampling is illustrated by its integration with sequential sampling for constructing spacefilling design in highly constrained regions. Extensive simulation results are provided to demonstrate the improved performance over existing state-of-the-art techniques.Chapter 4 presents a novel sequential design technique for efficiently calibrating (optimizing) the parameters of a functional output model. The proposed algorithm improves over the standard Bayesian optimization by (i) utilizing the generalized chi-square distribution as a more appropriate predictive distribution for the squared distance objective function in the calibration problems, and (ii) applying functional principal component analysis to reduce the dimensionality of the functional response data, which allows for efficient approximation of the predictive distribution and subsequent computation of the expected improvement acquisition function.Finally, Chapter 5 applies the variance-based global sensitivity analysis for factor importance computation, one of the fundamental problems in statistics and machine learning. Many existing works focus on the model-based importance, but an important feature in one learning algorithm may hold little significance in another learning algorithm. Hence, a factor importance measure ought to characterize the feature's predictive potential without relying on a specific prediction algorithm. To bypass the modeling step, the equivalence between predictive potential and total Sobol' indices is drawn, and a consistent estimator using only the noisy data is proposed. Integrating with forward selection and backward elimination gives rise to a novel algorithm for factor importance ranking and selection. The effectiveness of the algorithm is demonstrated in simulations and real world examples.
- 일반주제명
- Friction
- 일반주제명
- Inverse problems
- 일반주제명
- Normal distribution
- 일반주제명
- Data science
- 일반주제명
- Energy
- 일반주제명
- Design techniques
- 일반주제명
- Mathematics
- 기본자료저록
- Dissertations Abstracts International. 87-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2024 us c eng d■001000017360414
■00520260202105521
■006m o d
■007cr#unu||||||||
■020 ▼a9798263341213
■035 ▼a(MiAaPQ)AAI32309571
■035 ▼a(MiAaPQ)GeorgiaTech77669
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a600
■1001 ▼aHuang, Chaofan.
■24510▼aNovel Experimental Design Techniques for Data Science
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a215 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: B.
■500 ▼aAdvisor: Joseph, Roshan V.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2024.
■520 ▼aExperimental design [1, 2, 3] is a fundamental area of statistics, with many fascinating techniques developed in recent years. Many of these methodologies possess broader applications beyond experimental design-related problems but are not yet explored. This thesis presents novel applications of experimental design techniques for data science problems, with a focus on sampling, optimization, and machine learning.Chapter 1 and Chapter 2 integrate principles from optimal space-filling design into the weighted resampling methods, an integral part of sequential sampling, survey sampling, etc. The most straightforward approach is to draw sample independently with replacement according to their weights. However, this may result in clustered resamples that provide duplicated information and fail to cover certain regions of the sample space. Chapter 1 introduces a novel deterministic weighted sampling scheme known as the Importance Support Points (ISP) resampling. ISP resampling selects optimal resamples that not only best represent the weighted samples in terms of energy distance but also ensures space-fillingness. We incorporate ISP resampling into sequential sampling methods, and demonstrate its empirical improvement over the existing weighted resampling techniques. However, the quadratic complexity of ISP computation restricts its practicality in large data settings. To address this shortcoming, Chapter 2 presents Weighted Twinning, a nearest-neighbor based heuristic algorithm that is orders of magnitude faster for computing ISP, making it applicable to a broader class of problems.Chapter 3 explores another deterministic resampling method based on the minimum energy design (MinED). MinED resampling also aims to find the set of space-filling resamples that are representative for the target distribution. The effectiveness of MinED resampling is illustrated by its integration with sequential sampling for constructing spacefilling design in highly constrained regions. Extensive simulation results are provided to demonstrate the improved performance over existing state-of-the-art techniques.Chapter 4 presents a novel sequential design technique for efficiently calibrating (optimizing) the parameters of a functional output model. The proposed algorithm improves over the standard Bayesian optimization by (i) utilizing the generalized chi-square distribution as a more appropriate predictive distribution for the squared distance objective function in the calibration problems, and (ii) applying functional principal component analysis to reduce the dimensionality of the functional response data, which allows for efficient approximation of the predictive distribution and subsequent computation of the expected improvement acquisition function.Finally, Chapter 5 applies the variance-based global sensitivity analysis for factor importance computation, one of the fundamental problems in statistics and machine learning. Many existing works focus on the model-based importance, but an important feature in one learning algorithm may hold little significance in another learning algorithm. Hence, a factor importance measure ought to characterize the feature's predictive potential without relying on a specific prediction algorithm. To bypass the modeling step, the equivalence between predictive potential and total Sobol' indices is drawn, and a consistent estimator using only the noisy data is proposed. Integrating with forward selection and backward elimination gives rise to a novel algorithm for factor importance ranking and selection. The effectiveness of the algorithm is demonstrated in simulations and real world examples.
■590 ▼aSchool code: 0078.
■650 4▼aFriction
■650 4▼aInverse problems
■650 4▼aNormal distribution
■650 4▼aData science
■650 4▼aEnergy
■650 4▼aDesign techniques
■650 4▼aMathematics
■690 ▼a0791
■690 ▼a0800
■690 ▼a0405
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-05B.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360414▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


