서브메뉴
검색
Statistical Learning for Recurrent Event and Complex Network Data
Statistical Learning for Recurrent Event and Complex Network Data
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105222
- ISBN
- 9798291566275
- DDC
- 310
- 저자명
- Meng, Bo.
- 서명/저자
- Statistical Learning for Recurrent Event and Complex Network Data
- 발행사항
- [Sl] : University of Michigan, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 260 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-03, Section: B.
- 주기사항
- Advisor: Xu, Gongjun;Zhu, Ji.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2025.
- 초록/해제
- 요약The rapid development of modern technology has led to an unprecedented deluge of complex data, presenting researchers with major challenges in scientific studies. These datasets are frequently large-scale, high-dimensional, possess complex structures, and are often incomplete. In these demanding scenarios, traditional statistical methods often fail, leading to incorrect conclusions or biased results. Additionally, many statistical models also suffer from computational inefficiency in the presence of large-scale data. Motivated by these challenges, this thesis develops novel statistical modeling frameworks and efficient computational methodologies tailored for large-scale recurrent event data and complex structured network data.Chapter 2 presents a general framework for analyzing recurrent event data by modeling the conditional mean function as the solution to an Ordinary Differential Equation (ODE). This approach covers a wide range of semi-parametric recurrent event models, including both non-homogeneous Poisson processes (NHPPs) and non-Poisson processes, while remaining scalable and easy-to-implement. Based on this framework, we propose a Sieve Maximum Pseudo-Likelihood Estimation (SMPLE) method, prove its consistency and asymptotic normality, and show that it achieves semi-parametric efficiency when the NHPP model is correct. We also develop an efficient resampling procedure to estimate the asymptotic covariance and demonstrate the method's performance through extensive simulation studies and an application to ICU readmission data.Chapter 3 focuses on signed network data, where relationships among entities exhibit a complex interplay of positive (e.g., liking and alliances) and negative (e.g., disliking and conflicts) interactions. This chapter proposes a novel latent space model that utilizes non-linear kernel functions to capture the sign-generating pattern of balanced signed networks, and identifies a new sufficient condition for a family of signed networks to achieve population-level balance. We develop efficient projected gradient descent (PGD) algorithms to estimate the latent variables, and establish non-asymptotic error rates for parameter estimation under both correctly specified and mis-specified settings. The efficiency and robustness of our methodology are validated through extensive simulation studies. Finally, we apply this method to an international relations dataset to visualize alliance and conflict patterns among nations during World War I.Chapter 4 introduces a latent space model for analyzing longitudinal network data. This approach employs multivariate counting processes to model interaction sequences between node pairs, where intensity functions depend on static latent variables, time-varying baseline intensities, and time-varying edge covariates. To estimate the model parameters, spline-based sieve estimators are utilized, and the objective function is maximized using an efficient projected gradient descent (PGD) algorithm. This chapter establishes the statistical convergence rate of the global maximizer of the objective function and the algorithmic error rate of the PGD estimator, demonstrating that these two rates match and jointly achieve the optimal error rates under both parametric and nonparametric settings, up to a logarithmic factor. Extensive simulation studies and an application to a real-world bike-sharing dataset demonstrate the efficiency and robustness of the proposed framework.
- 일반주제명
- Statistics
- 일반주제명
- Applied mathematics
- 키워드
- Signed networks
- 기타저자
- University of Michigan Statistics
- 기본자료저록
- Dissertations Abstracts International. 87-03B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359837
■00520260202105222
■006m o d
■007cr#unu||||||||
■020 ▼a9798291566275
■035 ▼a(MiAaPQ)AAI32271818
■035 ▼a(MiAaPQ)umichrackham006239
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aMeng, Bo.
■24510▼aStatistical Learning for Recurrent Event and Complex Network Data
■260 ▼a[Sl]▼bUniversity of Michigan▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a260 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-03, Section: B.
■500 ▼aAdvisor: Xu, Gongjun;Zhu, Ji.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2025.
■520 ▼aThe rapid development of modern technology has led to an unprecedented deluge of complex data, presenting researchers with major challenges in scientific studies. These datasets are frequently large-scale, high-dimensional, possess complex structures, and are often incomplete. In these demanding scenarios, traditional statistical methods often fail, leading to incorrect conclusions or biased results. Additionally, many statistical models also suffer from computational inefficiency in the presence of large-scale data. Motivated by these challenges, this thesis develops novel statistical modeling frameworks and efficient computational methodologies tailored for large-scale recurrent event data and complex structured network data.Chapter 2 presents a general framework for analyzing recurrent event data by modeling the conditional mean function as the solution to an Ordinary Differential Equation (ODE). This approach covers a wide range of semi-parametric recurrent event models, including both non-homogeneous Poisson processes (NHPPs) and non-Poisson processes, while remaining scalable and easy-to-implement. Based on this framework, we propose a Sieve Maximum Pseudo-Likelihood Estimation (SMPLE) method, prove its consistency and asymptotic normality, and show that it achieves semi-parametric efficiency when the NHPP model is correct. We also develop an efficient resampling procedure to estimate the asymptotic covariance and demonstrate the method's performance through extensive simulation studies and an application to ICU readmission data.Chapter 3 focuses on signed network data, where relationships among entities exhibit a complex interplay of positive (e.g., liking and alliances) and negative (e.g., disliking and conflicts) interactions. This chapter proposes a novel latent space model that utilizes non-linear kernel functions to capture the sign-generating pattern of balanced signed networks, and identifies a new sufficient condition for a family of signed networks to achieve population-level balance. We develop efficient projected gradient descent (PGD) algorithms to estimate the latent variables, and establish non-asymptotic error rates for parameter estimation under both correctly specified and mis-specified settings. The efficiency and robustness of our methodology are validated through extensive simulation studies. Finally, we apply this method to an international relations dataset to visualize alliance and conflict patterns among nations during World War I.Chapter 4 introduces a latent space model for analyzing longitudinal network data. This approach employs multivariate counting processes to model interaction sequences between node pairs, where intensity functions depend on static latent variables, time-varying baseline intensities, and time-varying edge covariates. To estimate the model parameters, spline-based sieve estimators are utilized, and the objective function is maximized using an efficient projected gradient descent (PGD) algorithm. This chapter establishes the statistical convergence rate of the global maximizer of the objective function and the algorithmic error rate of the PGD estimator, demonstrating that these two rates match and jointly achieve the optimal error rates under both parametric and nonparametric settings, up to a logarithmic factor. Extensive simulation studies and an application to a real-world bike-sharing dataset demonstrate the efficiency and robustness of the proposed framework.
■590 ▼aSchool code: 0127.
■650 4▼aStatistics
■650 4▼aApplied mathematics
■653 ▼aRecurrent event analysis
■653 ▼aSieve maximum likelihood estimator
■653 ▼aSigned networks
■653 ▼aLatent space models
■653 ▼aLongitudinal network
■690 ▼a0463
■690 ▼a0364
■71020▼aUniversity of Michigan▼bStatistics.
■7730 ▼tDissertations Abstracts International▼g87-03B.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359837▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


