본문

서브메뉴

On Statistical Learning for Structural Data: Data Fusion and Semi-Supervised Learning : 结构化数据中的统计学习:数据融合与半监督学习
On Statistical Learning for Structural Data: Data Fusion and Semi-Supervised Learning  : 结...
On Statistical Learning for Structural Data: Data Fusion and Semi-Supervised Learning : 结构化数据中的统计学习:数据融合与半监督学习

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103822
ISBN  
9798280719101
DDC  
310
저자명  
Jiang, Yicong.
서명/저자  
On Statistical Learning for Structural Data: Data Fusion and Semi-Supervised Learning : 结构化数据中的统计学习:数据融合与半监督学习
발행사항  
[Sl] : Harvard University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
242 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
주기사항  
Includes supplementary digital materials.
주기사항  
Advisor: Janson, Lucas;Ke, Zheng Tracy.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2025.
초록/해제  
요약Nowadays, with the rise of the large data era, data tends to become more complex and structural. The data may incorporate various different sources, with distinct data quality, sample size, or set of covariates. For instance, in ecological inference and validation studies of epidemiology, it is common that the predictor and response of interest are gathered in different datasets. However, many statistical estimands of interest (e.g., in regression or causality) are functions of the joint distribution of multiple random variables. In this scenario, the only possible approach is one of data fusion, where multiple independent data sets, each measuring a subset of the random variables of interest, are combined for inference. In general, since all random variables are never observed jointly, their joint distribution, and hence also the estimand which is a function of it, is only partially identifiable. Unfortunately, the endpoints of the partially identifiable region depend in general on entire conditional distributions, rendering them hard both operationally and statistically to estimate. Inspired by this challenge, in the first chapter, we present a novel outer-bound on the region of partial identifiability (and establish conditions under which it is tight) that depends only on certain conditional first and second moments. This allows us to derive semiparametrically efficient estimators of our endpoint outer-bounds that only require the standard machine learning toolbox which learns conditional means. We prove asymptotic normality and semiparametric efficiency of our estimators and provide consistent estimators of their variances, enabling asymptotically valid confidence interval construction for our original partially identifiable estimand. We demonstrate the utility of our method in simulations and a data fusion problem from economics.Beyond multi-source datasets, specialized formats--such as network data--are also a crucial component of structured data. With the large amount of networks and graphs in the modern data age, such as social networks (Facebook, LinkedIn, etc.), citation networks, and medical networks (e.g., propagation networks of treatment, infection networks of viruses), a proper theory for unveiling the underlying network sub-structures like communities becomes increasingly essential. Motivated by social network analysis and network-based recommendation systems, in the second chapter, we study a semi-supervised community detection problem in which the objective is to estimate the community label of a new node using the network topology and partially observed community labels of existing nodes. The network is modeled using a degree-corrected stochastic block model, which allows for severe degree heterogeneity and potentially non-assortative communities. We propose an algorithm that computes a `structural similarity metric' between the new node and each of the K communities by aggregating labeled and unlabeled data. The estimated label of the new node corresponds to the value of k that maximizes this similarity metric. Our method is fast and numerically outperforms existing semi-supervised algorithms. Theoretically, we derive explicit bounds for the misclassification error and show the efficiency of our method by comparing it with an ideal classifier. Our findings highlight, to the best of our knowledge, the first semi-supervised community detection algorithm that offers theoretical guarantees.An extension of community detection is mixed membership estimation (MME), which is a classical problem in network data analysis. It extends community detection by allowing a node to have fractional memberships in multiple communities. One major challenge in mixed membership estimation is to discover the underlying simplex structure of the high-dimensional data, which corresponds to the membership of each individual. Previous literature mainly focuses on leveraging the data itself to identify the simplex in an unsupervised learning fashion. However, in many scenarios, prior information may be available. For instance, the past user history of certain individuals in a social network may provide a clue to their membership. To incorporate this class of knowledge, in the third chapter, we consider a semi-supervised setting where the membership vectors of a subset of nodes are given. This problem is significantly more challenging than semi-supervised community detection, and to the best of our knowledge, it has not been studied in most previous literature. We discover an insightful structural equation for utilizing the known labels, which inspires a delicate way of extending unsupervised mixed-membership estimation algorithms to the semi-supervised setting. Compared to unsupervised algorithms for identifying the vertex structure, our method does not require the existence of pure nodes and needs fewer regularity conditions on the vertices. Additionally, assuming a degree-corrected mixed membership model, we provide theoretical guarantees of our algorithm, and show that it is efficient in the sense that its error rate is the same as a least squares problem with more prior information on the vertices. To the best of our knowledge, this is the first semi-supervised vertex hunting algorithm with a theoretical guarantee. We also demonstrate the excellent performance of our algorithm in several empirical studies, illustrating that with only a tiny fraction of the label information, our method can dramatically outperform unsupervised algorithms.
일반주제명  
Statistics
일반주제명  
Computer science
키워드  
Community detection
키워드  
Data fusion
키워드  
Double machine learning
키워드  
Mixed-membership estimation
키워드  
Semi-supervised learning
키워드  
Semiparametric
기타저자  
Harvard University Statistics
기본자료저록  
Dissertations Abstracts International. 86-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017358261
■00520260202103822
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798280719101
■035    ▼a(MiAaPQ)AAI32041390
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a310
■1001  ▼aJiang,  Yicong.▼0(orcid)0009-0001-4354-4732
■24510▼aOn  Statistical  Learning  for  Structural  Data:  Data  Fusion  and  Semi-Supervised  Learning  ▼b结构化数据中的统计学习:数据融合与半监督学习
■260    ▼a[Sl]▼bHarvard  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a242  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-12,  Section:  B.
■500    ▼aIncludes  supplementary  digital  materials.
■500    ▼aAdvisor:  Janson,  Lucas;Ke,  Zheng  Tracy.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2025.
■520    ▼aNowadays,  with  the  rise  of  the  large  data  era,  data  tends  to  become  more  complex  and  structural.  The  data  may  incorporate  various  different  sources,  with  distinct  data  quality,  sample  size,  or  set  of  covariates.  For  instance,  in  ecological  inference  and  validation  studies  of  epidemiology,  it  is  common  that  the  predictor  and  response  of  interest  are  gathered  in  different  datasets.  However,  many  statistical  estimands  of  interest  (e.g.,  in  regression  or  causality)  are  functions  of  the  joint  distribution  of  multiple  random  variables.  In  this  scenario,  the  only  possible  approach  is  one  of  data  fusion,  where  multiple  independent  data  sets,  each  measuring  a  subset  of  the  random  variables  of  interest,  are  combined  for  inference.  In  general,  since  all  random  variables  are  never  observed  jointly,  their  joint  distribution,  and  hence  also  the  estimand  which  is  a  function  of  it,  is  only  partially  identifiable.  Unfortunately,  the  endpoints  of  the  partially  identifiable  region  depend  in  general  on  entire  conditional  distributions,  rendering  them  hard  both  operationally  and  statistically  to  estimate.  Inspired  by  this  challenge,  in  the  first  chapter,  we  present  a  novel  outer-bound  on  the  region  of  partial  identifiability  (and  establish  conditions  under  which  it  is  tight)  that  depends  only  on  certain  conditional  first  and  second  moments.  This  allows  us  to  derive  semiparametrically  efficient  estimators  of  our  endpoint  outer-bounds  that  only  require  the  standard  machine  learning  toolbox  which  learns  conditional  means.  We  prove  asymptotic  normality  and  semiparametric  efficiency  of  our  estimators  and  provide  consistent  estimators  of  their  variances,  enabling  asymptotically  valid  confidence  interval  construction  for  our  original  partially  identifiable  estimand.  We  demonstrate  the  utility  of  our  method  in  simulations  and  a  data  fusion  problem  from  economics.Beyond  multi-source  datasets,  specialized  formats--such  as  network  data--are  also  a  crucial  component  of  structured  data.  With  the  large  amount  of  networks  and  graphs  in  the  modern  data  age,  such  as  social  networks  (Facebook,  LinkedIn,  etc.),  citation  networks,  and  medical  networks  (e.g.,  propagation  networks  of  treatment,  infection  networks  of  viruses),  a  proper  theory  for  unveiling  the  underlying  network  sub-structures  like  communities  becomes  increasingly  essential.  Motivated  by  social  network  analysis  and  network-based  recommendation  systems,  in  the  second  chapter,  we  study  a  semi-supervised  community  detection  problem  in  which  the  objective  is  to  estimate  the  community  label  of  a  new  node  using  the  network  topology  and  partially  observed  community  labels  of  existing  nodes.  The  network  is  modeled  using  a  degree-corrected  stochastic  block  model,  which  allows  for  severe  degree  heterogeneity  and  potentially  non-assortative  communities.  We  propose  an  algorithm  that  computes  a  `structural  similarity  metric'  between  the  new  node  and  each  of  the  K  communities  by  aggregating  labeled  and  unlabeled  data.  The  estimated  label  of  the  new  node  corresponds  to  the  value  of  k  that  maximizes  this  similarity  metric.  Our  method  is  fast  and  numerically  outperforms  existing  semi-supervised  algorithms.  Theoretically,  we  derive  explicit  bounds  for  the  misclassification  error  and  show  the  efficiency  of  our  method  by  comparing  it  with  an  ideal  classifier.  Our  findings  highlight,  to  the  best  of  our  knowledge,  the  first  semi-supervised  community  detection  algorithm  that  offers  theoretical  guarantees.An  extension  of  community  detection  is  mixed  membership  estimation  (MME),  which  is  a  classical  problem  in  network  data  analysis.  It  extends  community  detection  by  allowing  a  node  to  have  fractional  memberships  in  multiple  communities.  One  major  challenge  in  mixed  membership  estimation  is  to  discover  the  underlying  simplex  structure  of  the  high-dimensional  data,  which  corresponds  to  the  membership  of  each  individual.  Previous  literature  mainly  focuses  on  leveraging  the  data  itself  to  identify  the  simplex  in  an  unsupervised  learning  fashion.  However,  in  many  scenarios,  prior  information  may  be  available.  For  instance,  the  past  user  history  of  certain  individuals  in  a  social  network  may  provide  a  clue  to  their  membership.  To  incorporate  this  class  of  knowledge,  in  the  third  chapter,  we  consider  a  semi-supervised  setting  where  the  membership  vectors  of  a  subset  of  nodes  are  given.  This  problem  is  significantly  more  challenging  than  semi-supervised  community  detection,  and  to  the  best  of  our  knowledge,  it  has  not  been  studied  in  most  previous  literature.  We  discover  an  insightful  structural  equation  for  utilizing  the  known  labels,  which  inspires  a  delicate  way  of  extending  unsupervised  mixed-membership  estimation  algorithms  to  the  semi-supervised  setting.  Compared  to  unsupervised  algorithms  for  identifying  the  vertex  structure,  our  method  does  not  require  the  existence  of  pure  nodes  and  needs  fewer  regularity  conditions  on  the  vertices.  Additionally,  assuming  a  degree-corrected  mixed  membership  model,  we  provide  theoretical  guarantees  of  our  algorithm,  and  show  that  it  is  efficient  in  the  sense  that  its  error  rate  is  the  same  as  a  least  squares  problem  with  more  prior  information  on  the  vertices.  To  the  best  of  our  knowledge,  this  is  the  first  semi-supervised  vertex  hunting  algorithm  with  a  theoretical  guarantee.  We  also  demonstrate  the  excellent  performance  of  our  algorithm  in  several  empirical  studies,  illustrating  that  with  only  a  tiny  fraction  of  the  label  information,  our  method  can  dramatically  outperform  unsupervised  algorithms.
■590    ▼aSchool  code:  0084.
■650  4▼aStatistics
■650  4▼aComputer  science
■653    ▼aCommunity  detection
■653    ▼aData  fusion
■653    ▼aDouble  machine  learning
■653    ▼aMixed-membership  estimation
■653    ▼aSemi-supervised  learning
■653    ▼aSemiparametric
■690    ▼a0463
■690    ▼a0984
■690    ▼a0501
■71020▼aHarvard  University▼bStatistics.
■7730  ▼tDissertations  Abstracts  International▼g86-12B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358261▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17499 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.