서브메뉴
검색
Essays in Methodology
Essays in Methodology
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104846
- ISBN
- 9798293893201
- DDC
- 310
- 서명/저자
- Essays in Methodology
- 발행사항
- [Sl] : University of California, Berkeley, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 163 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-04, Section: B.
- 주기사항
- Advisor: Dunning, Thad.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Berkeley, 2025.
- 초록/해제
- 요약This dissertation studies three problems in statistical methodology.How should researchers select experimental sites when the deployment population may differ from observed data? The first paper, "Site Selection under Distribution shift via Optimal Transport and Wasserstein Distributionally-Robust Optimization'', formulates the problem of experimental site selection as an optimal transport problem, and develops methods to minimize downstream estimation error by choosing sites that minimize Wasserstein distances between population and sample covariate distributions. I develop new theoretical upper bounds on PATE and CATE estimation errors, and show that these different objectives lead to different site selection strategies: both approaches form balanced, representative partitions of the support of the covariates, but use different penalties, which place a larger emphasis on site selections that place greater weight on minimizing downstream bias (PATE) and and variance (CATE) respectively. I extend this approach by using Wasserstein Distributionally Robust Optimization to guard against distribution shift when observed sites may not represent the target population, and develop a novel, data-driven procedure for uncertainty radius selection. I develop a cutting-plane algorithm that solves the resulting minimax problem by exploiting its sequential game structure, combined with a data-adaptive procedure for calibrating robustness parameters without requiring arbitrary assumptions about distributional uncertainty. Simulation evidence, and a reanalysis of a randomized microcredit experiment in Morocco, show that these methods outperform random and stratified sampling of sites, and alternative optimization methods i) for moderate-to-large size problem instances ii) when covariates are moderately informative about treatment effects, and iii) under induced distribution shift.The second paper, "Collusive and Adversarial Replication'', studies a game in which social ties between members of a research community may discourage prospective replicators from debunking papers that misreport results. Here, replication is an entrance decision, as a Replicator chooses whether or not to Replicate a given paper. A high level of social connectedness between members of a research community increases the field-wise False Discovery Rate, a measure of the social welfare associated with a healthy publication process. The moral is that larger, more diverse academic fields with fewer social ties are more likely to have an adversarial culture around replication, and that this improves social welfare. I consider three proposals to improve replication practices: Random auditing, or police-patrol replication; automated unit tests; and a recent proposal to lower the threshold for statistical significance. I argue that random auditing and automated unit tests can improve social welfare, but that the effect of lowering the statistical significance threshold is ambiguous.The third paper, "Why LLMs Hallucinate'', argues that LLMs hallucinate because their output is not constrained to be synonymous with claims for which they have evidence: a condition I call evidential closure. Information about the truth or falsity of sentences is not statistically identified in the standard neural language generation setup, and so cannot be conditioned on to generate new strings. We then show how to constrain LLMs to produce output that satisfies evidential closure. A multimodal LLM must learn about the external world (perceptual learning); it must learn a mapping from strings to states of the world (extensional learning); and, to achieve fluency when generalizing beyond a body of evidence, it must learn mappings from strings to their synonyms (intensional learning). The output of a unimodal LLM must be synonymous with strings in a validated evidence set. Finally, I present a heuristic procedure, Learn-Babble-Prune, that yields faithful output from an LLM by rejecting output that is not synonymous with claims for which the LLM has evidence.
- 일반주제명
- Statistics
- 일반주제명
- Political science
- 일반주제명
- Computer science
- 키워드
- Causal inference
- 기타저자
- University of California, Berkeley Political Science
- 기본자료저록
- Dissertations Abstracts International. 87-04B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017359187
■00520260202104846
■006m o d
■007cr#unu||||||||
■020 ▼a9798293893201
■035 ▼a(MiAaPQ)AAI32173835
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a310
■1001 ▼aBouyamourn, Adam.
■24510▼aEssays in Methodology
■260 ▼a[Sl]▼bUniversity of California, Berkeley▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a163 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-04, Section: B.
■500 ▼aAdvisor: Dunning, Thad.
■5021 ▼aThesis (Ph.D.)--University of California, Berkeley, 2025.
■520 ▼aThis dissertation studies three problems in statistical methodology.How should researchers select experimental sites when the deployment population may differ from observed data? The first paper, "Site Selection under Distribution shift via Optimal Transport and Wasserstein Distributionally-Robust Optimization'', formulates the problem of experimental site selection as an optimal transport problem, and develops methods to minimize downstream estimation error by choosing sites that minimize Wasserstein distances between population and sample covariate distributions. I develop new theoretical upper bounds on PATE and CATE estimation errors, and show that these different objectives lead to different site selection strategies: both approaches form balanced, representative partitions of the support of the covariates, but use different penalties, which place a larger emphasis on site selections that place greater weight on minimizing downstream bias (PATE) and and variance (CATE) respectively. I extend this approach by using Wasserstein Distributionally Robust Optimization to guard against distribution shift when observed sites may not represent the target population, and develop a novel, data-driven procedure for uncertainty radius selection. I develop a cutting-plane algorithm that solves the resulting minimax problem by exploiting its sequential game structure, combined with a data-adaptive procedure for calibrating robustness parameters without requiring arbitrary assumptions about distributional uncertainty. Simulation evidence, and a reanalysis of a randomized microcredit experiment in Morocco, show that these methods outperform random and stratified sampling of sites, and alternative optimization methods i) for moderate-to-large size problem instances ii) when covariates are moderately informative about treatment effects, and iii) under induced distribution shift.The second paper, "Collusive and Adversarial Replication'', studies a game in which social ties between members of a research community may discourage prospective replicators from debunking papers that misreport results. Here, replication is an entrance decision, as a Replicator chooses whether or not to Replicate a given paper. A high level of social connectedness between members of a research community increases the field-wise False Discovery Rate, a measure of the social welfare associated with a healthy publication process. The moral is that larger, more diverse academic fields with fewer social ties are more likely to have an adversarial culture around replication, and that this improves social welfare. I consider three proposals to improve replication practices: Random auditing, or police-patrol replication; automated unit tests; and a recent proposal to lower the threshold for statistical significance. I argue that random auditing and automated unit tests can improve social welfare, but that the effect of lowering the statistical significance threshold is ambiguous.The third paper, "Why LLMs Hallucinate'', argues that LLMs hallucinate because their output is not constrained to be synonymous with claims for which they have evidence: a condition I call evidential closure. Information about the truth or falsity of sentences is not statistically identified in the standard neural language generation setup, and so cannot be conditioned on to generate new strings. We then show how to constrain LLMs to produce output that satisfies evidential closure. A multimodal LLM must learn about the external world (perceptual learning); it must learn a mapping from strings to states of the world (extensional learning); and, to achieve fluency when generalizing beyond a body of evidence, it must learn mappings from strings to their synonyms (intensional learning). The output of a unimodal LLM must be synonymous with strings in a validated evidence set. Finally, I present a heuristic procedure, Learn-Babble-Prune, that yields faithful output from an LLM by rejecting output that is not synonymous with claims for which the LLM has evidence.
■590 ▼aSchool code: 0028.
■650 4▼aStatistics
■650 4▼aPolitical science
■650 4▼aComputer science
■653 ▼aCausal inference
■653 ▼aSelective inference
■653 ▼aPerceptual learning
■690 ▼a0463
■690 ▼a0615
■690 ▼a0984
■71020▼aUniversity of California, Berkeley▼bPolitical Science.
■7730 ▼tDissertations Abstracts International▼g87-04B.
■790 ▼a0028
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17359187▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


