서브메뉴
검색
Computational Methods for Organizational Health Literacy: Risks, Opportunities, and Future Directions
Computational Methods for Organizational Health Literacy: Risks, Opportunities, and Future Directions
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103545
- ISBN
- 9798280717107
- DDC
- 614
- 서명/저자
- Computational Methods for Organizational Health Literacy: Risks, Opportunities, and Future Directions
- 발행사항
- [Sl] : Harvard University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 137 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
- 주기사항
- Advisor: Emmons, Karen M.
- 학위논문주기
- Thesis (Ph.D.)--Harvard University, 2025.
- 초록/해제
- 요약IntroductionHealth literacy emerged as a field of research and practice in the late 20th century, initially emphasizing individual capacity to obtain, process, and understand health information. Over time, the field expanded to focus on organizations' capacity to provide clear, accessible health information, i.e. organizational health literacy. This shift in focus led to the development of tools to support organizational health literacy, such as the CDC Clear Communication Index. By helping organizations reduce the burden their communication materials place on target audiences, tools like the CDC Clear Communication Index can play a vital role in addressing today's pressing public health issues of mistrust and misinformation. However, there are significant gaps in knowledge about how these tools have been used so far and how they might interact with computational methods better suited for high-volume communication online than traditional manual scoring practices.This dissertation focused on the CDC Clear Communication Index, describing its use in research and experimenting with computational methods to apply it at scale. We chose the CDC Clear Communication Index due to its potential for impact at scale online. It is composed of 20 binary items, resulting in a percentage score. It is structured to suit a wide variety of media formats, message topics, and message lengths. It encompasses a holistic set of evidence-based communication guidelines, covering core message components, behavioral recommendations, use of numbers, and discussions of risk.This dissertation focuses on two key types of computational methods often referred to as "artificial intelligence" (AI) applications. The first is supervised machine learning, in which a particular training algorithm is used to identify patterns in labeled data useful for predicting labels on new data. The second is generative AI, leveraging large language models to generate content based on text inputs.MethodsThis dissertation includes three separate studies examining the potential of the CDC Clear Communication Index through distinct methods.Chapter 1 uses a scoping review methodology to describe the use of the Index as an assessment tool in descriptive research. Primary data analysis focused on study design. Secondary data analysis focused on reporting of results and methods.Chapter 2 introduces the 4-Factor Framework to Assess the Suitability of AI in Health Communication. We demonstrate its utility through a hypothetical use case: a US federal health agency assessing the suitability of open-source academic machine learning models to rate public health social media posts according to guidelines from the CDC Clear Communication Index. We trained 8 bag-of-words models and fine-tuned 8 BERT models to predict expert raters' labels of social media posts, based on a training dataset of US state health agencies' pandemic social media posts. We used qualitative process document review and quantitative analysis of the training data to assess models' explainability. We used qualitative process document review and qualitative analysis of model outputs to assess flexibility. We used a mix of quantitative metrics to describe models' performance in accurately predicting the labels that trained human raters assigned to social media posts. We used concise comparative summaries to identify words that differentiated social media posts that received the most incorrect model predictions from those that models performed perfectly on.Chapter 3 examines the performance of ChatGPT in applying binary items from the CDC Clear Communication Index to social media posts. We compare the results of prompt engineering between front-end user and back-end developer perspectives. To do so, we used 12 different prompting styles, varying in framing and length, to apply binary Index items to a test dataset of 27 social media posts. Using F1 and MCC as performance metrics, we compared each prompting style's performance in accurately predicting labels from expert raters. We used these metrics to identify the optimal prompting style for each item and the overall best performing style across all items, from each stakeholder perspective. We then compared the performance of different model versions of ChatGPT on the same tasks using a validation dataset of 260 social media posts, using the optimal prompt styles we identified during prompt engineering from the back-end developer perspective.ResultsOur scoping review in Chapter 1 identified a wide breadth of research contexts in which the CDC Clear Communication Index has been applied. However, we also uncovered major gaps in study design and reporting. Despite largely employing purposive samples, studies using the CDC Clear Communication Index focused on quantitative assessments, making interpretation of results difficult. Despite this quantitative focus, studies often lacked key details in reporting, such as mean and median Index scores. We also found a prominent focus on materials in text formats, from government, academic, and nonprofit authors. These results point to the potential of the Index, demonstrating its reach thus far, as well as the need for researchers to expand its impact through more varied study design and more robust reporting.Our analysis in Chapter 2 revealed significant tradeoffs between different facets of suitability that highlight the role institutional priorities play in implementation, especially in light of the potential for unfair model performance across social contexts. We found a tradeoff between model explainability and fairness. Further, while fine-tuned BERT models outperformed bag-of-words models, both demonstrated potential bias with social media posts containing references to time-sensitive local knowledge. In our hypothetical use case, we would recommend limited implementation of such models, supported through continuous evaluation. This study also demonstrates the transparency and documentation necessary for assessing AI suitability in health communication, with implications for implementation policies that should restrict the use of proprietary, closed-source tools.Our analysis in Chapter 3 revealed several long-term complications that would stem from the use of off-the-shelf large language models as health communication tools. We found that back-end developers and front-end users would reach different conclusions about how to best prompt ChatGPT-3.5 to accurately assess posts using items from the CDC Clear Communication Index. This highlights the need for participatory evaluation methods to shape generative AI implementation and policy. In our analysis of performance across model versions, we found inconsistent performance, highlighting the need for continual evaluation of generative AI tools. Our results highlight the need for investment in health communication infrastructure, even in the face of generative AI tools like ChatGPT.ConclusionThough these studies focused on just one organizational health literacy tool, these conclusions are likely transferable to other aspects of health communication. Multiple assessment tools draw on the same kinds of evidence-based practices as the CDC Clear Communication Index, which are not defined in terms of readily quantifiable variables. Further, application of AI in content generation, text summary, and simulated conversations would all require evaluation to ensure acceptable performance across various contexts over time.
- 일반주제명
- Public health
- 일반주제명
- Public administration
- 키워드
- Health literacy
- 키워드
- Social media
- 기타저자
- Harvard University Population Health Sciences
- 기본자료저록
- Dissertations Abstracts International. 86-12B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017357676
■00520260202103545
■006m o d
■007cr#unu||||||||
■020 ▼a9798280717107
■035 ▼a(MiAaPQ)AAI32041233
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a614
■1001 ▼aMendez, Samuel R.▼0(orcid)0000-0003-4402-1885
■24510▼aComputational Methods for Organizational Health Literacy: Risks, Opportunities, and Future Directions
■260 ▼a[Sl]▼bHarvard University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a137 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-12, Section: B.
■500 ▼aAdvisor: Emmons, Karen M.
■5021 ▼aThesis (Ph.D.)--Harvard University, 2025.
■520 ▼aIntroductionHealth literacy emerged as a field of research and practice in the late 20th century, initially emphasizing individual capacity to obtain, process, and understand health information. Over time, the field expanded to focus on organizations' capacity to provide clear, accessible health information, i.e. organizational health literacy. This shift in focus led to the development of tools to support organizational health literacy, such as the CDC Clear Communication Index. By helping organizations reduce the burden their communication materials place on target audiences, tools like the CDC Clear Communication Index can play a vital role in addressing today's pressing public health issues of mistrust and misinformation. However, there are significant gaps in knowledge about how these tools have been used so far and how they might interact with computational methods better suited for high-volume communication online than traditional manual scoring practices.This dissertation focused on the CDC Clear Communication Index, describing its use in research and experimenting with computational methods to apply it at scale. We chose the CDC Clear Communication Index due to its potential for impact at scale online. It is composed of 20 binary items, resulting in a percentage score. It is structured to suit a wide variety of media formats, message topics, and message lengths. It encompasses a holistic set of evidence-based communication guidelines, covering core message components, behavioral recommendations, use of numbers, and discussions of risk.This dissertation focuses on two key types of computational methods often referred to as "artificial intelligence" (AI) applications. The first is supervised machine learning, in which a particular training algorithm is used to identify patterns in labeled data useful for predicting labels on new data. The second is generative AI, leveraging large language models to generate content based on text inputs.MethodsThis dissertation includes three separate studies examining the potential of the CDC Clear Communication Index through distinct methods.Chapter 1 uses a scoping review methodology to describe the use of the Index as an assessment tool in descriptive research. Primary data analysis focused on study design. Secondary data analysis focused on reporting of results and methods.Chapter 2 introduces the 4-Factor Framework to Assess the Suitability of AI in Health Communication. We demonstrate its utility through a hypothetical use case: a US federal health agency assessing the suitability of open-source academic machine learning models to rate public health social media posts according to guidelines from the CDC Clear Communication Index. We trained 8 bag-of-words models and fine-tuned 8 BERT models to predict expert raters' labels of social media posts, based on a training dataset of US state health agencies' pandemic social media posts. We used qualitative process document review and quantitative analysis of the training data to assess models' explainability. We used qualitative process document review and qualitative analysis of model outputs to assess flexibility. We used a mix of quantitative metrics to describe models' performance in accurately predicting the labels that trained human raters assigned to social media posts. We used concise comparative summaries to identify words that differentiated social media posts that received the most incorrect model predictions from those that models performed perfectly on.Chapter 3 examines the performance of ChatGPT in applying binary items from the CDC Clear Communication Index to social media posts. We compare the results of prompt engineering between front-end user and back-end developer perspectives. To do so, we used 12 different prompting styles, varying in framing and length, to apply binary Index items to a test dataset of 27 social media posts. Using F1 and MCC as performance metrics, we compared each prompting style's performance in accurately predicting labels from expert raters. We used these metrics to identify the optimal prompting style for each item and the overall best performing style across all items, from each stakeholder perspective. We then compared the performance of different model versions of ChatGPT on the same tasks using a validation dataset of 260 social media posts, using the optimal prompt styles we identified during prompt engineering from the back-end developer perspective.ResultsOur scoping review in Chapter 1 identified a wide breadth of research contexts in which the CDC Clear Communication Index has been applied. However, we also uncovered major gaps in study design and reporting. Despite largely employing purposive samples, studies using the CDC Clear Communication Index focused on quantitative assessments, making interpretation of results difficult. Despite this quantitative focus, studies often lacked key details in reporting, such as mean and median Index scores. We also found a prominent focus on materials in text formats, from government, academic, and nonprofit authors. These results point to the potential of the Index, demonstrating its reach thus far, as well as the need for researchers to expand its impact through more varied study design and more robust reporting.Our analysis in Chapter 2 revealed significant tradeoffs between different facets of suitability that highlight the role institutional priorities play in implementation, especially in light of the potential for unfair model performance across social contexts. We found a tradeoff between model explainability and fairness. Further, while fine-tuned BERT models outperformed bag-of-words models, both demonstrated potential bias with social media posts containing references to time-sensitive local knowledge. In our hypothetical use case, we would recommend limited implementation of such models, supported through continuous evaluation. This study also demonstrates the transparency and documentation necessary for assessing AI suitability in health communication, with implications for implementation policies that should restrict the use of proprietary, closed-source tools.Our analysis in Chapter 3 revealed several long-term complications that would stem from the use of off-the-shelf large language models as health communication tools. We found that back-end developers and front-end users would reach different conclusions about how to best prompt ChatGPT-3.5 to accurately assess posts using items from the CDC Clear Communication Index. This highlights the need for participatory evaluation methods to shape generative AI implementation and policy. In our analysis of performance across model versions, we found inconsistent performance, highlighting the need for continual evaluation of generative AI tools. Our results highlight the need for investment in health communication infrastructure, even in the face of generative AI tools like ChatGPT.ConclusionThough these studies focused on just one organizational health literacy tool, these conclusions are likely transferable to other aspects of health communication. Multiple assessment tools draw on the same kinds of evidence-based practices as the CDC Clear Communication Index, which are not defined in terms of readily quantifiable variables. Further, application of AI in content generation, text summary, and simulated conversations would all require evaluation to ensure acceptable performance across various contexts over time.
■590 ▼aSchool code: 0084.
■650 4▼aPublic health
■650 4▼aPublic administration
■653 ▼aHealth communication
■653 ▼aHealth literacy
■653 ▼aNatural language processing
■653 ▼aSocial media
■653 ▼aHealth information
■690 ▼a0573
■690 ▼a0617
■690 ▼a0800
■71020▼aHarvard University▼bPopulation Health Sciences.
■7730 ▼tDissertations Abstracts International▼g86-12B.
■790 ▼a0084
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357676▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


