서브메뉴
검색
Understanding the Effects of Increased Transparency on Data Preprocessing Through In-Process Visualizations
Understanding the Effects of Increased Transparency on Data Preprocessing Through In-Process Visualizations
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104726
- ISBN
- 9798291555507
- DDC
- 020
- 저자명
- Su, William.
- 서명/저자
- Understanding the Effects of Increased Transparency on Data Preprocessing Through In-Process Visualizations
- 발행사항
- [Sl] : The University of North Carolina at Chapel Hill, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 114 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
- 주기사항
- Advisor: Wang, Yue;Gotz, David.
- 학위논문주기
- Thesis (Ph.D.)--The University of North Carolina at Chapel Hill, 2025.
- 초록/해제
- 요약Most work on evaluating bias in data science workflows tends to focus on the model. However, the training data fed into the model and the data preprocessing step that produces it can also have significant impact on model results. While there has been work on editing the data in data preprocessing to mitigate bias, the impact of conventional data preprocessing operations has been understudied. My dissertation delves into how the data preprocessing step can be improved to help analysts better understand the impact of the step and lead to smarter data science decisions. I first study the needs of data scientists when conducting data preprocessing through a small-scale interview study and compared the results with a literature survey of current preprocessing tools. The comparison analysis identified several key gaps between practice and theory. I utilized of result of the analysis to develop the Preprocess Analyzer (PPA) tool, which is designed to address some of the gaps by being integrated into existing data science work environments and provided users with a deeper insight into their data. I conducted a user study to evaluate the ability of PPA to aid with data preprocessing. The study results found that compared to existing popular tools, data scientists gained a better understanding of their data preprocessing workflow when utilizing PPA. Participants generally agreed that PPA included many helpful features such as the ability to quickly display useful statistics, highlight areas of concern, and integration into familiar work environments. I believe the results of this dissertation can guide the design of future data preprocessing tools to better meet the needs of the end user.
- 일반주제명
- Information science
- 일반주제명
- Computer science
- 일반주제명
- Library science
- 기타저자
- The University of North Carolina at Chapel Hill Information and Library Science
- 기본자료저록
- Dissertations Abstracts International. 87-02B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358611
■00520260202104726
■006m o d
■007cr#unu||||||||
■020 ▼a9798291555507
■035 ▼a(MiAaPQ)AAI32122278
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a020
■1001 ▼aSu, William.
■24510▼aUnderstanding the Effects of Increased Transparency on Data Preprocessing Through In-Process Visualizations
■260 ▼a[Sl]▼bThe University of North Carolina at Chapel Hill▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a114 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-02, Section: B.
■500 ▼aAdvisor: Wang, Yue;Gotz, David.
■5021 ▼aThesis (Ph.D.)--The University of North Carolina at Chapel Hill, 2025.
■520 ▼aMost work on evaluating bias in data science workflows tends to focus on the model. However, the training data fed into the model and the data preprocessing step that produces it can also have significant impact on model results. While there has been work on editing the data in data preprocessing to mitigate bias, the impact of conventional data preprocessing operations has been understudied. My dissertation delves into how the data preprocessing step can be improved to help analysts better understand the impact of the step and lead to smarter data science decisions. I first study the needs of data scientists when conducting data preprocessing through a small-scale interview study and compared the results with a literature survey of current preprocessing tools. The comparison analysis identified several key gaps between practice and theory. I utilized of result of the analysis to develop the Preprocess Analyzer (PPA) tool, which is designed to address some of the gaps by being integrated into existing data science work environments and provided users with a deeper insight into their data. I conducted a user study to evaluate the ability of PPA to aid with data preprocessing. The study results found that compared to existing popular tools, data scientists gained a better understanding of their data preprocessing workflow when utilizing PPA. Participants generally agreed that PPA included many helpful features such as the ability to quickly display useful statistics, highlight areas of concern, and integration into familiar work environments. I believe the results of this dissertation can guide the design of future data preprocessing tools to better meet the needs of the end user.
■590 ▼aSchool code: 0153.
■650 4▼aInformation science
■650 4▼aComputer science
■650 4▼aLibrary science
■653 ▼aData preprocessing operations
■653 ▼aPreprocess Analyzer tool
■653 ▼aData science decisions
■690 ▼a0723
■690 ▼a0984
■690 ▼a0399
■71020▼aThe University of North Carolina at Chapel Hill▼bInformation and Library Science.
■7730 ▼tDissertations Abstracts International▼g87-02B.
■790 ▼a0153
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358611▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


