서브메뉴
검색
Enabling Responsible Data Science Through Multi-Dimensional Data Management
Enabling Responsible Data Science Through Multi-Dimensional Data Management
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202103651
- ISBN
- 9798314875780
- DDC
- 004
- 저자명
- Lin, Yin.
- 서명/저자
- Enabling Responsible Data Science Through Multi-Dimensional Data Management
- 발행사항
- [Sl] : University of Michigan, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 158 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-11, Section: A.
- 주기사항
- Advisor: Jagadish, H. V.
- 학위논문주기
- Thesis (Ph.D.)--University of Michigan, 2025.
- 초록/해제
- 요약In today's world, data is collected and utilized at an unprecedented scale, profoundly influencing society. Big data enables analyses that guide high stakes decisions and supports data-driven systems. While the benefits of big data are significant, the challenges extend beyond efficient processing and storage; as data scientists, we are responsible for ensuring that data applications ethically benefit society. Data science technologies can cause harm if they reinforce inequities, particularly when sensitive data-such as data linked to protected characteristics like race and gender is mishandled. This dissertation contributes to responsible data science by proposing comprehensive data management techniques to address challenges throughout the big data lifecycle. These approaches aim to enhance fairness, transparency, and accountability in data systems while considering the complexities associated with multiple protected characteristics. Firstly, data acquisition often results in the underrepresentation of certain populations, which risks perpetuating unfair treatment and oversight of these groups. Obtaining representative samples becomes particularly challenging when dealing with intersectional subgroups. We propose coverage analysis techniques to efficiently identify representation bias in multi-table databases, guiding data users toward obtaining more representative samples. Secondly, biases embedded in historical decisions can propagate into downstream machine learning tasks, resulting in unfair predictions. We emphasize the importance of addressing the root causes of unfairness in the training data. We propose model agnostic data pre-processing techniques to effectively detect and mitigate biased data collection, thereby enhancing ma- chine learning fairness across subgroups. Thirdly, data analytics based on cherry-pick generalizations can lead to misleading insights, diminishing the experiences of certain subgroups in decision-making. We refine these generalizations across multiple attributes to develop a framework evaluating their appropriateness, identify subgroup discrepancies, and promote more accurate and inclusive representations of data. Lastly, when a data analysis pipeline produces unexpected outputs, it is the responsibility of data scientists to interpret the potential sources of error. We propose a row-level data lineage approach to enhance pipeline transparency, enabling them to trace the origins of issues.
- 일반주제명
- Computer science
- 일반주제명
- Information science
- 키워드
- Data management
- 키워드
- Database
- 키워드
- AI fairness
- 기타저자
- University of Michigan Computer Science & Engineering
- 기본자료저록
- Dissertations Abstracts International. 86-11A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358149
■00520260202103651
■006m o d
■007cr#unu||||||||
■020 ▼a9798314875780
■035 ▼a(MiAaPQ)AAI32092697
■035 ▼a(MiAaPQ)umichrackham006062
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aLin, Yin.
■24510▼aEnabling Responsible Data Science Through Multi-Dimensional Data Management
■260 ▼a[Sl]▼bUniversity of Michigan▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a158 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-11, Section: A.
■500 ▼aAdvisor: Jagadish, H. V.
■5021 ▼aThesis (Ph.D.)--University of Michigan, 2025.
■520 ▼aIn today's world, data is collected and utilized at an unprecedented scale, profoundly influencing society. Big data enables analyses that guide high stakes decisions and supports data-driven systems. While the benefits of big data are significant, the challenges extend beyond efficient processing and storage; as data scientists, we are responsible for ensuring that data applications ethically benefit society. Data science technologies can cause harm if they reinforce inequities, particularly when sensitive data-such as data linked to protected characteristics like race and gender is mishandled. This dissertation contributes to responsible data science by proposing comprehensive data management techniques to address challenges throughout the big data lifecycle. These approaches aim to enhance fairness, transparency, and accountability in data systems while considering the complexities associated with multiple protected characteristics. Firstly, data acquisition often results in the underrepresentation of certain populations, which risks perpetuating unfair treatment and oversight of these groups. Obtaining representative samples becomes particularly challenging when dealing with intersectional subgroups. We propose coverage analysis techniques to efficiently identify representation bias in multi-table databases, guiding data users toward obtaining more representative samples. Secondly, biases embedded in historical decisions can propagate into downstream machine learning tasks, resulting in unfair predictions. We emphasize the importance of addressing the root causes of unfairness in the training data. We propose model agnostic data pre-processing techniques to effectively detect and mitigate biased data collection, thereby enhancing ma- chine learning fairness across subgroups. Thirdly, data analytics based on cherry-pick generalizations can lead to misleading insights, diminishing the experiences of certain subgroups in decision-making. We refine these generalizations across multiple attributes to develop a framework evaluating their appropriateness, identify subgroup discrepancies, and promote more accurate and inclusive representations of data. Lastly, when a data analysis pipeline produces unexpected outputs, it is the responsibility of data scientists to interpret the potential sources of error. We propose a row-level data lineage approach to enhance pipeline transparency, enabling them to trace the origins of issues.
■590 ▼aSchool code: 0127.
■650 4▼aComputer science
■650 4▼aInformation science
■653 ▼aResponsible data science
■653 ▼aData management
■653 ▼aDatabase
■653 ▼aAI fairness
■690 ▼a0984
■690 ▼a0723
■690 ▼a0800
■71020▼aUniversity of Michigan▼bComputer Science & Engineering.
■7730 ▼tDissertations Abstracts International▼g86-11A.
■790 ▼a0127
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358149▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Подробнее информация.
- Бронирование
- не существует
- моя папка
- Первый запрос зрения
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


