본문

서브메뉴

Enabling Responsible Data Science Through Multi-Dimensional Data Management
Enabling Responsible Data Science Through Multi-Dimensional Data Management
Enabling Responsible Data Science Through Multi-Dimensional Data Management

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103651
ISBN  
9798314875780
DDC  
004
저자명  
Lin, Yin.
서명/저자  
Enabling Responsible Data Science Through Multi-Dimensional Data Management
발행사항  
[Sl] : University of Michigan, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
158 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-11, Section: A.
주기사항  
Advisor: Jagadish, H. V.
학위논문주기  
Thesis (Ph.D.)--University of Michigan, 2025.
초록/해제  
요약In today's world, data is collected and utilized at an unprecedented scale, profoundly influencing society. Big data enables analyses that guide high stakes decisions and supports data-driven systems. While the benefits of big data are significant, the challenges extend beyond efficient processing and storage; as data scientists, we are responsible for ensuring that data applications ethically benefit society. Data science technologies can cause harm if they reinforce inequities, particularly when sensitive data-such as data linked to protected characteristics like race and gender is mishandled. This dissertation contributes to responsible data science by proposing comprehensive data management techniques to address challenges throughout the big data lifecycle. These approaches aim to enhance fairness, transparency, and accountability in data systems while considering the complexities associated with multiple protected characteristics. Firstly, data acquisition often results in the underrepresentation of certain populations, which risks perpetuating unfair treatment and oversight of these groups. Obtaining representative samples becomes particularly challenging when dealing with intersectional subgroups. We propose coverage analysis techniques to efficiently identify representation bias in multi-table databases, guiding data users toward obtaining more representative samples. Secondly, biases embedded in historical decisions can propagate into downstream machine learning tasks, resulting in unfair predictions. We emphasize the importance of addressing the root causes of unfairness in the training data. We propose model agnostic data pre-processing techniques to effectively detect and mitigate biased data collection, thereby enhancing ma- chine learning fairness across subgroups. Thirdly, data analytics based on cherry-pick generalizations can lead to misleading insights, diminishing the experiences of certain subgroups in decision-making. We refine these generalizations across multiple attributes to develop a framework evaluating their appropriateness, identify subgroup discrepancies, and promote more accurate and inclusive representations of data. Lastly, when a data analysis pipeline produces unexpected outputs, it is the responsibility of data scientists to interpret the potential sources of error. We propose a row-level data lineage approach to enhance pipeline transparency, enabling them to trace the origins of issues.
일반주제명  
Computer science
일반주제명  
Information science
키워드  
Responsible data science
키워드  
Data management
키워드  
Database
키워드  
AI fairness
기타저자  
University of Michigan Computer Science & Engineering
기본자료저록  
Dissertations Abstracts International. 86-11A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017358149
■00520260202103651
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798314875780
■035    ▼a(MiAaPQ)AAI32092697
■035    ▼a(MiAaPQ)umichrackham006062
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aLin,  Yin.
■24510▼aEnabling  Responsible  Data  Science  Through  Multi-Dimensional  Data  Management
■260    ▼a[Sl]▼bUniversity  of  Michigan▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a158  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-11,  Section:  A.
■500    ▼aAdvisor:  Jagadish,  H.  V.
■5021  ▼aThesis  (Ph.D.)--University  of  Michigan,  2025.
■520    ▼aIn  today's  world,  data  is  collected  and  utilized  at  an  unprecedented  scale,  profoundly  influencing  society.  Big  data  enables  analyses  that  guide  high  stakes  decisions  and  supports  data-driven  systems.  While  the  benefits  of  big  data  are  significant,  the  challenges  extend  beyond  efficient  processing  and  storage;  as  data  scientists,  we  are  responsible  for  ensuring  that  data  applications  ethically  benefit  society.  Data  science  technologies  can  cause  harm  if  they  reinforce  inequities,  particularly  when  sensitive  data-such  as  data  linked  to  protected  characteristics  like  race  and  gender  is  mishandled.  This  dissertation  contributes  to  responsible  data  science  by  proposing  comprehensive  data  management  techniques  to  address  challenges  throughout  the  big  data  lifecycle.  These  approaches  aim  to  enhance  fairness,  transparency,  and  accountability  in  data  systems  while  considering  the  complexities  associated  with  multiple  protected  characteristics.  Firstly,  data  acquisition  often  results  in  the  underrepresentation  of  certain  populations,  which  risks  perpetuating  unfair  treatment  and  oversight  of  these  groups.  Obtaining  representative  samples  becomes  particularly  challenging  when  dealing  with  intersectional  subgroups.  We  propose  coverage  analysis  techniques  to  efficiently  identify  representation  bias  in  multi-table  databases,  guiding  data  users  toward  obtaining  more  representative  samples.  Secondly,  biases  embedded  in  historical  decisions  can  propagate  into  downstream  machine  learning  tasks,  resulting  in  unfair  predictions.  We  emphasize  the  importance  of  addressing  the  root  causes  of  unfairness  in  the  training  data.  We  propose  model  agnostic  data  pre-processing  techniques  to  effectively  detect  and  mitigate  biased  data  collection,  thereby  enhancing  ma-  chine  learning  fairness  across  subgroups.  Thirdly,  data  analytics  based  on  cherry-pick  generalizations  can  lead  to  misleading  insights,  diminishing  the  experiences  of  certain  subgroups  in  decision-making.  We  refine  these  generalizations  across  multiple  attributes  to  develop  a  framework  evaluating  their  appropriateness,  identify  subgroup  discrepancies,  and  promote  more  accurate  and  inclusive  representations  of  data.  Lastly,  when  a  data  analysis  pipeline  produces  unexpected  outputs,  it  is  the  responsibility  of  data  scientists  to  interpret  the  potential  sources  of  error.  We  propose  a  row-level  data  lineage  approach  to  enhance  pipeline  transparency,  enabling  them  to  trace  the  origins  of  issues.
■590    ▼aSchool  code:  0127.
■650  4▼aComputer  science
■650  4▼aInformation  science
■653    ▼aResponsible  data  science
■653    ▼aData  management
■653    ▼aDatabase
■653    ▼aAI  fairness
■690    ▼a0984
■690    ▼a0723
■690    ▼a0800
■71020▼aUniversity  of  Michigan▼bComputer  Science  &  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-11A.
■790    ▼a0127
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358149▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF16187 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.