본문

서브메뉴

User-Centered Programmatic Data Labeling
User-Centered Programmatic Data Labeling
User-Centered Programmatic Data Labeling

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202105556
ISBN  
9798263399115
DDC  
000
저자명  
Wu, Renzhi.
서명/저자  
User-Centered Programmatic Data Labeling
발행사항  
[Sl] : Georgia Institute of Technology, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
168 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
주기사항  
Advisor: Chu, Xu.
학위논문주기  
Thesis (Ph.D.)--Georgia Institute of Technology, 2024.
초록/해제  
요약This dissertation addresses the critical challenge of labeled data scarcity in machine learning (ML), particularly in the context of deep learning, by advancing the paradigm of programmatic data labeling from a user-centered perspective. Traditional methods of obtaining labeled data through human annotation are costly and unscalable, prompting a shift towards programmatic data labeling, where noisy labels generated by various sources are utilized. Programmatic data labeling utilizes the Labeling Function (LF) abstraction, a small program that takes in a data point and outputs a weak label. Each supervision source is then expressed by a LF to automatically generate noisy labels from data points, which are then aggregated to infer ground-truth labels for training ML models.Programmatic data labeling invovles two major steps: LF development and LF aggregation with a label model. The current process of LF development relies on the expertise of the user and can be inaccessible for non-experts, particularly when dealing video data. LF aggregation through existing label models can also be non-trivial, requiring users to navigate challenging hyperparameter settings and unsupervised training procedures.To address these challenges, this dissertation makes three contributions. First, we explore how to improve usability by specializing programmatic data labeling to the task at hand. We use the task of entity matching as an example application to develop an Integrated Development Environment (IDE), facilitating the development, debugging, and management of LFs. The IDE also features a label model tailored for entity matching, SIMPLE-EM, which outperforms existing models in accuracy and efficiency by leveraging entity matching-specific properties.Second, we reformulate the task of writing LFs for video data as a video data retrieval task, so that users can develop LFs on video data by writing video retrieval queries. We then present SketchQL, a novel visual query interface for video data retrieval that allows users to construct retrieval queries through simple mouse drag-and-drop actions, improving usability greatly. This system demonstrates superior performance in video data retrieval compared to state-of-the-art methods.Third, we propose HyperLM, a hyper label model that eliminates the need for hyperparameter tuning and dataset-specific training, offering deterministic, accurate, and efficient label aggregation. We present the first ever analytical solution with optimalities for the task of label aggregation and design our hyper label model to approximate the analytical solution which is intractable to be directly used. Our hyper label model showcases significant improvements over existing methods in both accuracy and computational efficiency.
일반주제명  
Usability
일반주제명  
Deep learning
일반주제명  
Writing
일반주제명  
Debugging
일반주제명  
User feedback
일반주제명  
Labeling
일반주제명  
Information retrieval
일반주제명  
Computer science
기타저자  
Georgia Institute of Technology.
기본자료저록  
Dissertations Abstracts International. 87-05B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2024        us                              c    eng  d
■001000017360614
■00520260202105556
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798263399115
■035    ▼a(MiAaPQ)AAI32315883
■035    ▼a(MiAaPQ)GeorgiaTech75209
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a000
■1001  ▼aWu,  Renzhi.
■24510▼aUser-Centered  Programmatic  Data  Labeling
■260    ▼a[Sl]▼bGeorgia  Institute  of  Technology▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a168  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-05,  Section:  B.
■500    ▼aAdvisor:  Chu,  Xu.
■5021  ▼aThesis  (Ph.D.)--Georgia  Institute  of  Technology,  2024.
■520    ▼aThis  dissertation  addresses  the  critical  challenge  of  labeled  data  scarcity  in  machine  learning  (ML),  particularly  in  the  context  of  deep  learning,  by  advancing  the  paradigm  of  programmatic  data  labeling  from  a  user-centered  perspective.  Traditional  methods  of  obtaining  labeled  data  through  human  annotation  are  costly  and  unscalable,  prompting  a  shift  towards  programmatic  data  labeling,  where  noisy  labels  generated  by  various  sources  are  utilized.  Programmatic  data  labeling  utilizes  the  Labeling  Function  (LF)  abstraction,  a  small  program  that  takes  in  a  data  point  and  outputs  a  weak  label.  Each  supervision  source  is  then  expressed  by  a  LF  to  automatically  generate  noisy  labels  from  data  points,  which  are  then  aggregated  to  infer  ground-truth  labels  for  training  ML  models.Programmatic  data  labeling  invovles  two  major  steps:  LF  development  and  LF  aggregation  with  a  label  model.  The  current  process  of  LF  development  relies  on  the  expertise  of  the  user  and  can  be  inaccessible  for  non-experts,  particularly  when  dealing  video  data.  LF  aggregation  through  existing  label  models  can  also  be  non-trivial,  requiring  users  to  navigate  challenging  hyperparameter  settings  and  unsupervised  training  procedures.To  address  these  challenges,  this  dissertation  makes  three  contributions.  First,  we  explore  how  to  improve  usability  by  specializing  programmatic  data  labeling  to  the  task  at  hand.  We  use  the  task  of  entity  matching  as  an  example  application  to  develop  an  Integrated  Development  Environment  (IDE),  facilitating  the  development,  debugging,  and  management  of  LFs.  The  IDE  also  features  a  label  model  tailored  for  entity  matching,  SIMPLE-EM,  which  outperforms  existing  models  in  accuracy  and  efficiency  by  leveraging  entity  matching-specific  properties.Second,  we  reformulate  the  task  of  writing  LFs  for  video  data  as  a  video  data  retrieval  task,  so  that  users  can  develop  LFs  on  video  data  by  writing  video  retrieval  queries.  We  then  present  SketchQL,  a  novel  visual  query  interface  for  video  data  retrieval  that  allows  users  to  construct  retrieval  queries  through  simple  mouse  drag-and-drop  actions,  improving  usability  greatly.  This  system  demonstrates  superior  performance  in  video  data  retrieval  compared  to  state-of-the-art  methods.Third,  we  propose  HyperLM,  a  hyper  label  model  that  eliminates  the  need  for  hyperparameter  tuning  and  dataset-specific  training,  offering  deterministic,  accurate,  and  efficient  label  aggregation.  We  present  the  first  ever  analytical  solution  with  optimalities  for  the  task  of  label  aggregation  and  design  our  hyper  label  model  to  approximate  the  analytical  solution  which  is  intractable  to  be  directly  used.  Our  hyper  label  model  showcases  significant  improvements  over  existing  methods  in  both  accuracy  and  computational  efficiency.
■590    ▼aSchool  code:  0078.
■650  4▼aUsability
■650  4▼aDeep  learning
■650  4▼aWriting
■650  4▼aDebugging
■650  4▼aUser  feedback
■650  4▼aLabeling
■650  4▼aInformation  retrieval
■650  4▼aComputer  science
■690    ▼a0800
■690    ▼a0984
■71020▼aGeorgia  Institute  of  Technology.
■7730  ▼tDissertations  Abstracts  International▼g87-05B.
■790    ▼a0078
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360614▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF15130 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.