서브메뉴
검색
User-Centered Programmatic Data Labeling
User-Centered Programmatic Data Labeling
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105556
- ISBN
- 9798263399115
- DDC
- 000
- 저자명
- Wu, Renzhi.
- 서명/저자
- User-Centered Programmatic Data Labeling
- 발행사항
- [Sl] : Georgia Institute of Technology, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 168 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-05, Section: B.
- 주기사항
- Advisor: Chu, Xu.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2024.
- 초록/해제
- 요약This dissertation addresses the critical challenge of labeled data scarcity in machine learning (ML), particularly in the context of deep learning, by advancing the paradigm of programmatic data labeling from a user-centered perspective. Traditional methods of obtaining labeled data through human annotation are costly and unscalable, prompting a shift towards programmatic data labeling, where noisy labels generated by various sources are utilized. Programmatic data labeling utilizes the Labeling Function (LF) abstraction, a small program that takes in a data point and outputs a weak label. Each supervision source is then expressed by a LF to automatically generate noisy labels from data points, which are then aggregated to infer ground-truth labels for training ML models.Programmatic data labeling invovles two major steps: LF development and LF aggregation with a label model. The current process of LF development relies on the expertise of the user and can be inaccessible for non-experts, particularly when dealing video data. LF aggregation through existing label models can also be non-trivial, requiring users to navigate challenging hyperparameter settings and unsupervised training procedures.To address these challenges, this dissertation makes three contributions. First, we explore how to improve usability by specializing programmatic data labeling to the task at hand. We use the task of entity matching as an example application to develop an Integrated Development Environment (IDE), facilitating the development, debugging, and management of LFs. The IDE also features a label model tailored for entity matching, SIMPLE-EM, which outperforms existing models in accuracy and efficiency by leveraging entity matching-specific properties.Second, we reformulate the task of writing LFs for video data as a video data retrieval task, so that users can develop LFs on video data by writing video retrieval queries. We then present SketchQL, a novel visual query interface for video data retrieval that allows users to construct retrieval queries through simple mouse drag-and-drop actions, improving usability greatly. This system demonstrates superior performance in video data retrieval compared to state-of-the-art methods.Third, we propose HyperLM, a hyper label model that eliminates the need for hyperparameter tuning and dataset-specific training, offering deterministic, accurate, and efficient label aggregation. We present the first ever analytical solution with optimalities for the task of label aggregation and design our hyper label model to approximate the analytical solution which is intractable to be directly used. Our hyper label model showcases significant improvements over existing methods in both accuracy and computational efficiency.
- 일반주제명
- Usability
- 일반주제명
- Deep learning
- 일반주제명
- Writing
- 일반주제명
- Debugging
- 일반주제명
- User feedback
- 일반주제명
- Labeling
- 일반주제명
- Information retrieval
- 일반주제명
- Computer science
- 기본자료저록
- Dissertations Abstracts International. 87-05B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2024 us c eng d■001000017360614
■00520260202105556
■006m o d
■007cr#unu||||||||
■020 ▼a9798263399115
■035 ▼a(MiAaPQ)AAI32315883
■035 ▼a(MiAaPQ)GeorgiaTech75209
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a000
■1001 ▼aWu, Renzhi.
■24510▼aUser-Centered Programmatic Data Labeling
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a168 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-05, Section: B.
■500 ▼aAdvisor: Chu, Xu.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2024.
■520 ▼aThis dissertation addresses the critical challenge of labeled data scarcity in machine learning (ML), particularly in the context of deep learning, by advancing the paradigm of programmatic data labeling from a user-centered perspective. Traditional methods of obtaining labeled data through human annotation are costly and unscalable, prompting a shift towards programmatic data labeling, where noisy labels generated by various sources are utilized. Programmatic data labeling utilizes the Labeling Function (LF) abstraction, a small program that takes in a data point and outputs a weak label. Each supervision source is then expressed by a LF to automatically generate noisy labels from data points, which are then aggregated to infer ground-truth labels for training ML models.Programmatic data labeling invovles two major steps: LF development and LF aggregation with a label model. The current process of LF development relies on the expertise of the user and can be inaccessible for non-experts, particularly when dealing video data. LF aggregation through existing label models can also be non-trivial, requiring users to navigate challenging hyperparameter settings and unsupervised training procedures.To address these challenges, this dissertation makes three contributions. First, we explore how to improve usability by specializing programmatic data labeling to the task at hand. We use the task of entity matching as an example application to develop an Integrated Development Environment (IDE), facilitating the development, debugging, and management of LFs. The IDE also features a label model tailored for entity matching, SIMPLE-EM, which outperforms existing models in accuracy and efficiency by leveraging entity matching-specific properties.Second, we reformulate the task of writing LFs for video data as a video data retrieval task, so that users can develop LFs on video data by writing video retrieval queries. We then present SketchQL, a novel visual query interface for video data retrieval that allows users to construct retrieval queries through simple mouse drag-and-drop actions, improving usability greatly. This system demonstrates superior performance in video data retrieval compared to state-of-the-art methods.Third, we propose HyperLM, a hyper label model that eliminates the need for hyperparameter tuning and dataset-specific training, offering deterministic, accurate, and efficient label aggregation. We present the first ever analytical solution with optimalities for the task of label aggregation and design our hyper label model to approximate the analytical solution which is intractable to be directly used. Our hyper label model showcases significant improvements over existing methods in both accuracy and computational efficiency.
■590 ▼aSchool code: 0078.
■650 4▼aUsability
■650 4▼aDeep learning
■650 4▼aWriting
■650 4▼aDebugging
■650 4▼aUser feedback
■650 4▼aLabeling
■650 4▼aInformation retrieval
■650 4▼aComputer science
■690 ▼a0800
■690 ▼a0984
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-05B.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360614▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


