서브메뉴
검색
Adapting Pre-trained Models and Leveraging Targeted Multilinguality for Under-Resourced and Endangered Language Processing
Adapting Pre-trained Models and Leveraging Targeted Multilinguality for Under-Resourced and Endangered Language Processing
Detailed Information
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211152002
- ISBN
- 9798383223857
- DDC
- 401
- 저자명
- Downey, C. M.
- 서명/저자
- Adapting Pre-trained Models and Leveraging Targeted Multilinguality for Under-Resourced and Endangered Language Processing
- 발행사항
- [Sl] : University of Washington, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 132 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 86-01, Section: B.
- 주기사항
- Advisor: Levow, Gina-Anne;Steinert-Threlkeld, Shane.
- 학위논문주기
- Thesis (Ph.D.)--University of Washington, 2024.
- 초록/해제
- 요약Advances in Natural Language Processing (NLP) over the past decade have largely been driven by the scale of data and computation used to train large neural network-based models. However, these techniques are inapplicable to the vast majority of the world's languages, which lack the vast digitized text datasets available for English and a few other very high-resource languages. In this dissertation, we present three case studies for extending NLP applications to under-resourced languages. These case studies include conducting unsupervised morphological segmentation for extremely low-resource languages via multilingual training and transfer, optimizing the vocabulary of a pre-trained cross-lingual model for specific target language(s), and specializing a pre-trained model for a low-resource language family (Uralic). Based on these case studies, we argue for three broad, guiding principles in extending NLP applications to under-resourced languages. First: where possible, robustly pre-trained models and representations should be leveraged. Second: components of pre-trained models that are not optimized for new languages should be substituted or substantially adapted. Third: targeted multilingual training provides a middle ground between the lack of adequate data to train models for individual under-resourced languages on one hand, and the diminishing returns of "massively multilingual" training on the other.
- 일반주제명
- Linguistics
- 일반주제명
- Computer science
- 일반주제명
- Language
- 키워드
- Uralic
- 키워드
- Multilinguality
- 키워드
- Vocabulary
- 기타저자
- University of Washington Linguistics
- 기본자료저록
- Dissertations Abstracts International. 86-01B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017162352
■00520250211152002
■006m o d
■007cr#unu||||||||
■020 ▼a9798383223857
■035 ▼a(MiAaPQ)AAI31330087
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a401
■1001 ▼aDowney, C. M.
■24510▼aAdapting Pre-trained Models and Leveraging Targeted Multilinguality for Under-Resourced and Endangered Language Processing
■260 ▼a[Sl]▼bUniversity of Washington▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a132 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 86-01, Section: B.
■500 ▼aAdvisor: Levow, Gina-Anne;Steinert-Threlkeld, Shane.
■5021 ▼aThesis (Ph.D.)--University of Washington, 2024.
■520 ▼aAdvances in Natural Language Processing (NLP) over the past decade have largely been driven by the scale of data and computation used to train large neural network-based models. However, these techniques are inapplicable to the vast majority of the world's languages, which lack the vast digitized text datasets available for English and a few other very high-resource languages. In this dissertation, we present three case studies for extending NLP applications to under-resourced languages. These case studies include conducting unsupervised morphological segmentation for extremely low-resource languages via multilingual training and transfer, optimizing the vocabulary of a pre-trained cross-lingual model for specific target language(s), and specializing a pre-trained model for a low-resource language family (Uralic). Based on these case studies, we argue for three broad, guiding principles in extending NLP applications to under-resourced languages. First: where possible, robustly pre-trained models and representations should be leveraged. Second: components of pre-trained models that are not optimized for new languages should be substituted or substantially adapted. Third: targeted multilingual training provides a middle ground between the lack of adequate data to train models for individual under-resourced languages on one hand, and the diminishing returns of "massively multilingual" training on the other.
■590 ▼aSchool code: 0250.
■650 4▼aLinguistics
■650 4▼aComputer science
■650 4▼aLanguage
■653 ▼aNatural Language Processing
■653 ▼aUralic
■653 ▼aMultilinguality
■653 ▼aVocabulary
■690 ▼a0290
■690 ▼a0984
■690 ▼a0679
■71020▼aUniversity of Washington▼bLinguistics.
■7730 ▼tDissertations Abstracts International▼g86-01B.
■790 ▼a0250
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162352▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
Preview
Export
ChatGPT Discussion
AI Recommended Related Books
Подробнее информация.
- Бронирование
- не существует
- моя папка
- Первый запрос зрения
- Non-Book Loan Application
- Nighttime Book Loan Application
Available after logging in.


