서브메뉴
검색
Responsible AI via Responsible Large Language Models- [electronic resource]
Responsible AI via Responsible Large Language Models- [electronic resource]
상세정보
- 자료유형
- 학위논문파일 국외
- 최종처리일시
- 20240214101253
- ISBN
- 9798380154154
- DDC
- 621.3
- 서명/저자
- Responsible AI via Responsible Large Language Models - [electronic resource]
- 발행사항
- [S.l.]: : University of California, Santa Barbara., 2023
- 발행사항
- Ann Arbor : : ProQuest Dissertations & Theses,, 2023
- 형태사항
- 1 online resource(142 p.)
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-02, Section: B.
- 주기사항
- Advisor: Wang, William Yang.
- 학위논문주기
- Thesis (Ph.D.)--University of California, Santa Barbara, 2023.
- 사용제한주기
- This item must not be sold to any third party vendors.
- 초록/해제
- 요약Large language models have advanced the state-of-the-art in natural language processing and achieved success in tasks such as summarization, question answering, and text classification. However, these models are trained on large-scale datasets, which may include harmful information. Studies have shown that as a result, the models can exhibit social biases and generate misinformation after training. This dissertation discusses research on analyzing and interpreting the risks of large language models across the areas of fairness, trustworthiness, and safety. The first part of this dissertation analyzes issues of fairness related to social biases in large language models. We first investigate issues of dialect bias pertaining to African American English and Standard American English within the context of text generation. We also analyze a more complex setting of fairness: cases in which multiple attributes affect each other to form compound biases. This is studied in relation to gender and seniority attributes.The second part focuses on trustworthiness and the spread of misinformation across different scopes: prevention, detection, and memorization. We describe an open-domain question-answering system for emergent domains that uses various retrieval and re-ranking techniques to provide users with information from trustworthy sources. This is demonstrated in the context of the emergent COVID-19 pandemic. We further work towards detecting potential online misinformation through the creation of a large-scale dataset that expands misinformation detection into the multimodal space of image and text. As misinformation can be both human-written and machine-written, we investigate the memorization and subsequent generation of misinformation through the lens of conspiracy theories.The final part of the dissertation describes recent work in AI safety regarding text that may lead to physical harm. This research analyzes covertly unsafe text across various language modeling tasks including generation, reasoning, and detection. Altogether, this work sheds light on the undiscovered and underrepresented risks in large language models. This can advance current research toward building safer and more equitable natural language processing systems. We conclude with discussions of future research in Responsible AI that expand upon work in the three areas.
- 일반주제명
- Computer engineering.
- 일반주제명
- Computer science.
- 키워드
- Machine learning
- 키워드
- Responsible AI
- 기타저자
- University of California, Santa Barbara Computer Science
- 기본자료저록
- Dissertations Abstracts International. 85-02B.
- 기본자료저록
- Dissertation Abstract International
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008240612s2023 us |||||||||||||||c||eng d■001000016933495
■00520240214101253
■006m o d
■007cr#unu||||||||
■020 ▼a9798380154154
■035 ▼a(MiAaPQ)AAI30529918
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a621.3
■1001 ▼aLevy, Sharon Gabriel.
■24510▼aResponsible AI via Responsible Large Language Models▼h[electronic resource]
■260 ▼a[S.l.]:▼bUniversity of California, Santa Barbara. ▼c2023
■260 1▼aAnn Arbor :▼bProQuest Dissertations & Theses, ▼c2023
■300 ▼a1 online resource(142 p.)
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-02, Section: B.
■500 ▼aAdvisor: Wang, William Yang.
■5021 ▼aThesis (Ph.D.)--University of California, Santa Barbara, 2023.
■506 ▼aThis item must not be sold to any third party vendors.
■520 ▼aLarge language models have advanced the state-of-the-art in natural language processing and achieved success in tasks such as summarization, question answering, and text classification. However, these models are trained on large-scale datasets, which may include harmful information. Studies have shown that as a result, the models can exhibit social biases and generate misinformation after training. This dissertation discusses research on analyzing and interpreting the risks of large language models across the areas of fairness, trustworthiness, and safety. The first part of this dissertation analyzes issues of fairness related to social biases in large language models. We first investigate issues of dialect bias pertaining to African American English and Standard American English within the context of text generation. We also analyze a more complex setting of fairness: cases in which multiple attributes affect each other to form compound biases. This is studied in relation to gender and seniority attributes.The second part focuses on trustworthiness and the spread of misinformation across different scopes: prevention, detection, and memorization. We describe an open-domain question-answering system for emergent domains that uses various retrieval and re-ranking techniques to provide users with information from trustworthy sources. This is demonstrated in the context of the emergent COVID-19 pandemic. We further work towards detecting potential online misinformation through the creation of a large-scale dataset that expands misinformation detection into the multimodal space of image and text. As misinformation can be both human-written and machine-written, we investigate the memorization and subsequent generation of misinformation through the lens of conspiracy theories.The final part of the dissertation describes recent work in AI safety regarding text that may lead to physical harm. This research analyzes covertly unsafe text across various language modeling tasks including generation, reasoning, and detection. Altogether, this work sheds light on the undiscovered and underrepresented risks in large language models. This can advance current research toward building safer and more equitable natural language processing systems. We conclude with discussions of future research in Responsible AI that expand upon work in the three areas.
■590 ▼aSchool code: 0035.
■650 4▼aComputer engineering.
■650 4▼aComputer science.
■653 ▼aMachine learning
■653 ▼aNatural language processing
■653 ▼aResponsible AI
■690 ▼a0800
■690 ▼a0984
■690 ▼a0464
■71020▼aUniversity of California, Santa Barbara▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g85-02B.
■773 ▼tDissertation Abstract International
■790 ▼a0035
■791 ▼aPh.D.
■792 ▼a2023
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T16933495▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.
■980 ▼a202402▼f2024
![Responsible AI via Responsible Large Language Models - [electronic resource]](/Users/Baul/Images/book.png)

