본문

서브메뉴

Information Extraction on Scientific Literature Under Limited Supervision
Information Extraction on Scientific Literature Under Limited Supervision
Information Extraction on Scientific Literature Under Limited Supervision

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260209102904
ISBN  
9798265401007
DDC  
574
저자명  
Bai, Fan.
서명/저자  
Information Extraction on Scientific Literature Under Limited Supervision
발행사항  
[Sl] : Georgia Institute of Technology, 2023
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2023
형태사항  
147 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-06, Section: B.
주기사항  
Advisor: Ritter, Alan.
학위논문주기  
Thesis (Ph.D.)--Georgia Institute of Technology, 2023.
초록/해제  
요약The exponential growth of scientific literature presents both challenges and opportunities for researchers across various disciplines. Effectively extracting pertinent information from this extensive corpus is crucial for advancing knowledge, enhancing collaboration, and driving innovation. However, manual extraction is a laborious and time-consuming process, underscoring the demand for automated solutions. Information extraction (IE), a subfield of natural language processing (NLP) focused on automatically extracting structured information from unstructured data sources, plays a crucial role in addressing this challenge. Despite their success, many IE methods often require substantial human-annotated data, which might not be easily accessible, particularly in specialized scientific domains. This highlights the need for adaptable and robust techniques capable of functioning with limited supervision.In this thesis, we study the task of information extraction on scientific literature, particularly addressing the challenge of limited (human) supervision. Specifically, our work has delved into four key dimensions of this problem. First, we explore the potential of harnessing easily accessible resources, like knowledge bases, to develop IE systems without direct human supervision. Second, we examine the use of pre-trained language models to create effective and efficient scientific IE systems, experimenting with various fine-tuning architectures and learning strategies. Next, we investigate the balance between the labor expenditure of human annotation and the computational cost linked with domain-specific pre-training, to achieve optimal performance under the budget constraints. Lastly, we capitalize on the emerging capabilities of large pre-trained language models by showcasing how information extraction can be achieved solely based on a human-crafted data schema. Through these explorations, this thesis aims to lay a solid foundation for the continued advancement of scientific IE under limited supervision.
일반주제명  
Adaptation
일반주제명  
Human performance
일반주제명  
Semantics
일반주제명  
Chemical synthesis
일반주제명  
Language
일반주제명  
Computer science
키워드  
Natural language processing
기타저자  
Georgia Institute of Technology.
기본자료저록  
Dissertations Abstracts International. 87-06B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260203s2023        us                              c    eng  d
■001000017365964
■00520260209102904
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798265401007
■035    ▼a(MiAaPQ)AAI32315616
■035    ▼a(MiAaPQ)GeorgiaTech73149
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574
■1001  ▼aBai,  Fan.
■24510▼aInformation  Extraction  on  Scientific  Literature  Under  Limited  Supervision
■260    ▼a[Sl]▼bGeorgia  Institute  of  Technology▼c2023
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2023
■300    ▼a147  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-06,  Section:  B.
■500    ▼aAdvisor:  Ritter,  Alan.
■5021  ▼aThesis  (Ph.D.)--Georgia  Institute  of  Technology,  2023.
■520    ▼aThe  exponential  growth  of  scientific  literature  presents  both  challenges  and  opportunities  for  researchers  across  various  disciplines.  Effectively  extracting  pertinent  information  from  this  extensive  corpus  is  crucial  for  advancing  knowledge,  enhancing  collaboration,  and  driving  innovation.  However,  manual  extraction  is  a  laborious  and  time-consuming  process,  underscoring  the  demand  for  automated  solutions.  Information  extraction  (IE),  a  subfield  of  natural  language  processing  (NLP)  focused  on  automatically  extracting  structured  information  from  unstructured  data  sources,  plays  a  crucial  role  in  addressing  this  challenge.  Despite  their  success,  many  IE  methods  often  require  substantial  human-annotated  data,  which  might  not  be  easily  accessible,  particularly  in  specialized  scientific  domains.  This  highlights  the  need  for  adaptable  and  robust  techniques  capable  of  functioning  with  limited  supervision.In  this  thesis,  we  study  the  task  of  information  extraction  on  scientific  literature,  particularly  addressing  the  challenge  of  limited  (human)  supervision.  Specifically,  our  work  has  delved  into  four  key  dimensions  of  this  problem.  First,  we  explore  the  potential  of  harnessing  easily  accessible  resources,  like  knowledge  bases,  to  develop  IE  systems  without  direct  human  supervision.  Second,  we  examine  the  use  of  pre-trained  language  models  to  create  effective  and  efficient  scientific  IE  systems,  experimenting  with  various  fine-tuning  architectures  and  learning  strategies.  Next,  we  investigate  the  balance  between  the  labor  expenditure  of  human  annotation  and  the  computational  cost  linked  with  domain-specific  pre-training,  to  achieve  optimal  performance  under  the  budget  constraints.  Lastly,  we  capitalize  on  the  emerging  capabilities  of  large  pre-trained  language  models  by  showcasing  how  information  extraction  can  be  achieved  solely  based  on  a  human-crafted  data  schema.  Through  these  explorations,  this  thesis  aims  to  lay  a  solid  foundation  for  the  continued  advancement  of  scientific  IE  under  limited  supervision.
■590    ▼aSchool  code:  0078.
■650  4▼aAdaptation
■650  4▼aHuman  performance
■650  4▼aSemantics
■650  4▼aChemical  synthesis
■650  4▼aLanguage
■650  4▼aComputer  science
■653    ▼aNatural  language  processing
■690    ▼a0984
■690    ▼a0679
■71020▼aGeorgia  Institute  of  Technology.
■7730  ▼tDissertations  Abstracts  International▼g87-06B.
■790    ▼a0078
■791    ▼aPh.D.
■792    ▼a2023
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17365964▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF15744 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.