본문

서브메뉴

Spectro-Temporal and Linguistic Processing of Speech in Artificial and Biological Neural Networks
Spectro-Temporal and Linguistic Processing of Speech in Artificial and Biological Neural N...
Spectro-Temporal and Linguistic Processing of Speech in Artificial and Biological Neural Networks

Detailed Information

자료유형  
 학위논문 서양
최종처리일시  
20250211151510
ISBN  
9798383596920
DDC  
616
저자명  
Keshishian, Menoua.
서명/저자  
Spectro-Temporal and Linguistic Processing of Speech in Artificial and Biological Neural Networks
발행사항  
[Sl] : Columbia University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
172 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-02, Section: A.
주기사항  
Advisor: Mesgarani, Nima.
학위논문주기  
Thesis (Ph.D.)--Columbia University, 2024.
초록/해제  
요약Humans possess the fascinating ability to communicate the most complex of ideas through spoken language, without requiring any external tools. This process has two sides-a speaker producing speech, and a listener comprehending it. While the two actions are intertwined in many ways, they entail differential activation of neural circuits in the brains of the speaker and the listener. Both processes are the active subject of artificial intelligence research, under the names of speech synthesis and automatic speech recognition, respectively. While the capabilities of these artificial models are approaching human levels, there are still many unanswered questions about how our brains do this task effortlessly. But the advances in these artificial models allow us the opportunity to study human speech recognition through a computational lens that we did not have before. This dissertation explores the intricate processes of speech perception and comprehension by drawing parallels between artificial and biological neural networks, through the use of computational frameworks that attempt to model either the brain circuits involved in speech recognition, or the process of speech recognition itself.There are two general types of analyses in this dissertation. The first type involves studying neural responses recorded directly through invasive electrophysiology from human participants listening to speech excerpts. The second type involves analyzing artificial neural networks trained to perform the same task of speech recognition, as a potential model for our brains. The first study introduces a novel framework leveraging deep neural networks (DNNs) for interpretable modeling of nonlinear sensory receptive fields, offering an enhanced understanding of auditory neural responses in humans. This approach not only predicts auditory neural responses with increased accuracy but also deciphers distinct nonlinear encoding properties, revealing new insights into the computational principles underlying sensory processing in the auditory cortex. The second study delves into the dynamics of temporal processing of speech in automatic speech recognition networks, elucidating how these systems learn to integrate information across various timescales, mirroring certain aspects of biological temporal processing. The third study presents a rigorous examination of the neural encoding of linguistic information of speech in the auditory cortex during speech comprehension. By analyzing neural responses to natural speech, we identify explicit, distributed neural encoding across multiple levels of linguistic processing, from phonetic features to semantic meaning. This multilevel linguistic analysis contributes to our understanding of the hierarchical and distributed nature of speech processing in the human brain. The final chapter of this dissertation compares linguistic encoding between an automatic speech recognition system and the human brain, elucidating their computational and representational similarities and differences. This comparison underscores the nuanced understanding of how linguistic information is processed and encoded across different systems, offering insights into both biological perception and artificial intelligence mechanisms in speech processing.Through this comprehensive examination, the dissertation advances our understanding of the computational and representational foundations of speech perception, demonstrating the potential of interdisciplinary approaches that bridge neuroscience and artificial intelligence to uncover the underlying mechanisms of speech processing in both artificial and biological systems.
일반주제명  
Neurosciences
일반주제명  
Linguistics
일반주제명  
Information technology
키워드  
Auditory neuroscience
키워드  
Electrocorticography
키워드  
Language processing
키워드  
Machine learning
키워드  
Speech processing
키워드  
Speech recognition
기타저자  
Columbia University Electrical Engineering
기본자료저록  
Dissertations Abstracts International. 86-02A.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017161979
■00520250211151510
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798383596920
■035    ▼a(MiAaPQ)AAI31299615
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a616
■1001  ▼aKeshishian,  Menoua.
■24510▼aSpectro-Temporal  and  Linguistic  Processing  of  Speech  in  Artificial  and  Biological  Neural  Networks
■260    ▼a[Sl]▼bColumbia  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a172  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-02,  Section:  A.
■500    ▼aAdvisor:  Mesgarani,  Nima.
■5021  ▼aThesis  (Ph.D.)--Columbia  University,  2024.
■520    ▼aHumans  possess  the  fascinating  ability  to  communicate  the  most  complex  of  ideas  through  spoken  language,  without  requiring  any  external  tools.  This  process  has  two  sides-a  speaker  producing  speech,  and  a  listener  comprehending  it.  While  the  two  actions  are  intertwined  in  many  ways,  they  entail  differential  activation  of  neural  circuits  in  the  brains  of  the  speaker  and  the  listener.  Both  processes  are  the  active  subject  of  artificial  intelligence  research,  under  the  names  of  speech  synthesis  and  automatic  speech  recognition,  respectively.  While  the  capabilities  of  these  artificial  models  are  approaching  human  levels,  there  are  still  many  unanswered  questions  about  how  our  brains  do  this  task  effortlessly.  But  the  advances  in  these  artificial  models  allow  us  the  opportunity  to  study  human  speech  recognition  through  a  computational  lens  that  we  did  not  have  before.  This  dissertation  explores  the  intricate  processes  of  speech  perception  and  comprehension  by  drawing  parallels  between  artificial  and  biological  neural  networks,  through  the  use  of  computational  frameworks  that  attempt  to  model  either  the  brain  circuits  involved  in  speech  recognition,  or  the  process  of  speech  recognition  itself.There  are  two  general  types  of  analyses  in  this  dissertation.  The  first  type  involves  studying  neural  responses  recorded  directly  through  invasive  electrophysiology  from  human  participants  listening  to  speech  excerpts.  The  second  type  involves  analyzing  artificial  neural  networks  trained  to  perform  the  same  task  of  speech  recognition,  as  a  potential  model  for  our  brains.  The  first  study  introduces  a  novel  framework  leveraging  deep  neural  networks  (DNNs)  for  interpretable  modeling  of  nonlinear  sensory  receptive  fields,  offering  an  enhanced  understanding  of  auditory  neural  responses  in  humans.  This  approach  not  only  predicts  auditory  neural  responses  with  increased  accuracy  but  also  deciphers  distinct  nonlinear  encoding  properties,  revealing  new  insights  into  the  computational  principles  underlying  sensory  processing  in  the  auditory  cortex.  The  second  study  delves  into  the  dynamics  of  temporal  processing  of  speech  in  automatic  speech  recognition  networks,  elucidating  how  these  systems  learn  to  integrate  information  across  various  timescales,  mirroring  certain  aspects  of  biological  temporal  processing.  The  third  study  presents  a  rigorous  examination  of  the  neural  encoding  of  linguistic  information  of  speech  in  the  auditory  cortex  during  speech  comprehension.  By  analyzing  neural  responses  to  natural  speech,  we  identify  explicit,  distributed  neural  encoding  across  multiple  levels  of  linguistic  processing,  from  phonetic  features  to  semantic  meaning.  This  multilevel  linguistic  analysis  contributes  to  our  understanding  of  the  hierarchical  and  distributed  nature  of  speech  processing  in  the  human  brain.  The  final  chapter  of  this  dissertation  compares  linguistic  encoding  between  an  automatic  speech  recognition  system  and  the  human  brain,  elucidating  their  computational  and  representational  similarities  and  differences.  This  comparison  underscores  the  nuanced  understanding  of  how  linguistic  information  is  processed  and  encoded  across  different  systems,  offering  insights  into  both  biological  perception  and  artificial  intelligence  mechanisms  in  speech  processing.Through  this  comprehensive  examination,  the  dissertation  advances  our  understanding  of  the  computational  and  representational  foundations  of  speech  perception,  demonstrating  the  potential  of  interdisciplinary  approaches  that  bridge  neuroscience  and  artificial  intelligence  to  uncover  the  underlying  mechanisms  of  speech  processing  in  both  artificial  and  biological  systems.
■590    ▼aSchool  code:  0054.
■650  4▼aNeurosciences
■650  4▼aLinguistics
■650  4▼aInformation  technology
■653    ▼aAuditory  neuroscience
■653    ▼aElectrocorticography
■653    ▼aLanguage  processing
■653    ▼aMachine  learning
■653    ▼aSpeech  processing
■653    ▼aSpeech  recognition
■690    ▼a0317
■690    ▼a0800
■690    ▼a0489
■690    ▼a0290
■71020▼aColumbia  University▼bElectrical  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g86-02A.
■790    ▼a0054
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17161979▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

Preview

Export

ChatGPT Discussion

AI Recommended Related Books


    New Books MORE
    Statistics for the past 3 years. Go to brief

    Подробнее информация.

    • Бронирование
    • не существует
    • моя папка
    • Первый запрос зрения
    • Non-Book Loan Application
    • Nighttime Book Loan Application
    материал
    Reg No. Количество платежных Местоположение статус Ленд информации
    TF12244 전자도서 대출가능 My Folder 부재도서신고 비도서대출신청 야간 도서대출신청

    * Бронирование доступны в заимствований книги. Чтобы сделать предварительный заказ, пожалуйста, нажмите кнопку бронирование

    Books borrowed together with this book

    Related Popular Books

    Available after logging in.