본문

서브메뉴

Odors as ''Natural Language'': Sparse Neural Networks in Mammalian Olfactory Systems and Large Language Models
Odors as ''Natural Language'': Sparse Neural Networks in Mammalian Olfactory Systems and L...
Odors as ''Natural Language'': Sparse Neural Networks in Mammalian Olfactory Systems and Large Language Models

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202103504
ISBN  
9798280717947
DDC  
574.191
저자명  
Liu, Bo.
서명/저자  
Odors as Natural Language: Sparse Neural Networks in Mammalian Olfactory Systems and Large Language Models
발행사항  
[Sl] : Harvard University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
155 p
주기사항  
Source: Dissertations Abstracts International, Volume: 86-12, Section: B.
주기사항  
Advisor: Murthy, Venkatesh N.
학위논문주기  
Thesis (Ph.D.)--Harvard University, 2025.
초록/해제  
요약The studies of physics, neuroscience, and artificial intelligence (AI) have a long intertwined history. Particularly, sparse connectivity is a common feature of the brain neural networks and a key focus in AI for efficient computation; notably, pruning trained networks for sparse connectivity has a long history, partially inspired by neuroscience. This thesis explores sparse neural networks through two linked research topics: one focused on the brain (bilateral alignment in olfactory systems), and the other on AI (pruning large language models for on-device AI assistants).For the first topic, inspired by mammalian dual nostrils creating two cortical neural representations of odors, in Chapter 1, we studied how to construct the inter-hemispheric projections aligning these representations. We hypothesized that this construction originates from online learning since mammals are constantly breathing. With a local Hebbian rule, we found that sparse interhemispheric projections suffice for bilateral alignment and discovered an inverse scaling that more cortical neurons allow sparser projections. Also, the local Hebbian rule was found to approximate the global stochastic gradient descent (SGD) rule since their update vectors align, suggesting that biologically plausible learning rules can approximate global learning rules if they contain the gradient information of the latter.The next chapter extends Chapter 1 from four perspectives: an analysis of the update vector alignment between Hebbian and SGD rules and how it depends on the network parameters; a simple theory that recurrent connections in olfactory cortex may improve the bilateral alignment, inspired by the Hopfield Networks (associative memory) 1 and similar to the design of Google Titans model that combines recurrent neural networks with Transformers; the dynamical properties of Hebbian learning; and finally, the geometric landscape of Hebbian learning.A similar inverse scaling has been discovered in the Transformer attention matrices used in large language models (LLMs), which motivated the second topic. Concretely, we pruned pretrained Meta Llama-2 and Llama-3 models to obtain models with fewer parameters and develop on-device AI assistants, explored their sparsity limits, and compared their performance at the limits. We found that more than 50% of the parameters in both models could be pruned, and Llama-3 produced fewer factual errors at the sparsity limit but required more parameters presumably due to its training settings and dataset.In summary, by studying sparsity in both biological and artificial neural networks, this thesis may provide valuable insights into the general bilateral alignment problem in neuroscience (across different modalities and brain regions such as the frontal cortex responsible for short-term and motor response and the medial entorhinal cortex for spatial memory), open the door to interesting theoretical questions, and inspire more efficient AI algorithms or applications.
일반주제명  
Biophysics
일반주제명  
Neurosciences
키워드  
Neural networks
키워드  
Stochastic gradient descent
키워드  
Hebbian learning
키워드  
Artificial neural networks
키워드  
Spatial memory
기타저자  
Harvard University Biology Molecular and Cellular
기본자료저록  
Dissertations Abstracts International. 86-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017357382
■00520260202103504
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798280717947
■035    ▼a(MiAaPQ)AAI32002155
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a574.191
■1001  ▼aLiu,  Bo.▼0(orcid)0000-0002-2819-608X
■24510▼aOdors  as  ''Natural  Language'':  Sparse  Neural  Networks  in  Mammalian  Olfactory  Systems  and  Large  Language  Models
■260    ▼a[Sl]▼bHarvard  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a155  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  86-12,  Section:  B.
■500    ▼aAdvisor:  Murthy,  Venkatesh  N.
■5021  ▼aThesis  (Ph.D.)--Harvard  University,  2025.
■520    ▼aThe  studies  of  physics,  neuroscience,  and  artificial  intelligence  (AI)  have  a  long  intertwined  history.  Particularly,  sparse  connectivity  is  a  common  feature  of  the  brain  neural  networks  and  a  key  focus  in  AI  for  efficient  computation;  notably,  pruning  trained  networks  for  sparse  connectivity  has  a  long  history,  partially  inspired  by  neuroscience.  This  thesis  explores  sparse  neural  networks  through  two  linked  research  topics:  one  focused  on  the  brain  (bilateral  alignment  in  olfactory  systems),  and  the  other  on  AI  (pruning  large  language  models  for  on-device  AI  assistants).For  the  first  topic,  inspired  by  mammalian  dual  nostrils  creating  two  cortical  neural  representations  of  odors,  in  Chapter  1,  we  studied  how  to  construct  the  inter-hemispheric  projections  aligning  these  representations.  We  hypothesized  that  this  construction  originates  from  online  learning  since  mammals  are  constantly  breathing.  With  a  local  Hebbian  rule,  we  found  that  sparse  interhemispheric  projections  suffice  for  bilateral  alignment  and  discovered  an  inverse  scaling  that  more  cortical  neurons  allow  sparser  projections.  Also,  the  local  Hebbian  rule  was  found  to  approximate  the  global  stochastic  gradient  descent  (SGD)  rule  since  their  update  vectors  align,  suggesting  that  biologically  plausible  learning  rules  can  approximate  global  learning  rules  if  they  contain  the  gradient  information  of  the  latter.The  next  chapter  extends  Chapter  1  from  four  perspectives:  an  analysis  of  the  update  vector  alignment  between  Hebbian  and  SGD  rules  and  how  it  depends  on  the  network  parameters;  a  simple  theory  that  recurrent  connections  in  olfactory  cortex  may  improve  the  bilateral  alignment,  inspired  by  the  Hopfield  Networks  (associative  memory)  1  and  similar  to  the  design  of  Google  Titans  model  that  combines  recurrent  neural  networks  with  Transformers;  the  dynamical  properties  of  Hebbian  learning;  and  finally,  the  geometric  landscape  of  Hebbian  learning.A  similar  inverse  scaling  has  been  discovered  in  the  Transformer  attention  matrices  used  in  large  language  models  (LLMs),  which  motivated  the  second  topic.  Concretely,  we  pruned  pretrained  Meta  Llama-2  and  Llama-3  models  to  obtain  models  with  fewer  parameters  and  develop  on-device  AI  assistants,  explored  their  sparsity  limits,  and  compared  their  performance  at  the  limits.  We  found  that  more  than  50%  of  the  parameters  in  both  models  could  be  pruned,  and  Llama-3  produced  fewer  factual  errors  at  the  sparsity  limit  but  required  more  parameters  presumably  due  to  its  training  settings  and  dataset.In  summary,  by  studying  sparsity  in  both  biological  and  artificial  neural  networks,  this  thesis  may  provide  valuable  insights  into  the  general  bilateral  alignment  problem  in  neuroscience  (across  different  modalities  and  brain  regions  such  as  the  frontal  cortex  responsible  for  short-term  and  motor  response  and  the  medial  entorhinal  cortex  for  spatial  memory),  open  the  door  to  interesting  theoretical  questions,  and  inspire  more  efficient  AI  algorithms  or  applications.
■590    ▼aSchool  code:  0084.
■650  4▼aBiophysics
■650  4▼aNeurosciences
■653    ▼aNeural  networks
■653    ▼aStochastic  gradient  descent
■653    ▼aHebbian  learning
■653    ▼aArtificial  neural  networks
■653    ▼aSpatial  memory
■690    ▼a0786
■690    ▼a0317
■690    ▼a0800
■71020▼aHarvard  University▼bBiology,  Molecular  and  Cellular.
■7730  ▼tDissertations  Abstracts  International▼g86-12B.
■790    ▼a0084
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17357382▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF17387 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.