본문

서브메뉴

The Recorded Voice: Computational Analysis of Trends and Expressions in Speech and Popular Song
The Recorded Voice: Computational Analysis of Trends and Expressions in Speech and Popular...
The Recorded Voice: Computational Analysis of Trends and Expressions in Speech and Popular Song

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20260202104646
ISBN  
9798291564370
DDC  
780
저자명  
Georgieva, Elena-Theodora.
서명/저자  
The Recorded Voice: Computational Analysis of Trends and Expressions in Speech and Popular Song
발행사항  
[Sl] : New York University, 2025
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2025
형태사항  
173 p
주기사항  
Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
주기사항  
Advisor: McFee, Brian;Ripolles, Pablo.
학위논문주기  
Thesis (Ph.D.)--New York University, 2025.
초록/해제  
요약This dissertation uses computational methods to analyze large corpora of vocal recordings, examining the relationship between recording, speech, and song. It begins with professionally recorded Popular music, examining how vocal pitch and timing have changed across genres and decades. Vocals in rap music stand out for their distinct pitch patterns and timing, while overall pitch variation and microtiming deviation have decreased over time-suggesting a possible decline in vocal complexity, at least in this dataset.The focus then shifts to amateur vocal recordings, both sung and spoken, captured in at-home environments. This dissertation introduces a new listener-evaluated dataset of 4,300 ratings of karaoke singing and audiobook narration recordings. It also introduces a baseline model for predicting listener ratings of recorded voices from audio alone. Unsurprisingly, objective features like intelligibility and recording quality are more predictable automatically.Building on this, a final study uses the CLAP audio-text model to identify and enhance degraded karaoke and audiobook recordings, comparing its recommendations to 4,600 human ratings collected for this research. While CLAP effectively detects some degradations, discrepancies with listener ratings highlight the continued value of human evaluations.
일반주제명  
Music
일반주제명  
Computer science
일반주제명  
Cognitive psychology
키워드  
Audio quality
키워드  
Listener evaluations
키워드  
Popular music
키워드  
Recording
키워드  
Voice
기타저자  
New York University Music and Performing Arts Professions
기본자료저록  
Dissertations Abstracts International. 87-02B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008260126s2025        us                              c    eng  d
■001000017358336
■00520260202104646
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798291564370
■035    ▼a(MiAaPQ)AAI32114661
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a780
■1001  ▼aGeorgieva,  Elena-Theodora.
■24510▼aThe  Recorded  Voice:  Computational  Analysis  of  Trends  and  Expressions  in  Speech  and  Popular  Song
■260    ▼a[Sl]▼bNew  York  University▼c2025
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2025
■300    ▼a173  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  87-02,  Section:  B.
■500    ▼aAdvisor:  McFee,  Brian;Ripolles,  Pablo.
■5021  ▼aThesis  (Ph.D.)--New  York  University,  2025.
■520    ▼aThis  dissertation  uses  computational  methods  to  analyze  large  corpora  of  vocal  recordings,  examining  the  relationship  between  recording,  speech,  and  song.  It  begins  with  professionally  recorded  Popular  music,  examining  how  vocal  pitch  and  timing  have  changed  across  genres  and  decades.  Vocals  in  rap  music  stand  out  for  their  distinct  pitch  patterns  and  timing,  while  overall  pitch  variation  and  microtiming  deviation  have  decreased  over  time-suggesting  a  possible  decline  in  vocal  complexity,  at  least  in  this  dataset.The  focus  then  shifts  to  amateur  vocal  recordings,  both  sung  and  spoken,  captured  in  at-home  environments.  This  dissertation  introduces  a  new  listener-evaluated  dataset  of  4,300  ratings  of  karaoke  singing  and  audiobook  narration  recordings.  It  also  introduces  a  baseline  model  for  predicting  listener  ratings  of  recorded  voices  from  audio  alone.  Unsurprisingly,  objective  features  like  intelligibility  and  recording  quality  are  more  predictable  automatically.Building  on  this,  a  final  study  uses  the  CLAP  audio-text  model  to  identify  and  enhance  degraded  karaoke  and  audiobook  recordings,  comparing  its  recommendations  to  4,600  human  ratings  collected  for  this  research.  While  CLAP  effectively  detects  some  degradations,  discrepancies  with  listener  ratings  highlight  the  continued  value  of  human  evaluations.
■590    ▼aSchool  code:  0146.
■650  4▼aMusic
■650  4▼aComputer  science
■650  4▼aCognitive  psychology
■653    ▼aAudio  quality
■653    ▼aListener  evaluations
■653    ▼aPopular  music
■653    ▼aRecording
■653    ▼aVoice
■690    ▼a0413
■690    ▼a0984
■690    ▼a0633
■71020▼aNew  York  University▼bMusic  and  Performing  Arts  Professions.
■7730  ▼tDissertations  Abstracts  International▼g87-02B.
■790    ▼a0146
■791    ▼aPh.D.
■792    ▼a2025
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358336▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF15215 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.