서브메뉴
검색
The Recorded Voice: Computational Analysis of Trends and Expressions in Speech and Popular Song
The Recorded Voice: Computational Analysis of Trends and Expressions in Speech and Popular Song
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202104646
- ISBN
- 9798291564370
- DDC
- 780
- 서명/저자
- The Recorded Voice: Computational Analysis of Trends and Expressions in Speech and Popular Song
- 발행사항
- [Sl] : New York University, 2025
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2025
- 형태사항
- 173 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-02, Section: B.
- 주기사항
- Advisor: McFee, Brian;Ripolles, Pablo.
- 학위논문주기
- Thesis (Ph.D.)--New York University, 2025.
- 초록/해제
- 요약This dissertation uses computational methods to analyze large corpora of vocal recordings, examining the relationship between recording, speech, and song. It begins with professionally recorded Popular music, examining how vocal pitch and timing have changed across genres and decades. Vocals in rap music stand out for their distinct pitch patterns and timing, while overall pitch variation and microtiming deviation have decreased over time-suggesting a possible decline in vocal complexity, at least in this dataset.The focus then shifts to amateur vocal recordings, both sung and spoken, captured in at-home environments. This dissertation introduces a new listener-evaluated dataset of 4,300 ratings of karaoke singing and audiobook narration recordings. It also introduces a baseline model for predicting listener ratings of recorded voices from audio alone. Unsurprisingly, objective features like intelligibility and recording quality are more predictable automatically.Building on this, a final study uses the CLAP audio-text model to identify and enhance degraded karaoke and audiobook recordings, comparing its recommendations to 4,600 human ratings collected for this research. While CLAP effectively detects some degradations, discrepancies with listener ratings highlight the continued value of human evaluations.
- 일반주제명
- Music
- 일반주제명
- Computer science
- 일반주제명
- Cognitive psychology
- 키워드
- Audio quality
- 키워드
- Popular music
- 키워드
- Recording
- 키워드
- Voice
- 기타저자
- New York University Music and Performing Arts Professions
- 기본자료저록
- Dissertations Abstracts International. 87-02B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2025 us c eng d■001000017358336
■00520260202104646
■006m o d
■007cr#unu||||||||
■020 ▼a9798291564370
■035 ▼a(MiAaPQ)AAI32114661
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a780
■1001 ▼aGeorgieva, Elena-Theodora.
■24510▼aThe Recorded Voice: Computational Analysis of Trends and Expressions in Speech and Popular Song
■260 ▼a[Sl]▼bNew York University▼c2025
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a173 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-02, Section: B.
■500 ▼aAdvisor: McFee, Brian;Ripolles, Pablo.
■5021 ▼aThesis (Ph.D.)--New York University, 2025.
■520 ▼aThis dissertation uses computational methods to analyze large corpora of vocal recordings, examining the relationship between recording, speech, and song. It begins with professionally recorded Popular music, examining how vocal pitch and timing have changed across genres and decades. Vocals in rap music stand out for their distinct pitch patterns and timing, while overall pitch variation and microtiming deviation have decreased over time-suggesting a possible decline in vocal complexity, at least in this dataset.The focus then shifts to amateur vocal recordings, both sung and spoken, captured in at-home environments. This dissertation introduces a new listener-evaluated dataset of 4,300 ratings of karaoke singing and audiobook narration recordings. It also introduces a baseline model for predicting listener ratings of recorded voices from audio alone. Unsurprisingly, objective features like intelligibility and recording quality are more predictable automatically.Building on this, a final study uses the CLAP audio-text model to identify and enhance degraded karaoke and audiobook recordings, comparing its recommendations to 4,600 human ratings collected for this research. While CLAP effectively detects some degradations, discrepancies with listener ratings highlight the continued value of human evaluations.
■590 ▼aSchool code: 0146.
■650 4▼aMusic
■650 4▼aComputer science
■650 4▼aCognitive psychology
■653 ▼aAudio quality
■653 ▼aListener evaluations
■653 ▼aPopular music
■653 ▼aRecording
■653 ▼aVoice
■690 ▼a0413
■690 ▼a0984
■690 ▼a0633
■71020▼aNew York University▼bMusic and Performing Arts Professions.
■7730 ▼tDissertations Abstracts International▼g87-02B.
■790 ▼a0146
■791 ▼aPh.D.
■792 ▼a2025
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17358336▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


