서브메뉴
검색
Advancements in Perceptual Quality Assessment for Interactive Media: From Mobile Cloud Gaming to Human Avatar Videos and Facial Expressions
Advancements in Perceptual Quality Assessment for Interactive Media: From Mobile Cloud Gaming to Human Avatar Videos and Facial Expressions
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260311091546.5
- ISBN
- 9798270232504
- DDC
- 006.696
- 서명/저자
- Advancements in Perceptual Quality Assessment for Interactive Media: From Mobile Cloud Gaming to Human Avatar Videos and Facial Expressions / Yu-Chih Berrie Chen
- 발행사항
- [Sl] : The University of Texas at Austin, 2025
- 형태사항
- 1 electronic resource (144 pages)
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-06, Section: A.
- 주기사항
- Advisors: Bovik, Alan C. Committee members: Kim, Hyeji; Geisler, Wilson S., III; de Veciana, Gustavo; Vikalo, Haris.
- 학위논문주기
- - Ph.D. : The University of Texas at Austin, 2025.
- 초록/해제
- 요약Over the past decade, interactive media-including mobile cloud gaming and virtual reality (VR) applications-have grown rapidly, necessitating efficient methods to assess and optimize perceptual visual quality. This dissertation presents a comprehensive study of human visual quality judgments across three major domains: gaming video streaming, rendered human avatars in VR/AR, and facial expression similarity in human faces. As streaming these media becomes increasingly prevalent, advanced video compression protocols and Video Quality Assessment (VQA) algorithms are crucial for balancing high-quality visual delivery and variable bandwidth conditions. In the realm of mobile cloud gaming, the computational load of rendering games is offloaded to remote servers, allowing users to play on lightweight devices like smartphones. However, this introduces unique challenges in latency, compression, and real-time distortion. To address this, we introduce GAMIVAL, a no-reference (NR) VQA model tailored for gaming content, which exhibits statistical characteristics distinct from naturalistic videos. GAMIVAL models spatial and temporal distortions, neural noise, and deep semantic features, achieving superior performance on the LIVE-Meta Mobile Cloud Gaming video quality database. Expanding to immersive VR-based experiences, we investigate the perceived quality of rendered human avatars-often referred to as "holograms." We introduce the LIVE-Meta Rendered Human Avatar VQA Database, which includes 720 avatar videos compressed with 20 encoding parameter combinations and annotated with human perceptual quality ratings collected via six-degrees-of-freedom VR headsets. This large-scale database enables the benchmarking of both full-reference and no-reference VQA models, including a newly proposed model, HoloQA. Our study addresses a critical gap in current VQA research: the lack of a dataset designed specifically for evaluating rendered human body avatar videos. Moreover, we explore perceptual similarity in human facial expressions, a key component of affective computing. Existing metrics often fail to capture the nuanced semantics of expressions. To address this, we propose FaceSIM, a novel similarity model built on contrastive learning principles and vision transformer backbones. The architecture comprises two complementary branches: a fixed content-aware encoder that extracts stable, identity-agnostic facial features, and a trainable similarity-aware encoder that learns to capture perceptual distinctions through contrastive supervision. FaceSIM is able to separate identities while modeling nuanced expression differences. A simplified one-way contrastive objective guides the training, encouraging perceptually similar facial expressions to cluster in the embedding space while pushing dissimilar ones apart. The resulting embeddings support both zero-shot and regression-based similarity estimation. Evaluations on a large-scale facial expression similarity dataset demonstrate that FaceSIM aligns closely with human perception across diverse expressions, identities, and intensities. Together, these contributions form a unified framework for understanding and improving human-centered visual quality assessment across a range of interactive media. By developing novel models and datasets, this dissertation enhances the perceptual fidelity of cloud gaming, immersive holographic communication, and facial expression reconstruction-advancing multimedia technologies for next-generation user experiences.
- 언어주기
- English
- 일반주제명
- Computer engineering
- 일반주제명
- Information technology
- 일반주제명
- Film studies
- 키워드
- Virtual reality
- 키워드
- Avatar videos
- 기타저자
- The University of Texas at Austin Electrical and Computer Engineering
- 기본자료저록
- Dissertations Abstracts International. 87-06A.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260311s2025 us eng d■001000017361231
■00520260311091546.5
■006m o d
■007cr|nu||||||||
■020 ▼a9798270232504
■040 ▼aMiAaPQD▼beng▼cMiAaPQD▼erda
■082 ▼a006.696
■1001 ▼aChen, Yu-Chih Berrie▼eauthor.
■24510▼aAdvancements in Perceptual Quality Assessment for Interactive Media: From Mobile Cloud Gaming to Human Avatar Videos and Facial Expressions ▼cYu-Chih Berrie Chen
■260 ▼a[Sl]▼bThe University of Texas at Austin▼c2025
■264 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2025
■300 ▼a1 electronic resource (144 pages)
■336 ▼atext▼btxt▼2rdacontent
■337 ▼acomputer▼bc▼2rdamedia
■338 ▼aonline resource▼bcr▼2rdacarrier
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-06, Section: A.
■500 ▼aAdvisors: Bovik, Alan C. Committee members: Kim, Hyeji; Geisler, Wilson S., III; de Veciana, Gustavo; Vikalo, Haris.
■5021 ▼bPh.D.▼cThe University of Texas at Austin▼d2025.
■520 ▼aOver the past decade, interactive media-including mobile cloud gaming and virtual reality (VR) applications-have grown rapidly, necessitating efficient methods to assess and optimize perceptual visual quality. This dissertation presents a comprehensive study of human visual quality judgments across three major domains: gaming video streaming, rendered human avatars in VR/AR, and facial expression similarity in human faces. As streaming these media becomes increasingly prevalent, advanced video compression protocols and Video Quality Assessment (VQA) algorithms are crucial for balancing high-quality visual delivery and variable bandwidth conditions. In the realm of mobile cloud gaming, the computational load of rendering games is offloaded to remote servers, allowing users to play on lightweight devices like smartphones. However, this introduces unique challenges in latency, compression, and real-time distortion. To address this, we introduce GAMIVAL, a no-reference (NR) VQA model tailored for gaming content, which exhibits statistical characteristics distinct from naturalistic videos. GAMIVAL models spatial and temporal distortions, neural noise, and deep semantic features, achieving superior performance on the LIVE-Meta Mobile Cloud Gaming video quality database. Expanding to immersive VR-based experiences, we investigate the perceived quality of rendered human avatars-often referred to as "holograms." We introduce the LIVE-Meta Rendered Human Avatar VQA Database, which includes 720 avatar videos compressed with 20 encoding parameter combinations and annotated with human perceptual quality ratings collected via six-degrees-of-freedom VR headsets. This large-scale database enables the benchmarking of both full-reference and no-reference VQA models, including a newly proposed model, HoloQA. Our study addresses a critical gap in current VQA research: the lack of a dataset designed specifically for evaluating rendered human body avatar videos. Moreover, we explore perceptual similarity in human facial expressions, a key component of affective computing. Existing metrics often fail to capture the nuanced semantics of expressions. To address this, we propose FaceSIM, a novel similarity model built on contrastive learning principles and vision transformer backbones. The architecture comprises two complementary branches: a fixed content-aware encoder that extracts stable, identity-agnostic facial features, and a trainable similarity-aware encoder that learns to capture perceptual distinctions through contrastive supervision. FaceSIM is able to separate identities while modeling nuanced expression differences. A simplified one-way contrastive objective guides the training, encouraging perceptually similar facial expressions to cluster in the embedding space while pushing dissimilar ones apart. The resulting embeddings support both zero-shot and regression-based similarity estimation. Evaluations on a large-scale facial expression similarity dataset demonstrate that FaceSIM aligns closely with human perception across diverse expressions, identities, and intensities. Together, these contributions form a unified framework for understanding and improving human-centered visual quality assessment across a range of interactive media. By developing novel models and datasets, this dissertation enhances the perceptual fidelity of cloud gaming, immersive holographic communication, and facial expression reconstruction-advancing multimedia technologies for next-generation user experiences.
■546 ▼aEnglish
■590 ▼aSchool code: 0227
■650 4▼aComputer engineering
■650 4▼aInformation technology
■650 4▼aFilm studies
■653 ▼aVirtual reality
■653 ▼aVideo Quality Assessment
■653 ▼aInteractive media
■653 ▼aMobile cloud gaming
■653 ▼aAvatar videos
■7102 ▼aThe University of Texas at Austin▼bElectrical and Computer Engineering.▼edegree granting institution.
■7201 ▼aBovik, Alan C.▼edegree supervisor.
■7730 ▼tDissertations Abstracts International▼g87-06A.
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17361231▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


