서브메뉴
검색
Deep Learning Applied to Image and Video Processing
Deep Learning Applied to Image and Video Processing
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20250211150955
- ISBN
- 9798381977776
- DDC
- 004
- 저자명
- Wang, Xijun.
- 서명/저자
- Deep Learning Applied to Image and Video Processing
- 발행사항
- [Sl] : Northwestern University, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 148 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 85-10, Section: B.
- 주기사항
- Advisor: Katsaggelos, Aggelos.
- 학위논문주기
- Thesis (Ph.D.)--Northwestern University, 2024.
- 초록/해제
- 요약This dissertation introduces deep learning (DL) methods applied to image and video processing, specifically concentrating on two domains: image and video restoration, and video classification with action localization.In the restoration domain, we introduce several innovative deep learning methodologies to address three challenges: super-resolution (SR), atmospheric turbulence (AT) correction, and motion blur (MB) removal.Generative Adversarial Networks (GANs) have demonstrated impressive performance in addressing super-resolution challenges, because of their ability to produce visually realistic images and video frames. However, previous GAN-based models frequently suffer from undesired side effects in their outputs, such as unexpected artifacts and noise. To mitigate these artifacts and enhance the perceptual quality of the results, in Chapter 1, we propose a general method that can be effectively used in most GAN-based super-resolution models by integrating essential spatial information into the training process. We extract spatial information from the input data and integrate it into the training loss, making the corresponding loss a spatially adaptive (SA) one. We show that the proposed approach is independent of the methods employed for spatial information extraction, as well as independent of SR tasks and models. This method consistently guides the training process towards generating visually pleasing SR images and video frames, substantially reducing artifacts and noise, and ultimately leading to enhanced perceptual quality. Besides including the spatial information through training loss, we also discover incorporating it through the model framework in Chapter 2. We design a new framework that incorporates two collaborative discriminators whose aim is to jointly improve the quality of the reconstructed video sequence. While one discriminator focuses on the general properties of the images, the second one specializes in obtaining realistically reconstructed features, such as edges. Experimental results demonstrate that the learned model outperforms current state-of-the-art models, yielding super-resolved frames with fine details, sharp edges, and reduced artifacts.Atmospheric turbulence, a common phenomenon in daily life, arises primarily due to the uneven heating of the Earth's surface. As a result, it causes distortion and blurring in acquired images or videos, significantly affecting downstream vision tasks, especially those dependent on capturing clear, stable images or videos from outdoor environments, such as accurate object detection or recognition. It is a challenging restoration task as it consists of two types of distortions: geometric distortion and spatially variant blur. In Chapter 3, we first propose a variational inference framework AT mitigation baseline, wherein we improve the performance by learning latent prior information from the input and degradation processes. Then we design a novel deep conditional diffusion model within the variational inference framework to further enhance the perceptual quality of output images. We demonstrate that the proposed framework achieves good quantitative and qualitative results on a comprehensive synthetic AT dataset. Though existing deep learning-based methods have achieved great performance in synthetic scenarios, they invariably exhibit a performance drop when applied to real-world cases. Therefore, in Chapter 4 we further propose a real-world atmospheric turbulence mitigation method under a domain adaptation framework, which connects supervised simulated atmospheric turbulence correction with unsupervised real-world atmospheric turbulence correction. We will show our proposed method enhances performance in real-world atmospheric turbulence scenarios, improving both image quality and downstream vision tasks.Respiratory motion and the resulting artifacts are considered to be a big problem in abdominal Magnetic Resonance Imaging (MRI). Many previous deep-learning techniques have been developed to address these respiratory motion artifacts. However, many models tend to oversmooth fine details, such as vessels in liver MRIs, while these details are most important for medical diagnosis. Thus, in Chapter 5, similar to our approach in SR, we propose a Generative Adversarial Networks (GAN)-based model for removing motion blur in abdominal MRI. We incorporate perceptual loss as part of our training loss to further enhance the perceptual quality of the images. Our model generates motion-reduced images with clearer and better fine-details, thereby providing radiologists with more realistic MRI images to aid in diagnosis.For the classification and action localization with videos, we present deep learning methods for solving avian-solar activity classification and weakly supervised action localization.Activity classification is essential in various real-life scenarios involving both humans and animals. The demand for precise activity classification concerning avian-solar interactions is rising, as the usage of solar energy facilities, such as photovoltaic array power stations, has been observed to impact bird species richness, behavior, and activity. However, there has been no effort to develop an automated system for monitoring and classifying avian-solar interactions. Current methods depend on human observers, which are time-consuming, resource-intensive, and prone to errors related to searcher efficiency. With the recent success of Deep Learning models in activity classification, in Chapter 6, we introduce a recurrent neural network-based model for automatically classifying six avian activities around solar energy facilities. Our model integrates crucial feature engineering metadata with video frame data, facilitating enhanced learning and more accurate activity classification. Furthermore, we address the challenge of data imbalance during training and demonstrate our model's effectiveness in detecting and classifying various activities within video tracks. Additionally, we analyze the saliency/backpropagation map of the trained proposed model and validate its decision-making rationale.Weakly-supervised temporal action localization aims to identify and localize the action instances in untrimmed videos with only video-level action labels. Humans can adapt abstract-level knowledge about actions in various video scenarios and detect the occurrence of actions. In Chapter 7, we mimic how humans do and introduce a new perspective for locating and identifying multiple actions in a video. We propose a network named VQK-Net with a video-specific query-key attention modeling, which learns a unique query for each action category in every input video. These learned queries encapsulate abstract-level features of actions and are capable of adapting this knowledge to the target video scenario, facilitating the detection of corresponding actions along the temporal dimension. To enhance the learning of these action category queries, we leverage not only the features of the current input video but also the correlations between different videos using a novel video-specific action category query learner worked with a query similarity loss. Finally, we conduct extensive experiments on three widely adopted datasets, achieving state-of-the-art performance.
- 일반주제명
- Computer science
- 일반주제명
- Computer engineering
- 일반주제명
- Information technology
- 키워드
- Deep learning
- 기타저자
- Northwestern University Computer Science
- 기본자료저록
- Dissertations Abstracts International. 85-10B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008250123s2024 us c eng d■001000017160311
■00520250211150955
■006m o d
■007cr#unu||||||||
■020 ▼a9798381977776
■035 ▼a(MiAaPQ)AAI30993511
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a004
■1001 ▼aWang, Xijun.
■24510▼aDeep Learning Applied to Image and Video Processing
■260 ▼a[Sl]▼bNorthwestern University▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a148 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 85-10, Section: B.
■500 ▼aAdvisor: Katsaggelos, Aggelos.
■5021 ▼aThesis (Ph.D.)--Northwestern University, 2024.
■520 ▼aThis dissertation introduces deep learning (DL) methods applied to image and video processing, specifically concentrating on two domains: image and video restoration, and video classification with action localization.In the restoration domain, we introduce several innovative deep learning methodologies to address three challenges: super-resolution (SR), atmospheric turbulence (AT) correction, and motion blur (MB) removal.Generative Adversarial Networks (GANs) have demonstrated impressive performance in addressing super-resolution challenges, because of their ability to produce visually realistic images and video frames. However, previous GAN-based models frequently suffer from undesired side effects in their outputs, such as unexpected artifacts and noise. To mitigate these artifacts and enhance the perceptual quality of the results, in Chapter 1, we propose a general method that can be effectively used in most GAN-based super-resolution models by integrating essential spatial information into the training process. We extract spatial information from the input data and integrate it into the training loss, making the corresponding loss a spatially adaptive (SA) one. We show that the proposed approach is independent of the methods employed for spatial information extraction, as well as independent of SR tasks and models. This method consistently guides the training process towards generating visually pleasing SR images and video frames, substantially reducing artifacts and noise, and ultimately leading to enhanced perceptual quality. Besides including the spatial information through training loss, we also discover incorporating it through the model framework in Chapter 2. We design a new framework that incorporates two collaborative discriminators whose aim is to jointly improve the quality of the reconstructed video sequence. While one discriminator focuses on the general properties of the images, the second one specializes in obtaining realistically reconstructed features, such as edges. Experimental results demonstrate that the learned model outperforms current state-of-the-art models, yielding super-resolved frames with fine details, sharp edges, and reduced artifacts.Atmospheric turbulence, a common phenomenon in daily life, arises primarily due to the uneven heating of the Earth's surface. As a result, it causes distortion and blurring in acquired images or videos, significantly affecting downstream vision tasks, especially those dependent on capturing clear, stable images or videos from outdoor environments, such as accurate object detection or recognition. It is a challenging restoration task as it consists of two types of distortions: geometric distortion and spatially variant blur. In Chapter 3, we first propose a variational inference framework AT mitigation baseline, wherein we improve the performance by learning latent prior information from the input and degradation processes. Then we design a novel deep conditional diffusion model within the variational inference framework to further enhance the perceptual quality of output images. We demonstrate that the proposed framework achieves good quantitative and qualitative results on a comprehensive synthetic AT dataset. Though existing deep learning-based methods have achieved great performance in synthetic scenarios, they invariably exhibit a performance drop when applied to real-world cases. Therefore, in Chapter 4 we further propose a real-world atmospheric turbulence mitigation method under a domain adaptation framework, which connects supervised simulated atmospheric turbulence correction with unsupervised real-world atmospheric turbulence correction. We will show our proposed method enhances performance in real-world atmospheric turbulence scenarios, improving both image quality and downstream vision tasks.Respiratory motion and the resulting artifacts are considered to be a big problem in abdominal Magnetic Resonance Imaging (MRI). Many previous deep-learning techniques have been developed to address these respiratory motion artifacts. However, many models tend to oversmooth fine details, such as vessels in liver MRIs, while these details are most important for medical diagnosis. Thus, in Chapter 5, similar to our approach in SR, we propose a Generative Adversarial Networks (GAN)-based model for removing motion blur in abdominal MRI. We incorporate perceptual loss as part of our training loss to further enhance the perceptual quality of the images. Our model generates motion-reduced images with clearer and better fine-details, thereby providing radiologists with more realistic MRI images to aid in diagnosis.For the classification and action localization with videos, we present deep learning methods for solving avian-solar activity classification and weakly supervised action localization.Activity classification is essential in various real-life scenarios involving both humans and animals. The demand for precise activity classification concerning avian-solar interactions is rising, as the usage of solar energy facilities, such as photovoltaic array power stations, has been observed to impact bird species richness, behavior, and activity. However, there has been no effort to develop an automated system for monitoring and classifying avian-solar interactions. Current methods depend on human observers, which are time-consuming, resource-intensive, and prone to errors related to searcher efficiency. With the recent success of Deep Learning models in activity classification, in Chapter 6, we introduce a recurrent neural network-based model for automatically classifying six avian activities around solar energy facilities. Our model integrates crucial feature engineering metadata with video frame data, facilitating enhanced learning and more accurate activity classification. Furthermore, we address the challenge of data imbalance during training and demonstrate our model's effectiveness in detecting and classifying various activities within video tracks. Additionally, we analyze the saliency/backpropagation map of the trained proposed model and validate its decision-making rationale.Weakly-supervised temporal action localization aims to identify and localize the action instances in untrimmed videos with only video-level action labels. Humans can adapt abstract-level knowledge about actions in various video scenarios and detect the occurrence of actions. In Chapter 7, we mimic how humans do and introduce a new perspective for locating and identifying multiple actions in a video. We propose a network named VQK-Net with a video-specific query-key attention modeling, which learns a unique query for each action category in every input video. These learned queries encapsulate abstract-level features of actions and are capable of adapting this knowledge to the target video scenario, facilitating the detection of corresponding actions along the temporal dimension. To enhance the learning of these action category queries, we leverage not only the features of the current input video but also the correlations between different videos using a novel video-specific action category query learner worked with a query similarity loss. Finally, we conduct extensive experiments on three widely adopted datasets, achieving state-of-the-art performance.
■590 ▼aSchool code: 0163.
■650 4▼aComputer science
■650 4▼aComputer engineering
■650 4▼aInformation technology
■653 ▼aSpatially adaptive
■653 ▼aAtmospheric turbulence
■653 ▼aDeep learning
■653 ▼aGenerative Adversarial Networks
■653 ▼aVideos processing
■690 ▼a0984
■690 ▼a0489
■690 ▼a0464
■71020▼aNorthwestern University▼bComputer Science.
■7730 ▼tDissertations Abstracts International▼g85-10B.
■790 ▼a0163
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160311▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


