본문

서브메뉴

Deep Learning Applied to Image and Video Processing
Deep Learning Applied to Image and Video Processing
Deep Learning Applied to Image and Video Processing

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211150955
ISBN  
9798381977776
DDC  
004
저자명  
Wang, Xijun.
서명/저자  
Deep Learning Applied to Image and Video Processing
발행사항  
[Sl] : Northwestern University, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
148 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-10, Section: B.
주기사항  
Advisor: Katsaggelos, Aggelos.
학위논문주기  
Thesis (Ph.D.)--Northwestern University, 2024.
초록/해제  
요약This dissertation introduces deep learning (DL) methods applied to image and video processing, specifically concentrating on two domains: image and video restoration, and video classification with action localization.In the restoration domain, we introduce several innovative deep learning methodologies to address three challenges: super-resolution (SR), atmospheric turbulence (AT) correction, and motion blur (MB) removal.Generative Adversarial Networks (GANs) have demonstrated impressive performance in addressing super-resolution challenges, because of their ability to produce visually realistic images and video frames. However, previous GAN-based models frequently suffer from undesired side effects in their outputs, such as unexpected artifacts and noise. To mitigate these artifacts and enhance the perceptual quality of the results, in Chapter 1, we propose a general method that can be effectively used in most GAN-based super-resolution models by integrating essential spatial information into the training process. We extract spatial information from the input data and integrate it into the training loss, making the corresponding loss a spatially adaptive (SA) one. We show that the proposed approach is independent of the methods employed for spatial information extraction, as well as independent of SR tasks and models. This method consistently guides the training process towards generating visually pleasing SR images and video frames, substantially reducing artifacts and noise, and ultimately leading to enhanced perceptual quality. Besides including the spatial information through training loss, we also discover incorporating it through the model framework in Chapter 2. We design a new framework that incorporates two collaborative discriminators whose aim is to jointly improve the quality of the reconstructed video sequence. While one discriminator focuses on the general properties of the images, the second one specializes in obtaining realistically reconstructed features, such as edges. Experimental results demonstrate that the learned model outperforms current state-of-the-art models, yielding super-resolved frames with fine details, sharp edges, and reduced artifacts.Atmospheric turbulence, a common phenomenon in daily life, arises primarily due to the uneven heating of the Earth's surface. As a result, it causes distortion and blurring in acquired images or videos, significantly affecting downstream vision tasks, especially those dependent on capturing clear, stable images or videos from outdoor environments, such as accurate object detection or recognition. It is a challenging restoration task as it consists of two types of distortions: geometric distortion and spatially variant blur. In Chapter 3, we first propose a variational inference framework AT mitigation baseline, wherein we improve the performance by learning latent prior information from the input and degradation processes. Then we design a novel deep conditional diffusion model within the variational inference framework to further enhance the perceptual quality of output images. We demonstrate that the proposed framework achieves good quantitative and qualitative results on a comprehensive synthetic AT dataset. Though existing deep learning-based methods have achieved great performance in synthetic scenarios, they invariably exhibit a performance drop when applied to real-world cases. Therefore, in Chapter 4 we further propose a real-world atmospheric turbulence mitigation method under a domain adaptation framework, which connects supervised simulated atmospheric turbulence correction with unsupervised real-world atmospheric turbulence correction. We will show our proposed method enhances performance in real-world atmospheric turbulence scenarios, improving both image quality and downstream vision tasks.Respiratory motion and the resulting artifacts are considered to be a big problem in abdominal Magnetic Resonance Imaging (MRI). Many previous deep-learning techniques have been developed to address these respiratory motion artifacts. However, many models tend to oversmooth fine details, such as vessels in liver MRIs, while these details are most important for medical diagnosis. Thus, in Chapter 5, similar to our approach in SR, we propose a Generative Adversarial Networks (GAN)-based model for removing motion blur in abdominal MRI. We incorporate perceptual loss as part of our training loss to further enhance the perceptual quality of the images. Our model generates motion-reduced images with clearer and better fine-details, thereby providing radiologists with more realistic MRI images to aid in diagnosis.For the classification and action localization with videos, we present deep learning methods for solving avian-solar activity classification and weakly supervised action localization.Activity classification is essential in various real-life scenarios involving both humans and animals. The demand for precise activity classification concerning avian-solar interactions is rising, as the usage of solar energy facilities, such as photovoltaic array power stations, has been observed to impact bird species richness, behavior, and activity. However, there has been no effort to develop an automated system for monitoring and classifying avian-solar interactions. Current methods depend on human observers, which are time-consuming, resource-intensive, and prone to errors related to searcher efficiency. With the recent success of Deep Learning models in activity classification, in Chapter 6, we introduce a recurrent neural network-based model for automatically classifying six avian activities around solar energy facilities. Our model integrates crucial feature engineering metadata with video frame data, facilitating enhanced learning and more accurate activity classification. Furthermore, we address the challenge of data imbalance during training and demonstrate our model's effectiveness in detecting and classifying various activities within video tracks. Additionally, we analyze the saliency/backpropagation map of the trained proposed model and validate its decision-making rationale.Weakly-supervised temporal action localization aims to identify and localize the action instances in untrimmed videos with only video-level action labels. Humans can adapt abstract-level knowledge about actions in various video scenarios and detect the occurrence of actions. In Chapter 7, we mimic how humans do and introduce a new perspective for locating and identifying multiple actions in a video. We propose a network named VQK-Net with a video-specific query-key attention modeling, which learns a unique query for each action category in every input video. These learned queries encapsulate abstract-level features of actions and are capable of adapting this knowledge to the target video scenario, facilitating the detection of corresponding actions along the temporal dimension. To enhance the learning of these action category queries, we leverage not only the features of the current input video but also the correlations between different videos using a novel video-specific action category query learner worked with a query similarity loss. Finally, we conduct extensive experiments on three widely adopted datasets, achieving state-of-the-art performance.
일반주제명  
Computer science
일반주제명  
Computer engineering
일반주제명  
Information technology
키워드  
Spatially adaptive
키워드  
Atmospheric turbulence
키워드  
Deep learning
키워드  
Generative Adversarial Networks
키워드  
Videos processing
기타저자  
Northwestern University Computer Science
기본자료저록  
Dissertations Abstracts International. 85-10B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017160311
■00520250211150955
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798381977776
■035    ▼a(MiAaPQ)AAI30993511
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aWang,  Xijun.
■24510▼aDeep  Learning  Applied  to  Image  and  Video  Processing
■260    ▼a[Sl]▼bNorthwestern  University▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a148  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-10,  Section:  B.
■500    ▼aAdvisor:  Katsaggelos,  Aggelos.
■5021  ▼aThesis  (Ph.D.)--Northwestern  University,  2024.
■520    ▼aThis  dissertation  introduces  deep  learning  (DL)  methods  applied  to  image  and  video  processing,  specifically  concentrating  on  two  domains:  image  and  video  restoration,  and  video  classification  with  action  localization.In  the  restoration  domain,  we  introduce  several  innovative  deep  learning  methodologies  to  address  three  challenges:  super-resolution  (SR),  atmospheric  turbulence  (AT)  correction,  and  motion  blur  (MB)  removal.Generative  Adversarial  Networks  (GANs)  have  demonstrated  impressive  performance  in  addressing  super-resolution  challenges,  because  of  their  ability  to  produce  visually  realistic  images  and  video  frames.  However,  previous  GAN-based  models  frequently  suffer  from  undesired  side  effects  in  their  outputs,  such  as  unexpected  artifacts  and  noise.  To  mitigate  these  artifacts  and  enhance  the  perceptual  quality  of  the  results,  in  Chapter  1,  we  propose  a  general  method  that  can  be  effectively  used  in  most  GAN-based  super-resolution  models  by  integrating  essential  spatial  information  into  the  training  process.  We  extract  spatial  information  from  the  input  data  and  integrate  it  into  the  training  loss,  making  the  corresponding  loss  a  spatially  adaptive  (SA)  one.  We  show  that  the  proposed  approach  is  independent  of  the  methods  employed  for  spatial  information  extraction,  as  well  as  independent  of  SR  tasks  and  models.  This  method  consistently  guides  the  training  process  towards  generating  visually  pleasing  SR  images  and  video  frames,  substantially  reducing  artifacts  and  noise,  and  ultimately  leading  to  enhanced  perceptual  quality.  Besides  including  the  spatial  information  through  training  loss,  we  also  discover  incorporating  it  through  the  model  framework  in  Chapter  2.  We  design  a  new  framework  that  incorporates  two  collaborative  discriminators  whose  aim  is  to  jointly  improve  the  quality  of  the  reconstructed  video  sequence.  While  one  discriminator  focuses  on  the  general  properties  of  the  images,  the  second  one  specializes  in  obtaining  realistically  reconstructed  features,  such  as  edges.  Experimental  results  demonstrate  that  the  learned  model  outperforms  current  state-of-the-art  models,  yielding  super-resolved  frames  with  fine  details,  sharp  edges,  and  reduced  artifacts.Atmospheric  turbulence,  a  common  phenomenon  in  daily  life,  arises  primarily  due  to  the  uneven  heating  of  the  Earth's  surface.  As  a  result,  it  causes  distortion  and  blurring  in  acquired  images  or  videos,  significantly  affecting  downstream  vision  tasks,  especially  those  dependent  on  capturing  clear,  stable  images  or  videos  from  outdoor  environments,  such  as  accurate  object  detection  or  recognition.  It  is  a  challenging  restoration  task  as  it  consists  of  two  types  of  distortions:  geometric  distortion  and  spatially  variant  blur.  In  Chapter  3,  we  first  propose  a  variational  inference  framework  AT  mitigation  baseline,  wherein  we  improve  the  performance  by  learning  latent  prior  information  from  the  input  and  degradation  processes.  Then  we  design  a  novel  deep  conditional  diffusion  model  within  the  variational  inference  framework  to  further  enhance  the  perceptual  quality  of  output  images.  We  demonstrate  that  the  proposed  framework  achieves  good  quantitative  and  qualitative  results  on  a  comprehensive  synthetic  AT  dataset.  Though  existing  deep  learning-based  methods  have  achieved  great  performance  in  synthetic  scenarios,  they  invariably  exhibit  a  performance  drop  when  applied  to  real-world  cases.  Therefore,  in  Chapter  4  we  further  propose  a  real-world  atmospheric  turbulence  mitigation  method  under  a  domain  adaptation  framework,  which  connects  supervised  simulated  atmospheric  turbulence  correction  with  unsupervised  real-world  atmospheric  turbulence  correction.  We  will  show  our  proposed  method  enhances  performance  in  real-world  atmospheric  turbulence  scenarios,  improving  both  image  quality  and  downstream  vision  tasks.Respiratory  motion  and  the  resulting  artifacts  are  considered  to  be  a  big  problem  in  abdominal  Magnetic  Resonance  Imaging  (MRI).  Many  previous  deep-learning  techniques  have  been  developed  to  address  these  respiratory  motion  artifacts.  However,  many  models  tend  to  oversmooth  fine  details,  such  as  vessels  in  liver  MRIs,  while  these  details  are  most  important  for  medical  diagnosis.  Thus,  in  Chapter  5,  similar  to  our  approach  in  SR,  we  propose  a  Generative  Adversarial  Networks  (GAN)-based  model  for  removing  motion  blur  in  abdominal  MRI.  We  incorporate  perceptual  loss  as  part  of  our  training  loss  to  further  enhance  the  perceptual  quality  of  the  images.  Our  model  generates  motion-reduced  images  with  clearer  and  better  fine-details,  thereby  providing  radiologists  with  more  realistic  MRI  images  to  aid  in  diagnosis.For  the  classification  and  action  localization  with  videos,  we  present  deep  learning  methods  for  solving  avian-solar  activity  classification  and  weakly  supervised  action  localization.Activity  classification  is  essential  in  various  real-life  scenarios  involving  both  humans  and  animals.  The  demand  for  precise  activity  classification  concerning  avian-solar  interactions  is  rising,  as  the  usage  of  solar  energy  facilities,  such  as  photovoltaic  array  power  stations,  has  been  observed  to  impact  bird  species  richness,  behavior,  and  activity.  However,  there  has  been  no  effort  to  develop  an  automated  system  for  monitoring  and  classifying  avian-solar  interactions.  Current  methods  depend  on  human  observers,  which  are  time-consuming,  resource-intensive,  and  prone  to  errors  related  to  searcher  efficiency.  With  the  recent  success  of  Deep  Learning  models  in  activity  classification,  in  Chapter  6,  we  introduce  a  recurrent  neural  network-based  model  for  automatically  classifying  six  avian  activities  around  solar  energy  facilities.  Our  model  integrates  crucial  feature  engineering  metadata  with  video  frame  data,  facilitating  enhanced  learning  and  more  accurate  activity  classification.  Furthermore,  we  address  the  challenge  of  data  imbalance  during  training  and  demonstrate  our  model's  effectiveness  in  detecting  and  classifying  various  activities  within  video  tracks.  Additionally,  we  analyze  the  saliency/backpropagation  map  of  the  trained  proposed  model  and  validate  its  decision-making  rationale.Weakly-supervised  temporal  action  localization  aims  to  identify  and  localize  the  action  instances  in  untrimmed  videos  with  only  video-level  action  labels.  Humans  can  adapt  abstract-level  knowledge  about  actions  in  various  video  scenarios  and  detect  the  occurrence  of  actions.  In  Chapter  7,  we  mimic  how  humans  do  and  introduce  a  new  perspective  for  locating  and  identifying  multiple  actions  in  a  video.  We  propose  a  network  named  VQK-Net  with  a  video-specific  query-key  attention  modeling,  which  learns  a  unique  query  for  each  action  category  in  every  input  video.  These  learned  queries  encapsulate  abstract-level  features  of  actions  and  are  capable  of  adapting  this  knowledge  to  the  target  video  scenario,  facilitating  the  detection  of  corresponding  actions  along  the  temporal  dimension.  To  enhance  the  learning  of  these  action  category  queries,  we  leverage  not  only  the  features  of  the  current  input  video  but  also  the  correlations  between  different  videos  using  a  novel  video-specific  action  category  query  learner  worked  with  a  query  similarity  loss.  Finally,  we  conduct  extensive  experiments  on  three  widely  adopted  datasets,  achieving  state-of-the-art  performance.
■590    ▼aSchool  code:  0163.
■650  4▼aComputer  science
■650  4▼aComputer  engineering
■650  4▼aInformation  technology
■653    ▼aSpatially  adaptive
■653    ▼aAtmospheric  turbulence
■653    ▼aDeep  learning
■653    ▼aGenerative  Adversarial  Networks
■653    ▼aVideos    processing
■690    ▼a0984
■690    ▼a0489
■690    ▼a0464
■71020▼aNorthwestern  University▼bComputer  Science.
■7730  ▼tDissertations  Abstracts  International▼g85-10B.
■790    ▼a0163
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17160311▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF13912 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.