본문

서브메뉴

Improving Articulated Pose Tracking and Contact Force Estimation for Qualitative Assessment of Human Actions
Improving Articulated Pose Tracking and Contact Force Estimation for Qualitative Assessmen...
Improving Articulated Pose Tracking and Contact Force Estimation for Qualitative Assessment of Human Actions

상세정보

자료유형  
 학위논문 서양
최종처리일시  
20250211152051
ISBN  
9798382738338
DDC  
004
저자명  
Louis, Nathan.
서명/저자  
Improving Articulated Pose Tracking and Contact Force Estimation for Qualitative Assessment of Human Actions
발행사항  
[Sl] : University of Michigan, 2024
발행사항  
Ann Arbor : ProQuest Dissertations & Theses, 2024
형태사항  
121 p
주기사항  
Source: Dissertations Abstracts International, Volume: 85-12, Section: B.
주기사항  
Advisor: Corso, Jason J.;Owens, Andrew.
학위논문주기  
Thesis (Ph.D.)--University of Michigan, 2024.
초록/해제  
요약Using video to automate human performance metrics or skill analysis is an important but under-explored task. Currently, measuring the quality of an action can be highly subjective, where even assessments from experts are affected by bias and inter-rater reliability. In contrast, Computer vision and AI have the potential to provide real-time non-intrusive solutions with increased objectivity, scalability, and repeatability across various domains. From video alone, we can automatically provide supplemental objective scoring of Olympic sports, evaluate the technical skill of surgeons for training purposes, or monitor the physical rehabilitation progress of a patient. Today we solve these problems with supervised learning, obtaining features that represent high correlation with our desired point of analysis. Supervised learning is powerful, data-driven, and sometimes the best available option. However alone, it may be sub-optimal in the presence of scarce data and insufficient when needed to generalize to varying conditions or to truly understand the target task.In this dissertation, the bases of our human analysis understanding are skeletal poses, namely hand poses and full body poses. For articulated hand poses, we improve tracking using our CondPose network to integrate prior detection confidences and encourage tracking consistency. While for human poses, we propose two physical simulation-based metrics for evaluating physical plausibility and perform external force estimation through predicted ground reaction forces (GRFs). However, in the human analysis domain, collecting and annotating data at the scale of other deep learning tasks is a recurring challenge. This limits our generalizability to different environments, procedures, and motions. We address this by exploring semi-supervised learning methods, such as contrastive pre-training and multi-task learning.We apply articulated hand pose tracking in the surgical environment for assessing surgical skill. By applying a time-shifted sampling augmentation, we introduce clip-contrastive pre-training on embedded hand features as an unsupervised learning step. We show that this contrastive pre-training improves performance when fine-tuned on surgical skill classification and assessment task. Unlike most prior work, we evaluate on open surgery videos rather than solely simulated environments. Specifically, we use videos of non-laparoscopic, collected through collaboration with the Cardiac Surgery department at Michigan Medicine. We use full body poses and contact force estimation to bridge the gap between visual observations and the physical world. This physically-grounded component is vital for understanding actions involving sports or physical rehab where humans interact with their environment. We leverage multi-task learning to perform 2D-to-3D human pose estimation and integrate other abundant sources of motion capture data, without requiring additional force plate supervision. Our experiments shows that this improves GRF estimation on unseen motions. To address data limitations, we also collect two novel datasets SurgicalHands and ForcePose. We use SurgicalHands in the surgical domain as a multi-instance articulated hand pose tracking dataset. It encompasses a high degree of complexity in appearance and movement, not present in prior datasets. ForcePose is a multi-view GRF dataset of tracked human poses and time-synchronized force plates, to our knowledge the largest and most varied of its kind. This dataset serves as a benchmark for mapping human body motion and physical forces, enabling physical grounding of specific actions.
일반주제명  
Computer science
일반주제명  
Electrical engineering
일반주제명  
Surgery
일반주제명  
Bioinformatics
키워드  
CondPose network
키워드  
Physical rehabilitation
키워드  
Multi-task learning
키워드  
SurgicalHands
키워드  
ForcePose
기타저자  
University of Michigan Electrical and Computer Engineering
기본자료저록  
Dissertations Abstracts International. 85-12B.
전자적 위치 및 접속  
로그인 후 원문을 볼 수 있습니다.

MARC

 008250123s2024        us                              c    eng  d
■001000017162762
■00520250211152051
■006m          o    d                
■007cr#unu||||||||
■020    ▼a9798382738338
■035    ▼a(MiAaPQ)AAI31348864
■035    ▼a(MiAaPQ)umichrackham005486
■040    ▼aMiAaPQ▼cMiAaPQ
■0820  ▼a004
■1001  ▼aLouis,  Nathan.
■24510▼aImproving  Articulated  Pose  Tracking  and  Contact  Force  Estimation  for  Qualitative  Assessment  of  Human  Actions
■260    ▼a[Sl]▼bUniversity  of  Michigan▼c2024
■260  1▼aAnn  Arbor▼bProQuest  Dissertations  &  Theses▼c2024
■300    ▼a121  p
■500    ▼aSource:  Dissertations  Abstracts  International,  Volume:  85-12,  Section:  B.
■500    ▼aAdvisor:  Corso,  Jason  J.;Owens,  Andrew.
■5021  ▼aThesis  (Ph.D.)--University  of  Michigan,  2024.
■520    ▼aUsing  video  to  automate  human  performance  metrics  or  skill  analysis  is  an  important  but  under-explored  task.  Currently,  measuring  the  quality  of  an  action  can  be  highly  subjective,  where  even  assessments  from  experts  are  affected  by  bias  and  inter-rater  reliability.  In  contrast,  Computer  vision  and  AI  have  the  potential  to  provide  real-time  non-intrusive  solutions  with  increased  objectivity,  scalability,  and  repeatability  across  various  domains.  From  video  alone,  we  can  automatically  provide  supplemental  objective  scoring  of  Olympic  sports,  evaluate  the  technical  skill  of  surgeons  for  training  purposes,  or  monitor  the  physical  rehabilitation  progress  of  a  patient.  Today  we  solve  these  problems  with  supervised  learning,  obtaining  features  that  represent  high  correlation  with  our  desired  point  of  analysis.  Supervised  learning  is  powerful,  data-driven,  and  sometimes  the  best  available  option.  However  alone,  it  may  be  sub-optimal  in  the  presence  of  scarce  data  and  insufficient  when  needed  to  generalize  to  varying  conditions  or  to  truly  understand  the  target  task.In  this  dissertation,  the  bases  of  our  human  analysis  understanding  are  skeletal  poses,  namely  hand  poses  and  full  body  poses.  For  articulated  hand  poses,  we  improve  tracking  using  our  CondPose  network  to  integrate  prior  detection  confidences  and  encourage  tracking  consistency.  While  for  human  poses,  we  propose  two  physical  simulation-based  metrics  for  evaluating  physical  plausibility  and  perform  external  force  estimation  through  predicted  ground  reaction  forces  (GRFs).  However,  in  the  human  analysis  domain,  collecting  and  annotating  data  at  the  scale  of  other  deep  learning  tasks  is  a  recurring  challenge.  This  limits  our  generalizability  to  different  environments,  procedures,  and  motions.  We  address  this  by  exploring  semi-supervised  learning  methods,  such  as  contrastive  pre-training  and  multi-task  learning.We  apply  articulated  hand  pose  tracking  in  the  surgical  environment  for  assessing  surgical  skill.  By  applying  a  time-shifted  sampling  augmentation,  we  introduce  clip-contrastive  pre-training  on  embedded  hand  features  as  an  unsupervised  learning  step.  We  show  that  this  contrastive  pre-training  improves  performance  when  fine-tuned  on  surgical  skill  classification  and  assessment  task.  Unlike  most  prior  work,  we  evaluate  on  open  surgery  videos  rather  than  solely  simulated  environments.  Specifically,  we  use  videos  of  non-laparoscopic,  collected  through  collaboration  with  the  Cardiac  Surgery  department  at  Michigan  Medicine. We  use  full  body  poses  and  contact  force  estimation  to  bridge  the  gap  between  visual  observations  and  the  physical  world.  This  physically-grounded  component  is  vital  for  understanding  actions  involving  sports  or  physical  rehab  where  humans  interact  with  their  environment.  We  leverage  multi-task  learning  to  perform  2D-to-3D  human  pose  estimation  and  integrate  other  abundant  sources  of  motion  capture  data,  without  requiring  additional  force  plate  supervision.  Our  experiments  shows  that  this  improves  GRF  estimation  on  unseen  motions. To  address  data  limitations,  we  also  collect  two  novel  datasets  SurgicalHands  and  ForcePose.  We  use  SurgicalHands  in  the  surgical  domain  as  a  multi-instance  articulated  hand  pose  tracking  dataset.  It  encompasses  a  high  degree  of  complexity  in  appearance  and  movement,  not  present  in  prior  datasets.  ForcePose  is  a  multi-view  GRF  dataset  of  tracked  human  poses  and  time-synchronized  force  plates,  to  our  knowledge  the  largest  and  most  varied  of  its  kind.  This  dataset  serves  as  a  benchmark  for  mapping  human  body  motion  and  physical  forces,  enabling  physical  grounding  of  specific  actions.
■590    ▼aSchool  code:  0127.
■650  4▼aComputer  science
■650  4▼aElectrical  engineering
■650  4▼aSurgery
■650  4▼aBioinformatics
■653    ▼aCondPose  network
■653    ▼aPhysical  rehabilitation
■653    ▼aMulti-task  learning
■653    ▼aSurgicalHands  
■653    ▼aForcePose
■690    ▼a0544
■690    ▼a0984
■690    ▼a0576
■690    ▼a0800
■690    ▼a0715
■71020▼aUniversity  of  Michigan▼bElectrical  and  Computer  Engineering.
■7730  ▼tDissertations  Abstracts  International▼g85-12B.
■790    ▼a0127
■791    ▼aPh.D.
■792    ▼a2024
■793    ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17162762▼nKERIS▼z이  자료의  원문은  한국교육학술정보원에서  제공합니다.

미리보기

내보내기

chatGPT토론

Ai 추천 관련 도서


    신착도서 더보기
    최근 3년간 통계입니다.

    소장정보

    • 예약
    • 소재불명신고
    • 나의폴더
    • 우선정리요청
    • 비도서대출신청
    • 야간 도서대출신청
    소장자료
    등록번호 청구기호 소장처 대출가능여부 대출정보
    TF09485 전자도서 대출가능 마이폴더 부재도서신고 비도서대출신청 야간 도서대출신청

    * 대출중인 자료에 한하여 예약이 가능합니다. 예약을 원하시면 예약버튼을 클릭하십시오.

    해당 도서를 다른 이용자가 함께 대출한 도서

    관련 인기도서

    로그인 후 이용 가능합니다.