서브메뉴
검색
Less Is More: Accelerating Vision by Eliminating Redundancy
Less Is More: Accelerating Vision by Eliminating Redundancy
상세정보
- 자료유형
- 학위논문 서양
- 최종처리일시
- 20260202105557
- ISBN
- 9798263395384
- DDC
- 006.31
- 저자명
- Bolya, Daniel.
- 서명/저자
- Less Is More: Accelerating Vision by Eliminating Redundancy
- 발행사항
- [Sl] : Georgia Institute of Technology, 2024
- 발행사항
- Ann Arbor : ProQuest Dissertations & Theses, 2024
- 형태사항
- 318 p
- 주기사항
- Source: Dissertations Abstracts International, Volume: 87-06, Section: B.
- 주기사항
- Advisor: Hoffman, Judy.
- 학위논문주기
- Thesis (Ph.D.)--Georgia Institute of Technology, 2024.
- 초록/해제
- 요약The key to modern machine learning is scale. With more data, bigger models, and more compute, as a community, we've found that the problems once deemed impossible for a computer to solve have rapidly become attainable-many even becoming easy with today's techniques. But as the scale of modern machine learning has ballooned, so too has its cost. Large transformer models, for instance, can require multiple hundreds of GPUs to train effectively and can be similarly unwieldy to deploy. In this dissertation, I aim to reduce those costs.Specifically, this work focuses on Vision Transformers (ViTs), which have been the dominant driving force in scaling machine learning for computer vision. Over the course of this dissertation, I show that these ViTs perform redundant computation, and that by exploiting these redundancies, we can greatly increase the efficiency of these systems, both during training and inference. In Part I, I show that we can reduce the amount of spatial computation required by these transformers without losing performance. In Part II, I show that certain architectural components are redundant and can be removed or greatly simplified. In Part III, I show how we can exploit redundant features within models to speed them up and to circumvent training. Finally, in Part IV, I show that these speed-ups can compound on each other, resulting in a much faster model.With the techniques presented in this work, I hope to greatly reduce the cost of modern vision models-both to make these models easier to use in practice and to enable the community to scale these models even further beyond what we can do now. Because sometimes, less is more.
- 키워드
- Machine learning
- 기본자료저록
- Dissertations Abstracts International. 87-06B.
- 전자적 위치 및 접속
- 로그인 후 원문을 볼 수 있습니다.
MARC
008260126s2024 us c eng d■001000017360627
■00520260202105557
■006m o d
■007cr#unu||||||||
■020 ▼a9798263395384
■035 ▼a(MiAaPQ)AAI32315911
■035 ▼a(MiAaPQ)GeorgiaTech75233
■040 ▼aMiAaPQ▼cMiAaPQ
■0820 ▼a006.31
■1001 ▼aBolya, Daniel.
■24510▼aLess Is More: Accelerating Vision by Eliminating Redundancy
■260 ▼a[Sl]▼bGeorgia Institute of Technology▼c2024
■260 1▼aAnn Arbor▼bProQuest Dissertations & Theses▼c2024
■300 ▼a318 p
■500 ▼aSource: Dissertations Abstracts International, Volume: 87-06, Section: B.
■500 ▼aAdvisor: Hoffman, Judy.
■5021 ▼aThesis (Ph.D.)--Georgia Institute of Technology, 2024.
■520 ▼aThe key to modern machine learning is scale. With more data, bigger models, and more compute, as a community, we've found that the problems once deemed impossible for a computer to solve have rapidly become attainable-many even becoming easy with today's techniques. But as the scale of modern machine learning has ballooned, so too has its cost. Large transformer models, for instance, can require multiple hundreds of GPUs to train effectively and can be similarly unwieldy to deploy. In this dissertation, I aim to reduce those costs.Specifically, this work focuses on Vision Transformers (ViTs), which have been the dominant driving force in scaling machine learning for computer vision. Over the course of this dissertation, I show that these ViTs perform redundant computation, and that by exploiting these redundancies, we can greatly increase the efficiency of these systems, both during training and inference. In Part I, I show that we can reduce the amount of spatial computation required by these transformers without losing performance. In Part II, I show that certain architectural components are redundant and can be removed or greatly simplified. In Part III, I show how we can exploit redundant features within models to speed them up and to circumvent training. Finally, in Part IV, I show that these speed-ups can compound on each other, resulting in a much faster model.With the techniques presented in this work, I hope to greatly reduce the cost of modern vision models-both to make these models easier to use in practice and to enable the community to scale these models even further beyond what we can do now. Because sometimes, less is more.
■590 ▼aSchool code: 0078.
■653 ▼aMachine learning
■690 ▼a0800
■71020▼aGeorgia Institute of Technology.
■7730 ▼tDissertations Abstracts International▼g87-06B.
■790 ▼a0078
■791 ▼aPh.D.
■792 ▼a2024
■793 ▼aEnglish
■85640▼uhttp://www.riss.kr/pdu/ddodLink.do?id=T17360627▼nKERIS▼z이 자료의 원문은 한국교육학술정보원에서 제공합니다.


