GI LogoGI Logo
  • Anmelden
Digitale Bibliothek
    • Gesamter Bestand

      • Bereiche & Sammlungen
      • Titel
      • Autor
      • Erscheinungsdatum
      • Schlagwort
    • Diese Sammlung

      • Titel
      • Autor
      • Erscheinungsdatum
      • Schlagwort
Digital Bibliothek der Gesellschaft für Informatik e.V.
GI-DL
    • English
    • Deutsch
  • Deutsch 
    • English
    • Deutsch
Dokumentanzeige 
  •   Startseite
  • Fachbereiche
  • Technische Informatik (TI)
  • PARS-Mitteilungen
  • PARS-Mitteilungen 2020
  • Dokumentanzeige
JavaScript is disabled for your browser. Some features of this site may not work without it.
  •   Startseite
  • Fachbereiche
  • Technische Informatik (TI)
  • PARS-Mitteilungen
  • PARS-Mitteilungen 2020
  • Dokumentanzeige

Weight Pruning for Deep Neural Networks on GPUs

Autor(en):
Hartenstein, Thomas [DBLP] ;
Maier, Daniel [DBLP] ;
Cosenza, Biagio [DBLP] ;
Juurlink, Ben [DBLP]
Zusammenfassung
Neural networks are getting more complex than ever before, leading to resource-demanding training processes that have been the target of optimization. With embedded real-time applications such as traffic identification in self-driving cars relying on neural networks, the inference latency is becoming more important. The size of the model has been identified as an important target of optimization, as smaller networks also require less computations for inference. A way to shrink a network in size is to remove small weights: weight pruning. This technique has been exploited in a number of ways and has shown to be able to significantly lower the number of weights, while maintaining a very close accuracy compared to the original network. However, current pruning techniques require the removal of up to 90% of the weights, requiring high amount of redundancy in the original network, to be able to speedup the inference as sparse data structures induce overhead. We propose a novel technique for the selection of the weights to be pruned. Our technique is specifically designed to take the architecture of GPUs into account. By selecting the weights to be removed in adjacent groups that are aligned to the memory architecture, we are able to fully exploit the memory bandwidth. Our results show that with the same amount of weights removed, our technique is able to speedup a neural network by a factor of 1:57 given a pruning rate of 90% while maintaining the same accuracy when compared to state-of-the-art pruning techniques.
  • Vollständige Referenz
  • BibTeX
Hartenstein, T., Maier, D., Cosenza, B. & Juurlink, B., (2020). Weight Pruning for Deep Neural Networks on GPUs.   PARS-Mitteilungen: Vol. 35, Nr. 1. Berlin: Gesellschaft für Informatik e.V., Fachgruppe PARS. (S. 51-62).
@article{mci/Hartenstein2020,
author = {Hartenstein, Thomas AND Maier, Daniel AND Cosenza, Biagio AND Juurlink, Ben},
title = {Weight Pruning for Deep Neural Networks on GPUs},
journal = {PARS-Mitteilungen},
volume = {35},
number = {1},
year = {2020},
,
pages = { 51-62 }
}
DateienGroesseFormatAnzeige
PARS2019_paper_9.pdf422.6Kb PDF Öffnen

Haben Sie fehlerhafte Angaben entdeckt? Sagen Sie uns Bescheid: Feedback abschicken

Mehr Information

ISSN: 0177-0454
Datum: 2020
Sprache: en (en)
Typ: Text/Journal Article
Sammlungen
  • PARS-Mitteilungen 2020 [12]

Zur Langanzeige


Über uns | FAQ | Hilfe | Impressum | Datenschutz

Gesellschaft für Informatik e.V. (GI), Kontakt: Geschäftsstelle der GI
Diese Digital Library basiert auf DSpace.

 

 


Über uns | FAQ | Hilfe | Impressum | Datenschutz

Gesellschaft für Informatik e.V. (GI), Kontakt: Geschäftsstelle der GI
Diese Digital Library basiert auf DSpace.