GI LogoGI Logo
  • Anmelden
Digitale Bibliothek
    • Gesamter Bestand

      • Bereiche & Sammlungen
      • Titel
      • Autor
      • Erscheinungsdatum
      • Schlagwort
    • Diese Sammlung

      • Titel
      • Autor
      • Erscheinungsdatum
      • Schlagwort
Digital Bibliothek der Gesellschaft für Informatik e.V.
GI-DL
    • English
    • Deutsch
  • Deutsch 
    • English
    • Deutsch
Dokumentanzeige 
  •   Startseite
  • Lecture Notes in Informatics
  • Proceedings
  • P124 - 9th Workshop on Parallel Systems and Algorythms
  • Dokumentanzeige
JavaScript is disabled for your browser. Some features of this site may not work without it.
  •   Startseite
  • Lecture Notes in Informatics
  • Proceedings
  • P124 - 9th Workshop on Parallel Systems and Algorythms
  • Dokumentanzeige

An optimized ZGEMM implementation for the Cell BE

Autor(en):
Schneider, Timo [DBLP] ;
Hoefler, Torsten [DBLP] ;
Wunderlich, Simon [DBLP] ;
Mehlan, Torsten [DBLP] ;
Rehm, Wolfgang [DBLP]
Zusammenfassung
The architecture of the IBM Cell BE processor represents a new approach for designing CPUs. The fast execution of legacy software has to stand back in order to achieve very high performance for new scientific software. The Cell BE consists of 9 independent cores and represents a new promising architecture for HPC systems. The programmer has to write parallel software that is distributed to the cores and executes subtasks of the program in parallel. The simplified Vector-CPU design achieves higher clock-rates and power efficiency and exhibits predictable behavior. But to exploit the capabilities of this upcoming CPU architecture it is necessary to provide optimized libraries for frequently used algorithms. The Basic Linear Algebra Subprograms (BLAS) provide functions that are crucial for many scientific applications. The routine ZGEMM, which computes a complex matrix–matrix–product, is one of these functions. This article describes strategies to implement the ZGEMM routine on the Cell BE processor. The main goal is achieve highest performance. We compare this optimized ZGEMM implementation with several math libraries on Cell and other modern architectures. Thus we are able to show that our ZGEMM algorithm performs best in comparison to the fastest publicly available ZGEMM and DGEMM implementations for Cell BE and reasonably well in the league of other BLAS implementations.
  • Vollständige Referenz
  • BibTeX
Schneider, T., Hoefler, T., Wunderlich, S., Mehlan, T. & Rehm, W., (2008). An optimized ZGEMM implementation for the Cell BE. In: Nagel, W. E., Hoffmann, R. & Koch, A. (Hrsg.), 9th workshop on parallel systems and algorithms – workshop of the GI/ITG special interest groups PARS and PARVA. Bonn: Gesellschaft für Informatik e. V.. (S. 113-122).
@inproceedings{mci/Schneider2008,
author = {Schneider, Timo AND Hoefler, Torsten AND Wunderlich, Simon AND Mehlan, Torsten AND Rehm, Wolfgang},
title = {An optimized ZGEMM implementation for the Cell BE},
booktitle = {9th workshop on parallel systems and algorithms – workshop of the GI/ITG special interest groups PARS and PARVA},
year = {2008},
editor = {Nagel, Wolfgang E. AND Hoffmann, Rolf AND Koch, Andreas} ,
pages = { 113-122 },
publisher = {Gesellschaft für Informatik e. V.},
address = {Bonn}
}
DateienGroesseFormatAnzeige
113.pdf196.5Kb PDF Öffnen

Haben Sie fehlerhafte Angaben entdeckt? Sagen Sie uns Bescheid: Feedback abschicken

Mehr Information

ISBN: 978-3-88579-218-5
ISSN: 1617-5468
Datum: 2008
Sprache: en (en)
Typ: Text/Conference Paper
Sammlungen
  • P124 - 9th Workshop on Parallel Systems and Algorythms [12]

Zur Langanzeige


Über uns | FAQ | Hilfe | Impressum | Datenschutz

Gesellschaft für Informatik e.V. (GI), Kontakt: Geschäftsstelle der GI
Diese Digital Library basiert auf DSpace.

 

 


Über uns | FAQ | Hilfe | Impressum | Datenschutz

Gesellschaft für Informatik e.V. (GI), Kontakt: Geschäftsstelle der GI
Diese Digital Library basiert auf DSpace.