<?xml version="1.0" encoding="UTF-8"?><rdf:RDF xmlns="http://purl.org/rss/1.0/" xmlns:rdf="http://www.w3.org/1999/02/22-rdf-syntax-ns#" xmlns:dc="http://purl.org/dc/elements/1.1/">
<channel rdf:about="http://dl.gi.de/handle/20.500.12116/1912">
<title>PARS-Mitteilungen 2013</title>
<link>http://dl.gi.de/handle/20.500.12116/1912</link>
<description/>
<items>
<rdf:Seq>
<rdf:li rdf:resource="http://dl.gi.de/handle/20.500.12116/8608"/>
<rdf:li rdf:resource="http://dl.gi.de/handle/20.500.12116/8605"/>
<rdf:li rdf:resource="http://dl.gi.de/handle/20.500.12116/8607"/>
<rdf:li rdf:resource="http://dl.gi.de/handle/20.500.12116/8606"/>
</rdf:Seq>
</items>
<dc:date>2026-07-23T19:18:43Z</dc:date>
</channel>
<item rdf:about="http://dl.gi.de/handle/20.500.12116/8608">
<title>Self-organizing Core Allocation</title>
<link>http://dl.gi.de/handle/20.500.12116/8608</link>
<description>Self-organizing Core Allocation
Ziermann, Tobias; Wildermann, Stefan; Teich, Jürgen
This paper deals with the problem of dynamic allocation of cores to parallel applications on homogeneous many-core systems such as, for example, MultiProcessor System-on-Chips (MPSoCs). For a given number of thread-parallel applications, the goal is to find a core assignment that maximizes the average speedup. However, the difficulty is that some applications may have a higher speedup variation than others when assigned additional cores. This paper first presents a centralized algorithm to calculate an optimal assignment for the above objective. However, as the number of cores and the dynamics of applications will significantly increase in the future, decentralized concepts are necessary to scale with this development. Therefore, a decentralized (self-organizing) algorithm is developed in order to minimize the amount of global information that has to be exchanged between applications. The experimental results show that this approach can reach the optimal result of the centralized version in average by 98.95%.
</description>
<dc:date>2013-01-01T00:00:00Z</dc:date>
</item>
<item rdf:about="http://dl.gi.de/handle/20.500.12116/8605">
<title>Comparison of PGAS Languages on a Linked Cell Algorithm</title>
<link>http://dl.gi.de/handle/20.500.12116/8605</link>
<description>Comparison of PGAS Languages on a Linked Cell Algorithm
Bauer, Martin; Kuschel, Christian; Ritter, Daniel; Sembritzki, Klaus
The intention of partitioned global address space (PGAS) languages is to decrease developing time of parallel programs by abstracting the view on the memory and communication. Despite the abstraction a decent speed-up is promised. In this paper the performance and implementation time of Co-Array Fortran (CAF), Unified Parallel C (UPC) and Cascade High Productivity Language (Chapel) are compared by means of a linked cell algorithm. An MPI parallel reference implementation in C is ported to CAF, Chapel and UPC, respectively, and is optimized with respect to the available features of the corresponding language. Our tests show parallel programs are developed faster with the above mentioned PGAS languages as compared to MPI. We experienced a performance penalty for the PGAS versions that can be reduced at the expense of a similar programming effort as for MPI. Programmers should be aware that the utilization of PGAS languages may lead to higher administrative effort for compiling and executing programs on different super-computers.
</description>
<dc:date>2013-01-01T00:00:00Z</dc:date>
</item>
<item rdf:about="http://dl.gi.de/handle/20.500.12116/8607">
<title>Acceleration of Optical Flow Computations on Tightly-Coupled Processor Arrays</title>
<link>http://dl.gi.de/handle/20.500.12116/8607</link>
<description>Acceleration of Optical Flow Computations on Tightly-Coupled Processor Arrays
Sousa, Éricles; Tanase, Alexandru; Lari, Vahid; Hannig, Frank; Teich, Jürgen; Paul, Johny; Stechele, Walter; Kröhnert, Manfred; Asfour, Tamin
Optical flow is widely used in many applications of portable mobile devices and automotive embedded systems for the determination of motion of objects in a visual scene. Also in robotics, it is used for motion detection, object segmentation, time-to-contact information, focus of expansion calculations, robot navigation, and automatic parking for vehicles. Similar to many other image processing algorithms, optical flow processes pixel operations repeatedly over whole image frames. Thus, it provides a high degree of fine-grained parallelism which can be efficiently exploited on massively parallel processor arrays. In this context, we propose to accelerate the computation of complex motion estimation vectors on programmable tightly-coupled processor arrays, which offer a high flexibility enabled by coarse-grained reconfiguration capabilities. Novel is also that the degree of parallelism may be adapted to the number of processors that are available to the application. Finally, we present an implementation that is 18 times faster when compared to (a) an FPGA-based soft processor implementation, and (b) may be adapted regarding different QoS requirements, hence, being more flexible than a dedicated hardware implementation.
</description>
<dc:date>2013-01-01T00:00:00Z</dc:date>
</item>
<item rdf:about="http://dl.gi.de/handle/20.500.12116/8606">
<title>Reduktion von False-Sharing in Software-Transactional-Memory</title>
<link>http://dl.gi.de/handle/20.500.12116/8606</link>
<description>Reduktion von False-Sharing in Software-Transactional-Memory
Kempf, Stefan; Veldema, Ronald; Philippsen, Michael
Software-Transactional-Memory (STM) erleichtert das parallele Programmieren, jedoch hat STM noch einen zu hohen Laufzeitaufwand, da gegenseitiger Ausschluss beim Zugriff auf gemeinsame Daten meist mittels einer Lock-Tabelle fester Gr¨ oße realisiert wird. F¨ ur Programme mit wenigen konkurrierenden Zugriffen und ¨ berwiegend Lesezugriffen ist diese Tabelle gr¨ u oßer als notwendig, so dass beim Commit einer Transaktion mehr Locks zur Konsistenzpr¨ ufung zu inspizieren sind als n¨ otig. F¨ ur große Datenmengen ist die Tabelle zu klein. Dann begrenzt False-Sharing (unterschiedliche Adressen werden auf das gleiche Lock abgebildet) die Parallelit¨ at, da sogar unabh¨ angige Transaktionen sich gegenseitig ausschließen. Diese Arbeit beschreibt eine Technik, die die Lock-Tabelle bei False-Sharing vergr¨ oßert. Zus¨ atzlich kann ein Programmierer mit Annotationen unterschiedliche Lock-Tabellen f¨ ur voneinander unabh¨ angige Daten verlangen, was die M¨ oglichkeit von False-Sharing und den Speicherbedarf f¨ ur die Locks weiter verringert In Benchmarks erreichen wir einen maximalen Speedup von 10.3 gegen¨ uber TL2, wobei die Lock-Tabelle bis zu 1024 mal kleiner ist.
</description>
<dc:date>2013-01-01T00:00:00Z</dc:date>
</item>
</rdf:RDF>
