<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" version="2.0">
<channel>
<title>PARS-Mitteilungen</title>
<link>http://dl.gi.de/handle/20.500.12116/1903</link>
<description/>
<pubDate>Tue, 21 Jul 2026 14:11:39 GMT</pubDate>
<dc:date>2026-07-21T14:11:39Z</dc:date>
<item>
<title>A Quantitative Analysis of Processor Memory Bandwidth of an FPGA-MPSoC</title>
<link>http://dl.gi.de/handle/20.500.12116/33868</link>
<description>A Quantitative Analysis of Processor Memory Bandwidth of an FPGA-MPSoC
Drehmel, Robert; Göbel, Matthias; Juurlink, Ben
System designers have to choose between a variety of different memories available on modern FPGA-MPSoCs. Our intention is to shed light on the achievable bandwidth when accessing them under diverse circumstances and to hint at their suitability for general-purpose applications. We conducted a systematic quantitative analysis of the memory bandwidth of two processing units using a sophisticated standalone bandwidth measurement tool. The results show a maximum cacheable memory bandwidth of 7.11 GiB/s for reads and 11.78 GiB/s for writes for the general-purpose processing unit, and 2.56 GiB/s for reads and 1.83 GiB/s writes for the special-purpose (real-time) processing unit. In contrast, the achieved non-cacheable read bandwidth lies between 60 MiB/s and 207 MiB/s, with an outlier of 2.67 GiB/s. We conclude that for most applications, relying on DRAM and hardware cache coherency management is the best choice in terms of benefit-cost ratio.
</description>
<pubDate>Wed, 01 Jan 2020 00:00:00 GMT</pubDate>
<guid isPermaLink="false">http://dl.gi.de/handle/20.500.12116/33868</guid>
<dc:date>2020-01-01T00:00:00Z</dc:date>
</item>
<item>
<title>Influence of Discretization of Frequencies and Processor Allocation on Static Scheduling of Parallelizable Tasks with Deadlines</title>
<link>http://dl.gi.de/handle/20.500.12116/33869</link>
<description>Influence of Discretization of Frequencies and Processor Allocation on Static Scheduling of Parallelizable Tasks with Deadlines
Litzinger, Sebastian; Keller, Jörg
Models for energy-efficient static scheduling of parallelizable tasks with deadlines onfrequency-scalable parallel machines comprise moldable vs. malleable tasks and continuous vs. discrete frequency levels. We investigate the tradeoff between scheduling time and energy efficiency when going from continuous to discrete processor allocation and frequency levels. To this end, we present a tool to convert a schedule computed for malleable tasks on machines with continuous frequency scaling (P. Sanders, J. Speck, Euro-Par 2012) into one for moldable tasks on a machine with discrete frequency levels. We compare the energy efficiency of the converted schedule to the energy consumed by a schedule produced by the integrated crown scheduler (N. Melot et al., ACM TACO 2015) for moldable tasks and a machine with discrete frequency levels. Our experiments indicate that the converted Sanders Speck schedules, while computed faster, consume more energy on average than crown schedules. Surprisingly, it is not the step from malleable to moldable tasks that is responsible, but the step from continuous to discrete frequency levels.
</description>
<pubDate>Wed, 01 Jan 2020 00:00:00 GMT</pubDate>
<guid isPermaLink="false">http://dl.gi.de/handle/20.500.12116/33869</guid>
<dc:date>2020-01-01T00:00:00Z</dc:date>
</item>
<item>
<title>Generating Optimized FPGA Based MPSoCs to Parallelize Legacy Embedded Software with Customizable Throughput</title>
<link>http://dl.gi.de/handle/20.500.12116/33867</link>
<description>Generating Optimized FPGA Based MPSoCs to Parallelize Legacy Embedded Software with Customizable Throughput
Heid, Kris; Hochberger, Christian
Executing legacy software on newly developed systems can lead to problems regarding the required throughput of the software. Automatic software parallelization can help to achieve a desired exection time even if a single core version would be to slow. In this contribution, we present a toolset that automatically parallelizes a given legacy software and distributes it to multiple soft-cores forming a processing pipeline. As a goal for the parallelization, the user can provide a minimum throughput that has to be achieved. Although this concept is limited to repetitive tasks, it can be well applied to most embedded system applications. The results show that the tool achieves remarkable speedups without any manual intervention or code restructuring for a sprectrum of benchmarks.
</description>
<pubDate>Wed, 01 Jan 2020 00:00:00 GMT</pubDate>
<guid isPermaLink="false">http://dl.gi.de/handle/20.500.12116/33867</guid>
<dc:date>2020-01-01T00:00:00Z</dc:date>
</item>
<item>
<title>GPU-beschleunigte Time Warping-Distanzen</title>
<link>http://dl.gi.de/handle/20.500.12116/33866</link>
<description>GPU-beschleunigte Time Warping-Distanzen
Bachmann, Jörg P.; Trogant, Kevin M.; Freytag, Johann-C.
Immer mehr Algorithmen konnten durch Implementierung auf GPUs um mehrere Größenordnungen beschleunigt werden. Insbesondere existieren hochparallele Implementierungen des im Bereich der Zeitreihenanalyse weit verbreiteten Algorithmus’ Dynamic Time Warping (DTW). Dieser Algorithmus berechnet einen Ähnlichkeitswert zweier Zeitreihen (z. B. Temperaturverläufe) unter Berücksichtigung zeitlicher Variationen wie z. B. zeitliche Verschiebungen. Leider können die existierenden GPU-Implementierungen von DTW nicht beliebige zeitliche Variationen berücksichtigen. In dieser Arbeit stellen wir Implementierungen für GPUs vor, die dieser Einschränkung nicht unterliegen. In unserer Evaluierung zeigen wir, dass sie einen Geschwindigkeitsvorteil von ca. zwei Größenordnungen gegenüber einer CPU-Implementierung erreichen.
</description>
<pubDate>Wed, 01 Jan 2020 00:00:00 GMT</pubDate>
<guid isPermaLink="false">http://dl.gi.de/handle/20.500.12116/33866</guid>
<dc:date>2020-01-01T00:00:00Z</dc:date>
</item>
</channel>
</rss>
