pith. sign in

arxiv: 1001.4108 · v2 · submitted 2010-01-23 · 💻 cs.DC · cs.PF

A Multi-Stage CUDA Kernel for Floyd-Warshall

classification 💻 cs.DC cs.PF
keywords algorithmcudafloyd-warshallachieveall-pairsallowappliedapproximately
0
0 comments X
read the original abstract

We present a new implementation of the Floyd-Warshall All-Pairs Shortest Paths algorithm on CUDA. Our algorithm runs approximately 5 times faster than the previously best reported algorithm. In order to achieve this speedup, we applied a new technique to reduce usage of on-chip shared memory and allow the CUDA scheduler to more effectively hide instruction latency.

This paper has not been read by Pith yet.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.