Fetching the paper…
Reading the bibliography…
Sorting is a primitive operation that is a building block for countless algorithms.
Shear sort: A true two-dimensional sorting techniques for vlsi networks
S. Sen, I. Sherson, and A. Shamir · 1986
Earlier work this paper cites.
The input/output coplexity of sorting and related problems
A. Aggarwal and J. Vitter · 1988
Earlier work this paper cites.
Weighted and unweighted selection algorithms for k sorted sequences
T. Hayashi, K. Nakano, and S. Olariu · 1997
Earlier work this paper cites.
Fundamental parallel algorithms for private-cache chip multiprocessors
L. Arge, M. Goodrich, M. Nelson, and N. Sitchinava · 2008
Earlier work this paper cites.
Fast scan algorithms on graphics processors
Y. Dotsenko, N. K. Govindaraju, P. Sloan, C. Boyd, and J. Manfedelli · 2008
Earlier work this paper cites.
Scalable parallel programming with CUDA
J. N. et al · 2008
Earlier work this paper cites.
Optimization principles and application performance evaluation of a multithreaded GPU using CUDA
S. Ryoo, C. I. Rodrigues, S. S. Baghsorkhi, S. S. Stone, D. B. Kirk, and W.-m. W. Hwu · 2008
Earlier work this paper cites.
Efficient parallel scan algorithms for GPUs
S. Sengupta, M. Harris, and M. Garland · 2008
Earlier work this paper cites.
An analytical model for a gpu architecture with memory-level and thread-level parallelism awareness
S. Hong and H. Kim · 2009
Earlier work this paper cites.
An analytical model for a gpu architecture with memory-level and thread-level parallelism awareness
S. Hong and H. Kim · 2009
Earlier work this paper cites.
A performance prediction model for the cuda gpgpu
K. Kothapalli, R. Mukherjee, S. Rehman, S. Patidar, P. Narayanan, and K. Srinathan · 2009
Earlier work this paper cites.
Parallel Scan for Stream Architectures
D. Merrill and A. Grimshaw · 2009
Earlier work this paper cites.
Deterministic sample sort for GPUs
F. Dehne and H. Zaboli · 2010
Earlier work this paper cites.
Thrust: A parallel template library, 2010
J. Hoberock and N. Bell · 2010
Earlier work this paper cites.
GPU sample sort
N. Leischner, V. Osipov, and P. Sanders · 2010
Cited alongside, same era.
The gpu computing era
J. Nickolls and W. Dally · 2010
Cited alongside, same era.
Demystifying GPU microarchitecture through microbenchmarking
H. Wong · 2010
Cited alongside, same era.
Experimental B + -tree for GPU
K. Kaczmarski · 2011
Cited alongside, same era.
A quantitative performance analysis model for GPU architectures
Y. Zhang and J. D. Owens · 2011
Cited alongside, same era.
GPU merge path: a GPU merging algorithm
O. Green, R. McColl, and D. A. Bader · 2012
Cited alongside, same era.
A decomposition for in-place matrix transposition
B. Catanzaro, A. Keller, and M. Garland · 2014
Later among the works it cites.
Implementation of the DWT in a GPU through a register-based strategy
P. Enfedaque, F. Auli-Llinas, and J. Moure · 2014
Later among the works it cites.
Merge path - A visually intuitive approach to parallel merging
O. Green, S. Odeh, and Y. Birk · 2014
Later among the works it cites.
A memory access model for highly-threaded many-core architectures
L. Ma, K. Agrawal, and R. D. Chamberlain · 2014
Later among the works it cites.
nVidia Tesla K40 specifications, 2014
NVIDIA · 2014
Later among the works it cites.
Sorting and permuting without bank conflicts on gpus
P. Afshani and N. Sitchinava · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
G. Greiner · 2012
Cited alongside, same era.
Programming Massively Parallel Processors
D. B. Kirk · 2012
Cited alongside, same era.
Simple memory machine models for GPUs
K. Nakano · 2012
Cited alongside, same era.
Discrete range searching primitive for the GPU and its applications
J. Soman, K. Kothapalli, and P. J. Narayanan · 2012
Cited alongside, same era.
Modern GPU, 2013
S. Baxter · 2013
Cited alongside, same era.
The hierarchical memory machine model for GPUs
K. Nakano · 2013
Cited alongside, same era.
Characterizing and enhancing global memory data coalescing on GPUs
N. Fauzia, L. N. Pouchet, and P. Sadayappan · 2015
Later among the works it cites.
Efficient batched predecessor search in shared memory on GPUs
B. Karsin, H. Casanova, and N. Sitchinava · 2015
Later among the works it cites.
Cub: Cuda unbound, 2015
D. Merrill · 2015
Later among the works it cites.
CUDA programming guide 7.0, 2015
NVIDIA · 2015
Later among the works it cites.
Nsight, 2015
NVIDIA · 2015
Later among the works it cites.
A fine-grained performance model for gpu architectures
N. Bombieri, F. Busato, and F. Fummi · 2016
Later among the works it cites.
A performance comparison of sort and scan libraries for GPUs
B. Merry · 2016
Later among the works it cites.