Fetching the paper…
Reading the bibliography…
Privatizing data is a useful strategy for increasing parallelism in a shared memory multithreaded program.
Parallel prefix computation
R. E. Ladner and M. J. Fischer · 1980
Earlier work this paper cites.
Advanced compiler optimizations for supercomputers
D. A. Padua and M. J. Wolfe · 1986
Earlier work this paper cites.
Automatic array privatization
P. Tu and D. A. Padua · 1994
Earlier work this paper cites.
Commutativity analysis: A new analysis technique for parallelizing compilers
M. C. Rinard and P. C. Diniz · 1997
Earlier work this paper cites.
The anatomy of a large-scale hypertextual web search engine
S. Brin and L. Page · 1998
Earlier work this paper cites.
Data speculation support for a chip multiprocessor
L. Hammond, M. Willey, and K. Olukotun · 1998
Earlier work this paper cites.
A scalable approach to thread-level speculation
J. G. Steffan, C. B. Colohan, A. Zhai, and T. C. Mowry · 2000
Earlier work this paper cites.
IEEE Standard 1003.1-2001, 2001
IEEE and The Open Group · 2001
Earlier work this paper cites.
Mapreduce: Simplified data processing on large clusters
J. Dean and S. Ghemawat · 2004
Earlier work this paper cites.
Programming with transactional coherence and consistency (tcc)
L. Hammond, B. D. Carlstrom, V. Wong, B. Hertzberg, M. Chen, C. Kozyrakis, and K. Olukotun · 2004
Earlier work this paper cites.
Unbounded transactional memory
C. S. Ananian, K. Asanovic, B. C. Kuszmaul, C. E. Leiserson, and S. Lie · 2005
Earlier work this paper cites.
Pin: Building customized program analysis tools with dynamic instrumentation
C.-K. Luk, R. Cohn, R. Muth, H. Patil, A. Klauser, G. Lowney, S. Wallace, V. J. Reddi, and K. Hazelwood · 2005
Earlier work this paper cites.
The stampede approach to thread-level speculation
J. G. Steffan, C. Colohan, A. Zhai, and T. C. Mowry · 2005
Earlier work this paper cites.
Bulksc: Bulk enforcement of sequential consistency
L. Ceze, J. Tuck, P. Montesinos, and J. Torrellas · 2007
Earlier work this paper cites.
Concurrent programming without locks
K. Fraser and T. Harris · 2007
Cited alongside, same era.
An integrated hardware-software approach to flexible transactional memory
A. Shriraman, M. F. Spear, H. Hossain, V. J. Marathe, S. Dwarkadas, and M. L. Scott · 2007
Cited alongside, same era.
Mechanisms for store-wait-free multiprocessors
T. F. Wenisch, A. Ailamaki, B. Falsafi, and A. Moshovos · 2007
Cited alongside, same era.
Mapreduce: Simplified data processing on large clusters
J. Dean and S. Ghemawat · 2008
Cited alongside, same era.
Advanced Synchronization Facility Proposed Architectural Specification, Publication 45432, Rev. 2.1, 2009
Advanced Micro Devices · 2009
Cited alongside, same era.
Grace: safe and efficient concurrent programming
E. Berger, T. Yang, T. Liu, D. Krishnan, and A. Novark · 2009
Cited alongside, same era.
Drfx: A simple and efficient memory model for concurrent programming languages
D. Marino, A. Singh, T. Millstein, M. Musuvathi, and S. Narayanasamy · 2010
Later among the works it cites.
Introducing the graph 500
R. C. Murphy, K. B. Wheeler, B. W. Barrett, and J. A. Ang · 2010
Later among the works it cites.
Two for the price of one: A model for parallel and incremental computation
S. Burckhardt, D. Leijen, C. Sadowski, J. Yi, and T. Ball · 2011
Later among the works it cites.
Commutative set: A language extension for implicit parallel programming
P. Prabhu, S. Ghosh, Y. Zhang, N. P. Johnson, and D. I. August · 2011
Later among the works it cites.
Managing performance vs. accuracy trade-offs with loop perforation
S. Sidiroglou-Douskos, S. Misailovic, H. Hoffmann, and M. Rinard · 2011
Later among the works it cites.
A Primer on Memory Consistency and Cache Coherence
D. J. Sorin, M. D. Hill, and D. A. Wood · 2011
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Invisifence: Performance-transparent memory ordering in conventional multiprocessors
C. Blundell, M. M. Martin, and T. F. Wenisch · 2009
Cited alongside, same era.
Rodinia: A benchmark suite for heterogeneous computing
S. Che, M. Boyer, J. Meng, D. Tarjan, J. W. Sheaffer, S.-H. Lee, and K. Skadron · 2009
Cited alongside, same era.
Reducers and other cilk++ hyperobjects
M. Frigo, P. Halpern, C. E. Leiserson, and S. Lewin-Berlin · 2009
Cited alongside, same era.
Cacti 6.0: A tool to model large caches
N. Muralimanohar, R. Balasubramonian, and N. P. Jouppi · 2009
Cited alongside, same era.
Coredet: a compiler and runtime system for deterministic multithreaded execution
T. Bergan, O. Anderson, J. Devietti, L. Ceze, and D. Grossman · 2010
Cited alongside, same era.
Retcon: Transactional repair without replay
C. Blundell, A. Raghavan, and M. M. Martin · 2010
Cited alongside, same era.
Later among the works it cites.
Efficient system-enforced deterministic parallelism
A. Aviram, S.-C. Weng, S. Hu, and B. Ford · 2012
Later among the works it cites.
OpenMP Application Programming Interface Version 4.0
OpenMP Architecture Review Board · 2013
Later among the works it cites.
Performance evaluation of intel(r) transactional synchronization extensions for high-performance computing
R. Yoo, C. Hughes, K. Lai, and R. Rajwar · 2013
Later among the works it cites.
General data structure expansion for multi-threading
H. Yu, H.-J. Ko, and Z. Li · 2013
Later among the works it cites.
S. Beamer, K. Asanovic, and D. A. Patterson · 2015
Later among the works it cites.
Valor: Efficient, software-only region conflict exceptions
S. Biswas, M. Zhang, M. D. Bond, and B. Lucia · 2015
Later among the works it cites.
Parallelizing user-defined aggregations using symbolic execution
V. Raychev, M. Musuvathi, and T. Mytkowicz · 2015
Later among the works it cites.
Exploiting commutativity to reduce the cost of updates to shared data in cache-coherent systems
G. Zhang, W. Horn, and D. Sanchez · 2015
Later among the works it cites.