Fetching the paper…
Reading the bibliography…
Large-scale machine learning and data mining applications require computer systems to perform massive matrix-vector and matrix-matrix multiplication operations that need to be parallelized across multiple nodes.
Parity Models: A General Framework for Coding-Based Resilience in ML Inference
Jack Kosaian, K. V. Rashmi, and Shivaram Venkataraman. 2019 · 1905
Earlier work this paper cites.
Algorithm-based Fault Tolerance for Matrix Operations
Kuang-Hua Huang et al · 1984
Earlier work this paper cites.
Matrix Algorithms on a Hypercube I: Matrix Multiplication
Geoffrey C. Fox, Steve W. Otto, and Anthony JG. Hey. 1987 · 1987
Earlier work this paper cites.
Approximate Analysis of Fork/Join Synchronization in Parallel Queues
R. Nelson and A. Tantawi. 1988 · 1988
Earlier work this paper cites.
Analysis of the Fork-Join Queue
C. Kim and A. K. Agrawala. 1989 · 1989
Earlier work this paper cites.
Introduction to Parallel Computing: Design and Analysis of Algorithms . Vol. 400
Vipin Kumar, Ananth Grama, Anshul Gupta, and George Karypis. 1994 · 1994
Earlier work this paper cites.
The PageRank Citation Ranking: Bringing order to the Web
Lawrence Page, Sergey Brin, Rajeev Motwani, and Terry Winograd. 1999 · 1999
Earlier work this paper cites.
LT codes. In null . IEEE, 271
Michael Luby. 2002 · 2002
Earlier work this paper cites.
Order statistics
H. A. David and H. N. Nagaraja. 2003 · 2003
Earlier work this paper cites.
Information theory, inference and learning algorithms
David JC MacKay. 2003 · 2003
Earlier work this paper cites.
Amazon Web Services EC2
Amazon. 2006 · 2006
Earlier work this paper cites.
Raptor codes
Amin Shokrollahi. 2006 · 2006
Earlier work this paper cites.
Dynamic load balancing of unbalanced computations using message passing. In Parallel and Distributed Processing Symposium, 2007. IPDPS 2007. IEEE International . IEEE, 1–8
James Dinan, Stephen Olivier, Gerald Sabin, Jan Prins, P Sadayappan, and Chau-Wen Tseng. 2007 · 2007
Earlier work this paper cites.
MapReduce: simplified data processing on large clusters
Jeffrey Dean and Sanjay Ghemawat. 2008 · 2008
Earlier work this paper cites.
The M/M/1 fork-join queue with variable sub-tasks
Elizabeth Varki, Arif Merchant, and Hui Chen. 2008 · 2008
Earlier work this paper cites.
Scalable work stealing. In High Performance Computing Networking, Storage and Analysis, Proceedings of the Conference on . IEEE, 1–11
James Dinan, D Brian Larkins, Ponnuswamy Sadayappan, Sriram Krishnamoorthy, and Jarek Nieplocha. 2009 · 2009
Earlier work this paper cites.
Reining in the Outliers in Map-Reduce Clusters using Mantri.. In USENIX Symposium on Operating Systems Design and Implementation (OSDI) , Vol. 10. 24
Ganesh Ananthanarayanan, Srikanth Kandula, Albert G Greenberg, Ion Stoica, Yi Lu, Bikas Saha, and Edward Harris. 2010 · 2010
Earlier work this paper cites.
Fountain codes. In Global telecommunications conference (GLOBECOM 2010) . 7–12
Gauri Joshi, Joong Bum Rhim, John Sun, and Da Wang. 2010 · 2010
Earlier work this paper cites.
Spark: Cluster computing with working sets
Matei Zaharia, Mosharaf Chowdhury, Michael J Franklin, Scott Shenker, and Ion Stoica. 2010 · 2010
Earlier work this paper cites.
An analysis of single-layer networks in unsupervised feature learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics . 215–223
Adam Coates, Andrew Ng, and Honglak Lee. 2011 · 2011
Earlier work this paper cites.
Raptor codes
Amin Shokrollahi, Michael Luby, et al · 2011
Earlier work this paper cites.
Codes can reduce queueing delay in data centers. In IEEE International Symposium on Information Theory Proceedings (ISIT) . 2766–2770
Longbo Huang, S. Pawar, Hao Zhang, and K. Ramchandran. 2012 · 2012
Cited alongside, same era.
Coding for fast content download. In Allerton Conference on Communication, Control, and Computing . IEEE, 326–333
Gauri Joshi, Yanpei Liu, and Emina Soljanin. 2012 · 2012
Cited alongside, same era.
Effective Straggler Mitigation: Attack of the Clones.. In USENIX Symposium on Networked Systems Design and Implementation (NSDI) , Vol. 13. 185–198
Ganesh Ananthanarayanan, Ali Ghodsi, Scott Shenker, and Ion Stoica. 2013 · 2013
Cited alongside, same era.
The tail at scale
Jeffrey Dean and Luiz André Barroso. 2013 · 2013
Cited alongside, same era.
Performance modeling and design of computer systems: queueing theory in action
Mor Harchol-Balter. 2013 · 2013
Cited alongside, same era.
On Delay-Optimal Scheduling in Queueing Systems with Replications
Yin Sun, Can Emre Koksal, and Ness B. Shroff. 2016 · 2016
Later among the works it cites.
Coded convolution for parallel and distributed computing within a deadline. In IEEE International Symposium on Information Theory (ISIT) . IEEE, 2403–2407
Sanghamitra Dutta, Viveck Cadambe, and Pulkit Grover. 2017 · 2017
Later among the works it cites.
Improving distributed gradient descent using reed-solomon codes
Wael Halbawi, Navid Azizan-Ruhi, Fariborz Salehi, and Babak Hassibi. 2017 · 2017
Later among the works it cites.
Boosting Service Capacity via Adaptive Task Replication
Gauri Joshi. 2017 · 2017
Later among the works it cites.
Efficient Redundancy Techniques for Latency Reduction in Cloud Systems
Gauri Joshi, Emina Soljanin, and Gregory Wornell. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Amazon Web Services Lambda
Amazon. 2014 · 2014
Cited alongside, same era.
Numerical Methods for Partial Differential Equations
William F Ames. 2014 · 2014
Cited alongside, same era.
On the Delay-Storage Trade-Off in Content Download from Coded Distributed Storage Systems
Gauri Joshi, Yanpei Liu, and Emina Soljanin. 2014 · 2014
Cited alongside, same era.
Efficient Task Replication for Fast Response times in Parallel Computation. In ACM SIGMETRICS Performance Evaluation Review , Vol. 42. ACM, 599–600
Da Wang, Gauri Joshi, and Gregory Wornell. 2014 · 2014
Cited alongside, same era.
High-performance Hardware for Machine Learning
William Dally. 2015 · 2015
Cited alongside, same era.
Kubernetes
Google. 2015 · 2015
Cited alongside, same era.
Queues with redundancy: Latency-cost analysis
Gauri Joshi, Emina Soljanin, and Gregory Wornell. 2015 · 2015
Cited alongside, same era.
Later among the works it cites.
Encoded distributed optimization. In IEEE International Symposium on Information Theory (ISIT) . IEEE, 2890–2894
Can Karakus, Yifan Sun, and Suhas Diggavi. 2017 · 2017
Later among the works it cites.
Speeding Up Distributed Machine Learning Using Codes
Kangwook Lee, Maximilian Lam, Ramtin Pedarsani, Dimitris Papailiopoulos, and Kannan Ramchandran. 2017a · 2017
Later among the works it cites.
The MDS Queue: Analysing the Latency Performance of Erasure Codes
Kangwook Lee, Nihar B. Shah, Longbo Huang, and Kannan Ramchandran. 2017b · 2017
Later among the works it cites.
Near-Optimal Straggler Mitigation for Distributed Gradient Methods
Songze Li, Seyed Mohammadreza Mousavi Kalan, A Salman Avestimehr, and Mahdi Soltanolkotabi. 2017 · 2017
Later among the works it cites.
Block-Diagonal and LT Codes for Distributed Computing With Straggling Servers
Albin Severinson, Alexandre Graell i Amat, and Eirik Rosnes. 2017 · 2017
Later among the works it cites.
Gradient Coding: Avoiding Stragglers in Synchronous Gradient Descent
Rashish Tandon, Qi Lei, Alexandros G Dimakis, and Nikos Karampatziakis. 2017 · 2017
Later among the works it cites.
Coded Distributed Computing for Inverse Problems. In Advances in Neural Information Processing Systems . 709–719
Yaoqing Yang, Pulkit Grover, and Soummya Kar. 2017 · 2017
Later among the works it cites.
Qian Yu, Mohammad Ali Maddah-Ali, and A Salman Avestimehr. 2017b · 2017
Later among the works it cites.
On the Optimal Recovery Threshold of Coded Matrix Multiplication
Sanghamitra Dutta, Mohammad Fahim, Farzin Haddadpour, Haewon Jeong, Viveck Cadambe, and Pulkit Grover. 2018 · 2018
Closest in time.
Oversketch: Approximate matrix multiplication for the cloud. In 2018 IEEE International Conference on Big Data (Big Data) . IEEE, 298–304
Vipul Gupta, Shusen Wang, Thomas Courtade, and Kannan Ramchandran. 2018 · 2018
Closest in time.
Synergy via Redundancy: Boosting Service Capacity with Adaptive Replication
Gauri Joshi. 2018 · 2018
Closest in time.
Learning a Code: Machine Learning for Approximate Non-Linear Coded Computation
Jack Kosaian, K. V. Rashmi, and Shivaram Venkataraman. 2018 · 2018
Closest in time.
numpywren: serverless linear algebra
Vaishaal Shankar, Karl Krauth, Qifan Pu, Eric Jonas, Shivaram Venkataraman, Ion Stoica, Benjamin Recht, and Jonathan Ragan-Kelley. 2018 · 2018
Closest in time.
Coded Sparse Matrix Multiplication
Sinong Wang, Jiashang Liu, and Ness Shroff. 2018 · 2018
Closest in time.
Efficient Straggler Replication in Large-Scale Parallel Computing
Da Wang, Gauri Joshi, and Gregory W. Wornell. 2019 · 2019
Closest in time.