Fetching the paper…
Reading the bibliography…
Coreset, which is a summary of the original dataset in the form of a small weighted set in the same sample space, provides a promising approach to enable machine learning over distributed data.
R. Fisher, “Iris data set,” https://archive.ics.uci.edu/ml/datasets/iris, 1936. [Online]. Available: https://archive.ics.uci.edu/ml/datasets/iris
1936
Earlier work this paper cites.
N. Megiddo and K. J. Supowit, “On the complexity of some common geometric location problems,” SIAM Journal of Computing , vol. 13, no. 1, pp. 182–196, 1984
1984
Earlier work this paper cites.
D. Wolpert, “Stacked generalization,” Neural Networks , vol. 5, no. 2, pp. 241–259, 1992
1992
Earlier work this paper cites.
E. Alpaydin and F. Alimoglu, “Pen-based recognition of handwritten digits data set,” https://archive.ics.uci.edu/ml/datasets/Pen-Based+Recognition+of+Handwritten+Digits, 1996. [Online]. Available: https://archive.ics.uci.edu/ml/datasets/Pen-Based+Recognition+of+Handwritten+Digits
1996
Earlier work this paper cites.
P. K. Chan and S. J. Stolfo, “Toward parallel and distributed learning by meta-learning,” in AAAI Workshop in Knowledge Discovery in Databases , 1997
1997
Earlier work this paper cites.
J. Kittler, M. Hatef, R. P. Duin, and J. Matas, “On combining classifiers,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 20, no. 3, pp. 226–239, March 1998
1998
Earlier work this paper cites.
Y. LeCun, C. Cortes, and C. Burges, “The MNIST database of handwritten digits,” http://yann.lecun.com/exdb/mnist/, 1998. [Online]. Available: http://yann.lecun.com/exdb/mnist/
1998
Earlier work this paper cites.
Y. Guo and J. Sutiwaraphun, “Probing knowledge in distributed data mining,” in Methodologies for Knowledge Discovery and Data Mining , 1999
1999
Earlier work this paper cites.
M. Bādoiu, S. Har-Peled, and P. Indyk, “Approximate clustering via core-sets,” in ACM STOC , 2002
2002
Earlier work this paper cites.
G. Tsoumakas and I. Vlahavas, “Effective stacking of distributed classifiers,” in ECAI , 2002
2002
Earlier work this paper cites.
A. Lazarevic and Z. Obradovic, “Boosting algorithms for parallel and distributed learning,” Distributed and Parallel Databases , vol. 11, no. 2, pp. 203–229, March 2002
2002
Earlier work this paper cites.
S. Har-Peled and K. R. Varadarajan, “Projective clustering in high dimensions using core-sets,” in SOCG , 2002
2002
Earlier work this paper cites.
M. Bādoiu and K. L. Clarkson, “Smaller core-sets for balls,” in SODA , 2003
2003
Earlier work this paper cites.
N. Chawla, L. Halla, K. Bowyer, and W. Kegelmeyer, “Learning ensembles from bites: A scalable and accurate approach,” Journal of Machine Learning Research , vol. 5, pp. 421–451, April 2004
2004
Earlier work this paper cites.
S. Har-Peled and S. Mazumdar, “On coresets for k-means and k-median clustering,” in STOC , 2004
2004
Earlier work this paper cites.
I. W. Tsang, J. T. Kwok, and P.-M. Cheung, “Core vector machines: Fast SVM training on very large data sets,” The Journal of Machine Learning Research , vol. 6, pp. 363–392, December 2005
2005
Earlier work this paper cites.
G. Frahling and C. Sohler, “Coresets in dynamic geometric data streams,” in STOC , 2005
2005
Earlier work this paper cites.
D. Feldman, A. Fiat, and M. Sharir, “Coresets for weighted facilities and their applications,” in FOCS , 2006
2006
Cited alongside, same era.
S. Har-Peled, D. Roth, and D. Zimak, “Maximum margin coresets for active and noise tolerant learning,” in IJCAI , 2007
2007
Cited alongside, same era.
D. Arthur and S. Vassilvitskii, “k-means++: The advantages of careful seeding,” in SODA , January 2007
2007
Cited alongside, same era.
D. Aloise, A. Deshpande, P. Hansen, and P. Popat, “NP-hardness of Euclidean sum-of-squares clustering,” Machine Learning , vol. 75, no. 2, pp. 245–248, May 2009
2009
Cited alongside, same era.
K. L. Clarkson, “Coresets, sparse greedy approximation, and the Frank-Wolfe algorithm,” ACM Transactions on Algorithms , vol. 6, no. 4, August 2010
2010
Cited alongside, same era.
2016
Later among the works it cites.
A. Barger and D. Feldman, “k-means for streaming and distributed big sparse data,” in SDM , 2016
2016
Later among the works it cites.
J. M. Phillips, “Coresets and sketches,” CoRR , vol. abs/1601.00617, 2016
2016
Later among the works it cites.
D. Feldman, M. Volkov, and D. Rus, “Dimensionality reduction of massive sparse datasets using coresets,” in NIPS , 2016
2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Langberg and L. J. Schulman, “Universal ϵ \epsilon approximators for integrals,” in SODA , 2010
2010
Cited alongside, same era.
D. Feldman and M. Langberg, “A unified framework for approximating and clustering data,” in STOC , June 2011
2011
Cited alongside, same era.
D. Feldman, A. Krause, and M. Faulkner, “Scalable training of mixture models via coresets,” in NIPS , 2011
2011
Cited alongside, same era.
D. Peteiro-Barral and B. Guijarro-Berdinas, “A survey of methods for distributed machine learning,” in Progress in Artificial Intelligence , November 2012
2012
Cited alongside, same era.
M. F. Balcan, S. Ehrlich, and Y. Liang, “Distributed k-means and k-median clustering on general topologies,” in NIPS , 2013
2013
Cited alongside, same era.
D. Feldman, M. Feigin, and N. Sochen, “Learning big (image) data via coresets for dictionaries,” Journal of Mathematical Imaging and Vision , vol. 46, no. 3, pp. 276–291, March 2013
2013
Cited alongside, same era.
D. Feldman, M. Schmidt, and C. Sohler, “Turning big data into tiny data: Constant-size coresets for k-means, PCA, and projective clustering,” in SODA , 2013
2013
Cited alongside, same era.
2016
Later among the works it cites.
M. Cohen, Y. T. Lee, G. Miller, J. Pachocki, and A. Sidford, “Geometric median in nearly linear time,” in STOC , 2016
2016
Later among the works it cites.
S. Moro, P. Rita, and B. Vala, “Facebook metrics data set,” https://archive.ics.uci.edu/ml/datasets/Facebook+metrics, 2016. [Online]. Available: https://archive.ics.uci.edu/ml/datasets/Facebook+metrics
2016
Later among the works it cites.
V. Smith, C.-K. Chiang, M. Sanjabi, and A. S. Talwalkar, “Federated multi-task learning,” in Advances in Neural Information Processing Systems , 2017, pp. 4424–4434
2017
Later among the works it cites.
J. Wang, J. D. Lee, M. Mahdavi, M. Kolar, N. Srebro et al. , “Sketching meets random projection in the dual: A provable recovery algorithm for big and high-dimensional data,” Electronic Journal of Statistics , vol. 11, no. 2, pp. 4896–4944, 2017
2017
Later among the works it cites.
S. Har-Peled, Geometric Approximation Algorithms . American Mathematical Society, 2011
2017
Later among the works it cites.
S. Wang, T. Tuor, T. Salonidis, K. K. Leung, C. Makaya, T. He, and K. Chan, “When edge meets learning: Adaptive control for resource-constrained distributed machine learning,” in IEEE INFOCOM , April 2018
2018
Later among the works it cites.
A. Munteanu and C. Schwiegelshohn, “Coresets-methods and history: A theoreticians design pattern for approximation and streaming algorithms,” KI - Künstliche Intelligenz , vol. 32, no. 1, pp. 37–53, 2018
2018
Later among the works it cites.
A. Molina, A. Munteanu, and K. Kersting, “Core dependency networks,” in AAAI , 2018
2018
Later among the works it cites.
A. Virmaux and K. Scaman, “Lipschitz regularity of deep neural networks: analysis and efficient estimation,” in Advances in Neural Information Processing Systems , 2018, pp. 3835–3844
2018
Later among the works it cites.
H. Lu, M. Li, T. He, S. Wang, V. Narayanan, and K. S. Chan, “Robust coreset construction for distributed machine learning,” in 2019 IEEE Global Communications Conference (GLOBECOM) , 2019
2019
Closest in time.
K. Makarychev, Y. Makarychev, and I. Razenshteyn, “Performance of johnson-lindenstrauss transform for k-means and k-medians clustering,” in Proceedings of the 51st Annual ACM SIGACT Symposium on Theory of Computing . ACM, 2019, pp. 1027–1038
2019
Closest in time.