Fetching the paper…
Reading the bibliography…
While most work on evaluating machine learning (ML) models focuses on computing accuracy on batches of data, tracking accuracy alone in a streaming setting (i.e., unbounded, timestamp-ordered datasets) fails to appropriately identify when models are performing unexpectedly.
Increasing the robustness of DNNs against image corruptions by playing the Game of Noise
Evgenia Rusak, Lukas Schott, Roland S. Zimmermann, Julian Bitterwolf, Oliver Bringmann, Matthias Bethge, and Wieland Brendel. 2020 · 2001
Earlier work this paper cites.
Covariate Shift Adaptation by Importance Weighted Cross Validation. In JMLR
Masashi Sugiyama et al · 2007
Earlier work this paper cites.
WILDS: A Benchmark of in-the-Wild Distribution Shifts
Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Sara Beery, Jure Leskovec, Anshul Kundaje, Emma Pierson, Sergey Levine, Chelsea Finn, and Percy Liang. 2020 · 2012
Earlier work this paper cites.
TLC Trip Record Data
2020 · 2020
Earlier work this paper cites.
The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution Generalization
Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Lixuan Zhu, Samyak Parajuli, Mike Guo, Dawn Xiaodong Song, Jacob Steinhardt, and Justin Gilmer. 2021 · 2021
Cited alongside, same era.
The CLEAR Benchmark: Continual LEArning on Real-World Imagery. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 2)
Zhiqiu Lin, Jia Shi, Deepak Pathak, and Deva Ramanan. 2021 · 2021
Cited alongside, same era.
Ease.ML: A Lifecycle Management System for MLDev and MLOps. In Conference on Innovative Data Systems Research
Leonel Aguilar Melgar, David Dao, Shaoduo Gan, Nezihe M Gürel, Nora Hollenstein, Jiawei Jiang, Bojan Karlaš, Thomas Lemmin, Tian Li, Yang Li, Susie Rao, Johannes Rausch, Cedric Renggli, Luka Rimanic, Maurice Weber, Shuai Zhang, Zhikuan Zhao, Kevin Schawinski, Wentao Wu, and Ce Zhang. 2021 · 2021
Cited alongside, same era.
{BREEDS}: Benchmarks for Subpopulation Shift. In International Conference on Learning Representations
Shibani Santurkar, Dimitris Tsipras, and Aleksander Madry. 2021 · 2021
Cited alongside, same era.
Towards Observability for Machine Learning Pipelines
Shreya Shankar and Aditya Parameswaran. 2021 · 2021
Later among the works it cites.
DABS: a Domain-Agnostic Benchmark for Self-Supervised Learning. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1)
Alex Tamkin, Vincent Liu, Rongfei Lu, Daniel Fein, Colin Schultz, and Noah Goodman. 2021 · 2021
Later among the works it cites.
Deconstructing Distributions: A Pointwise Framework of Learning
Gal Kaplun, Nikhil Ghosh, S. Garg, Boaz Barak, and Preetum Nakkiran. 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…