Fetching the paper…
Reading the bibliography…
Data collection and labeling are critical bottlenecks in the deployment of machine learning applications.
H. S. Seung, M. Opper, and H. Sompolinsky, “Query by committee,” in Proceedings of the Fifth Annual Workshop on Computational Learning Theory , ser. COLT ’92. New York, NY, USA: Association for Computing Machinery, 1992, p. 287–294. [Online]. Available: https://doi.org/10.1145/130385.130417
1992
Earlier work this paper cites.
J. Madhavan, P. Bernstein, A. Doan, and A. Halevy, “Corpus-based schema matching,” in 21st International Conference on Data Engineering (ICDE’05) , 2005, pp. 57–68
2005
Earlier work this paper cites.
A. Kapoor, E. Horvitz, and S. Basu, “Selective supervision: Guiding supervised learning with decision-theoretic active learning,” in IJCAI’07 Proceedings of the 20th international joint conference on Artifical intelligence . Morgan Kaufmann Publishers Inc., January 2007, pp. 877–882. [Online]. Available: https://www.microsoft.com/en-us/research/publication/selective-supervision-guiding-supervised-learning-decision-theoretic-active-learning/
2007
Earlier work this paper cites.
J. S. Hamid, P. Hu, N. M. Roslin, V. Ling, C. M. Greenwood, and J. Beyene, “Data integration in genetics and genomics: methods and challenges,” Human genomics and proteomics: HGP , vol. 2009, 2009
2009
Earlier work this paper cites.
H. Gonzalez, A. Halevy, C. S. Jensen, A. Langen, J. Madhavan, R. Shapley, and W. Shen, “Google fusion tables: data management, integration and collaboration in the cloud,” in Proceedings of the 1st ACM symposium on Cloud computing , 2010, pp. 175–180
2010
Earlier work this paper cites.
G. Little, L. B. Chilton, M. Goldman, and R. C. Miller, “Turkit: human computation algorithms on mechanical turk,” in Proceedings of the 23nd Annual ACM Symposium on User Interface Software and Technology , ser. UIST ’10. New York, NY, USA: Association for Computing Machinery, 2010, p. 57–66. [Online]. Available: https://doi.org/10.1145/1866029.1866040
2010
Earlier work this paper cites.
A. Marcus, E. Wu, D. R. Karger, S. Madden, and R. C. Miller, “Demonstration of qurk: a query processor for humanoperators,” in Proceedings of the 2011 ACM SIGMOD International Conference on Management of Data , ser. SIGMOD ’11. New York, NY, USA: Association for Computing Machinery, 2011, p. 1315–1318. [Online]. Available: https://doi.org/10.1145/1989323.1989486
2011
Earlier work this paper cites.
C. E. Brodley, U. Rebbapragada, K. Small, and B. Wallace, “Challenges and opportunities in applied machine learning,” Ai Magazine , vol. 33, no. 1, pp. 11–24, 2012
2012
Earlier work this paper cites.
M. Yakout, K. Ganjam, K. Chakrabarti, and S. Chaudhuri, “Infogather: entity augmentation and attribute discovery by holistic matching with web tables,” in Proceedings of the 2012 ACM SIGMOD International Conference on Management of Data , 2012, pp. 97–108
2012
Earlier work this paper cites.
D. W. Barowy, C. Curtsinger, E. D. Berger, and A. McGregor, “Automan: a platform for integrating human-based and digital computation,” SIGPLAN Not. , vol. 47, no. 10, p. 639–654, oct 2012. [Online]. Available: https://doi.org/10.1145/2398857.2384663
2012
Earlier work this paper cites.
A. G. Parameswaran, H. Park, H. Garcia-Molina, N. Polyzotis, and J. Widom, “Deco: declarative crowdsourcing,” in Proceedings of the 21st ACM International Conference on Information and Knowledge Management , ser. CIKM ’12. New York, NY, USA: Association for Computing Machinery, 2012, p. 1203–1212. [Online]. Available: https://doi.org/10.1145/2396761.2398421
2012
Earlier work this paper cites.
J. Wang, T. Kraska, M. J. Franklin, and J. Feng, “Crowder: Crowdsourcing entity resolution,” 2012
2012
Earlier work this paper cites.
J. Winn, “Open data and the academy: An evaluation of ckan for research data management,” 2013
2013
Earlier work this paper cites.
H. Janssen, “Monte-carlo based uncertainty analysis: Sampling efficiency and sampling convergence,” Reliability Engineering System Safety , vol. 109, pp. 123–132, 2013. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S0951832012001536
2013
Earlier work this paper cites.
X.-W. Chen and X. Lin, “Big data deep learning: challenges and perspectives,” IEEE access , vol. 2, pp. 514–525, 2014
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
J. Pennington, R. Socher, and C. D. Manning, “Glove: Global vectors for word representation,” in Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP) , 2014, pp. 1532–1543
2014
Earlier work this paper cites.
B. Kämpgen, T. Weller, S. O’Riain, C. Weber, and A. Harth, “Accepting the xbrl challenge with linked data for financial data integration,” in The Semantic Web: Trends and Challenges , V. Presutti, C. d’Amato, F. Gandon, M. d’Aquin, S. Staab, and A. Tordai, Eds. Cham: Springer International Publishing, 2014, pp. 595–610
2014
Earlier work this paper cites.
Y. Tong, C. C. Cao, C. J. Zhang, Y. Li, and L. Chen, “Crowdcleaner: Data cleaning for multi-version data on the web via crowdsourcing,” in 2014 IEEE 30th International Conference on Data Engineering , 2014, pp. 1182–1185
2014
Earlier work this paper cites.
X. Dong, E. Gabrilovich, G. Heitz, W. Horn, N. Lao, K. Murphy, T. Strohmann, S. Sun, and W. Zhang, “Knowledge vault: a web-scale approach to probabilistic knowledge fusion,” in Proceedings of the 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , ser. KDD ’14. New York, NY, USA: Association for Computing Machinery, 2014, p. 601–610. [Online]. Available: https://doi.org/10.1145/2623330.2623623
2014
Earlier work this paper cites.
Y. LeCun, Y. Bengio, and G. Hinton, “Deep learning,” nature , vol. 521, no. 7553, pp. 436–444, 2015
2015
Earlier work this paper cites.
A. Halevy, F. Korn, N. F. Noy, C. Olston, N. Polyzotis, S. Roy, and S. E. Whang, “Goods: Organizing google’s datasets,” in Proceedings of the 2016 International Conference on Management of Data , 2016, pp. 795–806
2016
Earlier work this paper cites.
D. Ritze, O. Lehmberg, Y. Oulabi, and C. Bizer, “Profiling the potential of web tables for augmenting cross-domain knowledge bases,” in Proceedings of the 25th international conference on world wide web , 2016, pp. 251–261
2016
Earlier work this paper cites.
X. Chu, I. F. Ilyas, S. Krishnan, and J. Wang, “Data cleaning: Overview and emerging challenges,” in Proceedings of the 2016 International Conference on Management of Data , ser. SIGMOD ’16. New York, NY, USA: Association for Computing Machinery, 2016, p. 2201–2206. [Online]. Available: https://doi.org/10.1145/2882903.2912574
2016
Earlier work this paper cites.
H. Liu, A. Kumar T.K., J. P. Thomas, and X. Hou, “Cleaning framework for bigdata: An interactive approach for data cleaning,” in 2016 IEEE Second International Conference on Big Data Computing Service and Applications (BigDataService) , 2016, pp. 174–181
2016
Cited alongside, same era.
T. John and P. Misra, Data lake for enterprises . Packt Publishing Ltd, 2017
2017
Cited alongside, same era.
K. W. Church, “Word2vec,” Natural Language Engineering , vol. 23, no. 1, pp. 155–162, 2017
2017
Cited alongside, same era.
Z. Wang, “Machine learning methods for finding textual features of depression from publications,” 2017
2017
Cited alongside, same era.
J. C. Chang, S. Amershi, and E. Kamar, “Revolt: Collaborative crowdsourcing for labeling machine learning datasets,” in Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems , ser. CHI ’17. New York, NY, USA: Association for Computing Machinery, 2017, p. 2334–2346. [Online]. Available: https://doi.org/10.1145/3025453.3026044
R. Wu, S. Chaba, S. Sawlani, X. Chu, and S. Thirumuruganathan, “Zeroer: Entity resolution using zero labeled examples,” in Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data , ser. SIGMOD ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 1149–1164. [Online]. Available: https://doi.org/10.1145/3318464.3389743
2020
Later among the works it cites.
B. Karlaš, P. Li, R. Wu, N. M. Gürel, X. Chu, W. Wu, and C. Zhang, “Nearest neighbor classifiers over incomplete information: From certain answers to certain predictions,” 2020
2020
Later among the works it cites.
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio, “Generative adversarial networks,” Commun. ACM , vol. 63, no. 11, p. 139–144, oct 2020. [Online]. Available: https://doi.org/10.1145/3422622
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
A. J. Ratner, S. H. Bach, H. R. Ehrenberg, and C. Ré, “Snorkel: Fast training set generation for information extraction,” in Proceedings of the 2017 ACM International Conference on Management of Data , ser. SIGMOD ’17. New York, NY, USA: Association for Computing Machinery, 2017, p. 1683–1686. [Online]. Available: https://doi.org/10.1145/3035918.3056442
2017
Cited alongside, same era.
S. Gupta and V. Giri, Practical Enterprise Data Lake Insights: Handle Data-Driven Challenges in an Enterprise Big Data Lake . Apress, 2018
2018
Cited alongside, same era.
M. Cafarella, A. Halevy, H. Lee, J. Madhavan, C. Yu, D. Z. Wang, and E. Wu, “Ten years of webtables,” Proceedings of the VLDB Endowment , vol. 11, no. 12, pp. 2140–2149, 2018
2018
Cited alongside, same era.
M. Stonebraker, D. J. Abadi, A. Batkin, X. Chen, M. Cherniack, M. Ferreira, E. Lau, A. Lin, S. Madden, E. O’Neil et al. , “C-store: a column-oriented dbms,” in Making Databases Work: the Pragmatic Wisdom of Michael Stonebraker , 2018, pp. 491–518
2018
Cited alongside, same era.
R. J. Miller, “Open data integration,” Proc. VLDB Endow. , vol. 11, no. 12, p. 2130–2139, aug 2018. [Online]. Available: https://doi.org/10.14778/3229863.3240491
2018
Cited alongside, same era.
T. Mitchell, W. Cohen, E. Hruschka, P. Talukdar, B. Yang, J. Betteridge, A. Carlson, B. Dalvi, M. Gardner, B. Kisiel, J. Krishnamurthy, N. Lao, K. Mazaitis, T. Mohamed, N. Nakashole, E. Platanios, A. Ritter, M. Samadi, B. Settles, R. Wang, D. Wijaya, A. Gupta, X. Chen, A. Saparov, M. Greaves, and J. Welling, “Never-ending learning,” Commun. ACM , vol. 61, no. 5, p. 103–115, apr 2018. [Online]. Available: https://doi.org/10.1145/3191513
2018
Cited alongside, same era.
J. Chen and X. Ran, “Deep learning with edge computing: A review,” Proceedings of the IEEE , vol. 107, no. 8, pp. 1655–1674, 2019
2019
Cited alongside, same era.
2020
Later among the works it cites.
C. Badue, R. Guidolini, R. V. Carneiro, P. Azevedo, V. B. Cardoso, A. Forechi, L. Jesus, R. Berriel, T. M. Paixao, F. Mutz et al. , “Self-driving cars: A survey,” Expert systems with applications , vol. 165, p. 113816, 2021
2021
Later among the works it cites.
P. Sawadogo and J. Darmont, “On data lake architectures and metadata management,” Journal of Intelligent Information Systems , vol. 56, no. 1, pp. 97–120, 2021
2021
Later among the works it cites.
Y. Li, J. Li, Y. Suhara, J. Wang, W. Hirota, and W.-C. Tan, “Deep entity matching: Challenges and opportunities,” Journal of Data and Information Quality (JDIQ) , vol. 13, no. 1, pp. 1–17, 2021
2021
Later among the works it cites.
——, “Deep entity matching: Challenges and opportunities,” J. Data and Information Quality , vol. 13, no. 1, jan 2021. [Online]. Available: https://doi.org/10.1145/3431816
2021
Later among the works it cites.
Z. Zhao, A. Kunar, R. Birke, and L. Y. Chen, “Ctab-gan: Effective table data synthesizing,” in Proceedings of The 13th Asian Conference on Machine Learning , ser. Proceedings of Machine Learning Research, V. N. Balasubramanian and I. Tsang, Eds., vol. 157. PMLR, 17–19 Nov 2021, pp. 97–112. [Online]. Available: https://proceedings.mlr.press/v157/zhao21a.html
2021
Later among the works it cites.
R. Wu, P. Sakala, P. Li, X. Chu, and Y. He, “Demonstration of panda: a weakly supervised entity matching system,” Proceedings of the VLDB Endowment , vol. 14, no. 12, p. 2735–2738, Jul. 2021. [Online]. Available: http://dx.doi.org/10.14778/3476311.3476332
2021
Later among the works it cites.
H. Chen, J. Chen, and J. Ding, “Data evaluation and enhancement for quality improvement of machine learning,” IEEE Transactions on Reliability , vol. 70, no. 2, pp. 831–847, 2021
2021
Later among the works it cites.
R. Wu, N. Das, S. Chaba, S. Gandhi, D. H. Chau, and X. Chu, “A cluster-then-label approach for few-shot learning with application to automatic image data labeling,” J. Data and Information Quality , vol. 14, no. 3, may 2022. [Online]. Available: https://doi.org/10.1145/3491232
2022
Later among the works it cites.
B. Čuš and D. Golec, “Data lakehouse: Benefits in small and medium enterprises,” Mednarodno inovativno poslovanje= Journal of Innovative Business and Management , vol. 14, no. 2, pp. 1–10, 2022
2022
Later among the works it cites.
A. Panwar, V. Bhatnagar, M. Khari, A. W. Salehi, and G. Gupta, “A blockchain framework to secure personal health record (phr) in ibm cloud-based data lake,” Computational Intelligence and Neuroscience , vol. 2022, no. 1, p. 3045107, 2022
2022
Later among the works it cites.
S. Hao, P. Li, R. Wu, and X. Chu, “A model-agnostic approach for learning with noisy labels of arbitrary distributions,” in 2022 IEEE 38th International Conference on Data Engineering (ICDE) , 2022, pp. 1219–1231
2022
Later among the works it cites.
Y. Kim and B. Shin, “In defense of core-set: A density-aware core-set selection for active learning,” in Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining , ser. KDD ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 804–812. [Online]. Available: https://doi.org/10.1145/3534678.3539476
2022
Later among the works it cites.
J. Zhang, C.-Y. Hsieh, Y. Yu, C. Zhang, and A. Ratner, “A survey on programmatic weak supervision,” 2022
2022
Later among the works it cites.
R. Wu, A. Bendeck, X. Chu, and Y. He, “Ground truth inference for weakly supervised entity matching,” Proceedings of the ACM on Management of Data , vol. 1, no. 1, pp. 1–28, 2023
2023
Later among the works it cites.
A. C. Odimarha, S. A. Ayodeji, and E. A. Abaku, “Machine learning’s influence on supply chain and logistics optimization in the oil and gas sector: a comprehensive analysis,” Computer Science & IT Research Journal , vol. 5, no. 3, pp. 725–740, 2024
2024
Closest in time.
“Quandl - alternativedata,” https://alternativedata.org/data_provider/quandl/ , (Accessed on 06/18/2024)
2024
Closest in time.
“Data market :: There is a way to manage technology,” https://www.datamarket.com.tr/en/ , (Accessed on 06/18/2024)
2024
Closest in time.
“Kaggle: Your machine learning and data science community,” https://www.kaggle.com/ , (Accessed on 06/18/2024)
2024
Closest in time.
“Amazon mechanical turk,” https://www.mturk.com/ , (Accessed on 06/18/2024)
2024
Closest in time.