Fetching the paper…
Reading the bibliography…
We propose TAROT, a targeted data selection framework grounded in optimal transport theory.
On the translocation of masses
Kantorovitch, L · 1958
Earlier work this paper cites.
Sinkhorn distances: Lightspeed computation of optimal transport
Cuturi, M · 2013
Earlier work this paper cites.
The cityscapes dataset for semantic urban scene understanding
Cordts, M., Omran, M., Ramos, S., Rehfeld, T., Enzweiler, M., Benenson, R., Franke, U., Roth, S., and Schiele, B · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., and Sun, J · 2016
Earlier work this paper cites.
Playing for data: Ground truth from computer games
Richter, S. R., Vineet, V., Roth, S., and Koltun, V · 2016
Earlier work this paper cites.
Rethinking atrous convolution for semantic image segmentation
Chen, L.-C · 2017
Earlier work this paper cites.
Geometric dataset distances via optimal transport
Alvarez-Melis, D. and Fusi, N · 2020
Earlier work this paper cites.
Coresets via bilevel optimization for continual learning and streaming
Borsos, Z., Mutny, M., and Krause, A · 2020
Earlier work this paper cites.
nuscenes: A multimodal dataset for autonomous driving
Caesar, H., Bankiti, V., Lang, A. H., Vora, S., Liong, V. E., Xu, Q., Krishnan, A., Pan, Y., Baldan, G., and Beijbom, O · 2020
Earlier work this paper cites.
Coresets for data-efficient training of machine learning models
Mirzasoleiman, B., Bilmes, J., and Leskovec, J · 2020
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent
Pruthi, G., Liu, F., Kale, S., and Sundararajan, M · 2020
Earlier work this paper cites.
Segment everything everywhere all at once
Zou, X., Yang, J., Zhang, H., Li, F., Li, L., Wang, J., Wang, L., Gao, J., and Lee, Y. J · 2020
Earlier work this paper cites.
Whitening for self-supervised representation learning
Ermolov, A., Siarohin, A., Sangineto, E., and Sebe, N · 2021
Earlier work this paper cites.
Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset
Ettinger, S., Cheng, S., Caine, B., Liu, C., Zhao, H., Pradhan, S., Chai, Y., Sapp, B., Qi, C. R., Zhou, Y., Yang, Z., Chouard, A., Sun, P., Ngiam, J., Vasudevan, V., McCauley, A., Shlens, J., and Anguelov, D · 2021
Earlier work this paper cites.
Latent variable sequential set transformers for joint multi-agent motion prediction
Girgis, R., Golemo, F., Codevilla, F., Weiss, M., D’Souza, J. A., Kahou, S. E., Heide, F., and Pal, C · 2021
Earlier work this paper cites.
Nuplan: A closed-loop ml-based planning benchmark for autonomous vehicles
H. Caesar, J. Kabzan, K. T. e. a · 2021
Cited alongside, same era.
Glister: Generalization based data subset selection for efficient and robust learning
Killamsetty, K., Sivasubramanian, D., Ramakrishnan, G., and Iyer, R · 2021
Cited alongside, same era.
Argoverse 2: Next generation datasets for self-driving perception and forecasting
Wilson, B., Qi, W., Agarwal, T., Lambert, J., Singh, J., Khandelwal, S., Pan, B., Kumar, R., Hartnett, A., Pontes, J. K., Ramanan, D., Carr, P., and Hays, J · 2021
Cited alongside, same era.
Deepcore: A comprehensive library for coreset selection in deep learning
Guo, C., Zhao, B., and Bai, Y · 2022
Cited alongside, same era.
Datamodels: Predicting predictions from training data
Ilyas, A., Park, S. M., Engstrom, L., Leclerc, G., and Madry, A · 2022
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models
Goal-lbp: Goal-based local behavior guided trajectory prediction for autonomous driving
Yao, Z., Li, X., Lang, B., and Chuah, M. C · 2023
Later among the works it cites.
A survey on data selection for language models
Albalak, A., Elazar, Y., Xie, S. M., Longpre, S., Lambert, N., Wang, X., Muennighoff, N., Hou, B., Pan, L., Jeong, H., et al · 2024
Closest in time.
What is your data worth to gpt? llm-scale data valuation with influence functions
Choe, S. K., Ahn, H., Bae, J., Zhao, K., Kang, M., Chung, Y., Pratapa, A., Neiswanger, W., Strubell, E., Mitamura, T., et al · 2024
Closest in time.
Efficient ensembles improve training data attribution
Deng, J., Li, T.-W., Zhang, S., and Ma, J · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Cited alongside, same era.
Moderate coreset: A universal method of data selection for real-world data-efficient deep learning
Xia, X., Liu, J., Yu, J., Shen, X., Han, B., and Liu, T · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al · 2023
Cited alongside, same era.
Free dolly: Introducing the world’s first truly open instruction-tuned llm
Conover, M., Hayes, M., Mathur, A., Xie, J., Wan, J., Shah, S., Ghodsi, A., Wendell, P., Zaharia, M., and Xin, R · 2023
Cited alongside, same era.
Lava: Data valuation without pre-specified learning algorithms
Just, H. A., Kang, F., Wang, T., Zeng, Y., Ko, M., Jin, M., and Jia, R · 2023
Cited alongside, same era.
The flan collection: Designing data and methods for effective instruction tuning
Longpre, S., Hou, L., Vu, T., Webson, A., Chung, H. W., Tay, Y., Zhou, D., Le, Q. V., Zoph, B., Wei, J., et al · 2023
Cited alongside, same era.
Wayformer: Motion forecasting via simple & efficient attention networks
Nayakanti, N., Al-Rfou, R., Zhou, A., Goel, K., Refaat, K. S., and Sapp, B · 2023
Cited alongside, same era.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Dsdm: Model-aware dataset selection with datamodels
Engstrom, L., Feldmann, A., and Madry, A · 2024
Closest in time.
Unitraj: A unified framework for scalable vehicle trajectory prediction
Feng, L., Bahari, M., Amor, K. M. B., Zablocki, É., Cord, M., and Alahi, A · 2024
Closest in time.
Most influential subset selection: Challenges, promises, and beyond
Hu, Y., Hu, P., Zhao, H., and Ma, J. W · 2024
Closest in time.
Performance scaling via optimal transport: Enabling data selection from partially revealed sources
Kang, F., Just, H. A., Sahu, A. K., and Jia, R · 2024
Closest in time.
Openassistant conversations-democratizing large language model alignment
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z. R., Stevens, K., Barhoum, A., Nguyen, D., Stanley, O., Nagyfi, R., et al · 2024
Closest in time.
Less is more: Data value estimation for visual instruction tuning
Liu, Z., Zhou, K., Zhao, W. X., Gao, D., Li, Y., and Wen, J.-R · 2024
Closest in time.
Sun, Z., Wang, Z., Halilaj, L., and Luettin, J · 2024
Closest in time.
Qwen2.5: A party of foundation models, September 2024
Team, Q · 2024
Closest in time.
Feature distribution matching by optimal transport for effective and robust coreset selection
Xiao, W., Chen, Y., Shan, Q., Wang, Y., and Su, J · 2024
Closest in time.
Metastore: Analyzing deep learning meta-data at scale
Zhang, H., Yan, B., Cao, L., Madden, S., and Rundensteiner, E · 2024
Closest in time.