Fetching the paper…
Reading the bibliography…
Web-scale pre-training datasets are the cornerstone of LLMs' success.
M. Hu and B. Liu, “Mining and summarizing customer reviews,” in Proceedings of the tenth ACM SIGKDD international conference on Knowledge discovery and data mining , 2004, pp. 168–177
2004
Earlier work this paper cites.
L. Fei-Fei, R. Fergus, and P. Perona, “Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,” in 2004 conference on computer vision and pattern recognition workshop . IEEE, 2004, pp. 178–178
2004
Earlier work this paper cites.
L. Fei-Fei, R. Fergus, and P. Perona, “Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories,” in 2004 conference on computer vision and pattern recognition workshop . IEEE, 2004, pp. 178–178
2004
Earlier work this paper cites.
B. Dolan and C. Brockett, “Automatically constructing a corpus of sentential paraphrases,” in Third international workshop on paraphrasing (IWP2005) , 2005
2005
Earlier work this paper cites.
M.-E. Nilsback and A. Zisserman, “Automated flower classification over a large number of classes,” in 2008 Sixth Indian conference on computer vision, graphics & image processing . IEEE, 2008, pp. 722–729
2008
Earlier work this paper cites.
A. Krizhevsky, G. Hinton et al. , “Learning multiple layers of features from tiny images,” 2009
2009
Earlier work this paper cites.
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE conference on computer vision and pattern recognition . Ieee, 2009, pp. 248–255
2009
Earlier work this paper cites.
S. Ben-David, J. Blitzer, K. Crammer, A. Kulesza, F. Pereira, and J. W. Vaughan, “A theory of learning from different domains,” Machine learning , vol. 79, pp. 151–175, 2010
2010
Earlier work this paper cites.
N. Srebro, K. Sridharan, and A. Tewari, “Smoothness, low noise and fast rates,” in Advances in Neural Information Processing Systems , J. Lafferty, C. Williams, J. Shawe-Taylor, R. Zemel, and A. Culotta, Eds., vol. 23. Curran Associates, Inc., 2010. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2010/file/76cf99d3614e23eabab16fb27e944bf9-Paper.pdf
2010
Earlier work this paper cites.
Y. Netzer, T. Wang, A. Coates, A. Bissacco, B. Wu, A. Y. Ng et al. , “Reading digits in natural images with unsupervised feature learning,” in NIPS workshop on deep learning and unsupervised feature learning , vol. 2011, no. 2. Granada, 2011, p. 4
2011
Earlier work this paper cites.
O. M. Parkhi, A. Vedaldi, A. Zisserman, and C. Jawahar, “Cats and dogs,” in 2012 IEEE conference on computer vision and pattern recognition . IEEE, 2012, pp. 3498–3505
2012
Earlier work this paper cites.
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts, “Recursive deep models for semantic compositionality over a sentiment treebank,” in Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing , D. Yarowsky, T. Baldwin, A. Korhonen, K. Livescu, and S. Bethard, Eds. Seattle, Washington, USA: Association for Computational Linguistics, Oct. 2013, pp. 1631–1642. [Online]. Available: https://aclanthology.org/D13-1170
2013
Earlier work this paper cites.
J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” in Proceedings of the IEEE international conference on computer vision workshops , 2013, pp. 554–561
2013
Earlier work this paper cites.
2013
Earlier work this paper cites.
S. Shalev-Shwartz and S. Ben-David, Understanding machine learning: From theory to algorithms . Cambridge university press, 2014
2014
Earlier work this paper cites.
D. Chen and C. Manning, “A fast and accurate dependency parser using neural networks,” in Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP) , A. Moschitti, B. Pang, and W. Daelemans, Eds. Doha, Qatar: Association for Computational Linguistics, Oct. 2014, pp. 740–750. [Online]. Available: https://aclanthology.org/D14-1082
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
M. Cimpoi, S. Maji, I. Kokkinos, S. Mohamed, and A. Vedaldi, “Describing textures in the wild,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2014, pp. 3606–3613
2014
Earlier work this paper cites.
G. Hinton, “Distilling the knowledge in a neural network,” arXiv preprint arXiv:1503.02531 , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 770–778
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
I. Loshchilov, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101 , 2017
2017
Earlier work this paper cites.
G. Cheng, J. Han, and X. Lu, “Remote sensing image scene classification: Benchmark and state of the art,” Proceedings of the IEEE , vol. 105, no. 10, pp. 1865–1883, 2017
2017
Earlier work this paper cites.
C. Sun, A. Shrivastava, S. Singh, and A. Gupta, “Revisiting unreasonable effectiveness of data in deep learning era,” in Proceedings of the IEEE international conference on computer vision , 2017, pp. 843–852
2017
Earlier work this paper cites.
B. S. Veeling, J. Linmans, J. Winkens, T. Cohen, and M. Welling, “Rotation equivariant cnns for digital pathology,” in Medical Image Computing and Computer Assisted Intervention–MICCAI 2018: 21st International Conference, Granada, Spain, September 16-20, 2018, Proceedings, Part II 11 . Springer, 2018, pp. 210–218
2018
Earlier work this paper cites.
A. Gokaslan, V. Cohen, E. Pavlick, and S. Tellex, “Openwebtext corpus,” 2019
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
T. Pires, E. Schlinger, and D. Garrette, “How multilingual is multilingual BERT?” in Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , A. Korhonen, D. Traum, and L. Màrquez, Eds. Florence, Italy: Association for Computational Linguistics, Jul. 2019, pp. 4996–5001. [Online]. Available: https://aclanthology.org/P19-1493
2019
Earlier work this paper cites.
X. Zhang, F. Chen, C.-T. Lu, and N. Ramakrishnan, “Mitigating uncertainty in document classification,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , J. Burstein, C. Doran, and T. Solorio, Eds. Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 3126–3136. [Online]. Available: https://aclanthology.org/N19-1316
2019
Earlier work this paper cites.
M. T. Pilehvar and J. Camacho-Collados, “WiC: the word-in-context dataset for evaluating context-sensitive meaning representations,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) , J. Burstein, C. Doran, and T. Solorio, Eds. Minneapolis, Minnesota: Association for Computational Linguistics, Jun. 2019, pp. 1267–1273. [Online]. Available: https://aclanthology.org/N19-1128
2019
Earlier work this paper cites.
P. Helber, B. Bischke, A. Dengel, and D. Borth, “Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,” IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing , vol. 12, no. 7, pp. 2217–2226, 2019
2019
Earlier work this paper cites.
M. Tan and Q. Le, “EfficientNet: Rethinking model scaling for convolutional neural networks,” in Proceedings of the 36th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, K. Chaudhuri and R. Salakhutdinov, Eds., vol. 97. PMLR, 09–15 Jun 2019, pp. 6105–6114. [Online]. Available: https://proceedings.mlr.press/v97/tan19a.html
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
E. A. Chi, J. Hewitt, and C. D. Manning, “Finding universal grammatical relations in multilingual BERT,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, Eds. Online: Association for Computational Linguistics, Jul. 2020, pp. 5564–5577. [Online]. Available: https://aclanthology.org/2020.acl-main.493
2020
Earlier work this paper cites.
Z. Wang, J. Wohlwend, and T. Lei, “Structured pruning of large language models,” in Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP) , B. Webber, T. Cohn, Y. He, and Y. Liu, Eds. Online: Association for Computational Linguistics, Nov. 2020, pp. 6151–6162. [Online]. Available: https://aclanthology.org/2020.emnlp-main.496
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” Journal of machine learning research , vol. 21, no. 140, pp. 1–67, 2020
2020
Earlier work this paper cites.
M. Lukasik, S. Bhojanapalli, A. Menon, and S. Kumar, “Does label smoothing mitigate label noise?” in Proceedings of the 37th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, H. D. III and A. Singh, Eds., vol. 119. PMLR, 13–18 Jul 2020, pp. 6448–6458. [Online]. Available: https://proceedings.mlr.press/v119/lukasik20a.html
2020
Earlier work this paper cites.
H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and T. Zhao, “SMART: Robust and efficient fine-tuning for pre-trained natural language models through principled regularized optimization,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics , D. Jurafsky, J. Chai, N. Schluter, and J. Tetreault, Eds. Online: Association for Computational Linguistics, Jul. 2020, pp. 2177–2190. [Online]. Available: https://aclanthology.org/2020.acl-main.197
2020
Earlier work this paper cites.
A. Samuels and J. Mcgonical, “News sentiment analysis,” arXiv preprint arXiv:2007.02238 , 2020
2020
Earlier work this paper cites.
P. Kavumba, N. Inoue, B. Heinzerling, K. Singh, P. Reisert, and K. Inui, “Balanced copa: Countering superficial cues in causal reasoning,” Association for Natural Language Processing , pp. 1105–1108, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
N. Saunshi, S. Malladi, and S. Arora, “A mathematical exploration of why language models help solve downstream tasks,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=vVjIW3sEc1s
2021
Earlier work this paper cites.
C. Wei, S. M. Xie, and T. Ma, “Why do pretrained language models help in downstream tasks? an analysis of head and prompt tuning,” in Advances in Neural Information Processing Systems , M. Ranzato, A. Beygelzimer, Y. Dauphin, P. Liang, and J. W. Vaughan, Eds., vol. 34. Curran Associates, Inc., 2021, pp. 16 158–16 170. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2021/file/86b3e165b8154656a71ffe8a327ded7d-Paper.pdf
2021
Earlier work this paper cites.
Z. Xie, L. Yuan, Z. Zhu, and M. Sugiyama, “Positive-negative momentum: Manipulating stochastic gradient noise to improve generalization,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 11 448–11 458. [Online]. Available: https://proceedings.mlr.press/v139/xie21h.html
2021
Earlier work this paper cites.
H. Hua, X. Li, D. Dou, C. Xu, and J. Luo, “Noise stability regularization for improving BERT fine-tuning,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies , K. Toutanova, A. Rumshisky, L. Zettlemoyer, D. Hakkani-Tur, I. Beltagy, S. Bethard, R. Cotterell, T. Chakraborty, and Y. Zhou, Eds. Online: Association for Computational Linguistics, Jun. 2021, pp. 3229–3241. [Online]. Available: https://aclanthology.org/2021.naacl-main.258
2021
Earlier work this paper cites.
Z. Xie, I. Sato, and M. Sugiyama, “A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=wXgk_iCiYGo
2021
Earlier work this paper cites.
C. Baldassi, C. Lauditi, E. M. Malatesta, G. Perugini, and R. Zecchina, “Unveiling the structure of wide flat minima in neural networks,” Phys. Rev. Lett. , vol. 127, p. 278301, Dec 2021. [Online]. Available: https://link.aps.org/doi/10.1103/PhysRevLett.127.278301
2021
Cited alongside, same era.
J. Li, Z. Du, L. Zhu, Z. Ding, K. Lu, and H. T. Shen, “Divergence-agnostic unsupervised domain adaptation by adversarial attacks,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 11, pp. 8196–8211, 2021
2021
Cited alongside, same era.
S. M. Xie, T. Ma, and P. Liang, “Composed fine-tuning: Freezing pre-trained denoising autoencoders for improved generalization,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 11 424–11 435. [Online]. Available: https://proceedings.mlr.press/v139/xie21f.html
2021
Cited alongside, same era.
S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I. S. Kweon, and S. Xie, “Convnext v2: Co-designing and scaling convnets with masked autoencoders,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 16 133–16 142
2023
Later among the works it cites.
S. Longpre, G. Yauney, E. Reif, K. Lee, A. Roberts, B. Zoph, D. Zhou, J. Wei, K. Robinson, D. Mimno, and D. Ippolito, “A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , K. Duh, H. Gomez, and S. Bethard, Eds. Mexico City, Mexico: Association for Computational Linguistics, Jun. 2024, pp. 3245–3276. [Online]. Available: https://aclanthology.org/2024.naacl-long.179
2024
Later among the works it cites.
Y. Elazar, A. Bhagia, I. H. Magnusson, A. Ravichander, D. Schwenk, A. Suhr, E. P. Walsh, D. Groeneveld, L. Soldaini, S. Singh, H. Hajishirzi, N. A. Smith, and J. Dodge, “What’s in my big data?” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=RvfPnOkPV4
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2021
Cited alongside, same era.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Proceedings of the 38th International Conference on Machine Learning , ser. Proceedings of Machine Learning Research, M. Meila and T. Zhang, Eds., vol. 139. PMLR, 18–24 Jul 2021, pp. 8748–8763. [Online]. Available: https://proceedings.mlr.press/v139/radford21a.html
2021
Cited alongside, same era.
D. Barrett and B. Dherin, “Implicit gradient regularization,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=3q5IqUrkcF
2021
Cited alongside, same era.
P. Foret, A. Kleiner, H. Mobahi, and B. Neyshabur, “Sharpness-aware minimization for efficiently improving generalization,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=6Tm1mposlrM
2021
Cited alongside, same era.
T. Ridnik, E. Ben-Baruch, A. Noy, and L. Zelnik-Manor, “Imagenet-21k pretraining for the masses,” in Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1) , 2021. [Online]. Available: https://openreview.net/forum?id=Zkj_VcZ6ol
2021
Cited alongside, same era.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF international conference on computer vision , 2021, pp. 10 012–10 022
2021
Cited alongside, same era.
S. Chen, C. Wang, Z. Chen, Y. Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Yoshioka, X. Xiao et al. , “Wavlm: Large-scale self-supervised pre-training for full stack speech processing,” IEEE Journal of Selected Topics in Signal Processing , vol. 16, no. 6, pp. 1505–1518, 2022
2022
Cited alongside, same era.
Z. Fang, Y. Li, J. Lu, J. Dong, B. Han, and F. Liu, “Is out-of-distribution detection learnable?” in Advances in Neural Information Processing Systems , S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh, Eds., vol. 35. Curran Associates, Inc., 2022, pp. 37 199–37 213. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2022/file/f0e91b1314fa5eabf1d7ef6d1561ecec-Paper-Conference.pdf
2022
Cited alongside, same era.
J.-F. Le Gall, Measure theory, probability, and stochastic processes . Springer, 2022
2022
Cited alongside, same era.
2024
Later among the works it cites.
Z. Allen-Zhu and Y. Li, “Physics of language models: Part 3.1, knowledge storage and extraction,” in Forty-first International Conference on Machine Learning , 2024. [Online]. Available: https://openreview.net/forum?id=5x788rqbcj
2024
Later among the works it cites.
I. Shumailov, Z. Shumaylov, Y. Zhao, N. Papernot, R. Anderson, and Y. Gal, “Ai models collapse when trained on recursively generated data,” Nature , vol. 631, no. 8022, pp. 755–759, 2024
2024
Later among the works it cites.
M. E. A. Seddik, S.-W. Chen, S. Hayou, P. Youssef, and M. A. DEBBAH, “How bad is training on synthetic data? a statistical analysis of language model collapse,” in First Conference on Language Modeling , 2024. [Online]. Available: https://openreview.net/forum?id=t3z6UlV09o
2024
Later among the works it cites.
2024
Later among the works it cites.
A. Albalak, Y. Elazar, S. M. Xie, S. Longpre, N. Lambert, X. Wang, N. Muennighoff, B. Hou, L. Pan, H. Jeong, C. Raffel, S. Chang, T. Hashimoto, and W. Y. Wang, “A survey on data selection for language models,” Transactions on Machine Learning Research , 2024, survey Certification. [Online]. Available: https://openreview.net/forum?id=XfHWcNTSHp
2024
Later among the works it cites.
L. Soldaini, R. Kinney, A. Bhagia, D. Schwenk, D. Atkinson, R. Authur, B. Bogin, K. Chandu, J. Dumas, Y. Elazar, V. Hofmann, A. Jha, S. Kumar, L. Lucy, X. Lyu, N. Lambert, I. Magnusson, J. Morrison, N. Muennighoff, A. Naik, C. Nam, M. Peters, A. Ravichander, K. Richardson, Z. Shen, E. Strubell, N. Subramani, O. Tafjord, E. Walsh, L. Zettlemoyer, N. Smith, H. Hajishirzi, I. Beltagy, D. Groeneveld, J. Dodge, and K. Lo, “Dolma: an open corpus of three trillion tokens for language model pretraining research,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , L.-W. Ku, A. Martins, and V. Srikumar, Eds. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 15 725–15 788. [Online]. Available: https://aclanthology.org/2024.acl-long.840
2024
Later among the works it cites.
H. Chen, J. Wang, A. Shah, R. Tao, H. Wei, X. Xie, M. Sugiyama, and B. Raj, “Understanding and mitigating the label noise in pre-training on downstream tasks,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=TjhUtloBZU
2024
Later among the works it cites.
B. Yang, F. Liu, Y. Zou, X. Wu, Y. Wang, and D. A. Clifton, “Zeronlg: Aligning and autoencoding domains for zero-shot multimodal and multilingual natural language generation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 8, pp. 5712–5724, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Zhang, H. Bai, H. Lin, J. Zhao, L. Hou, and C. V. Cannistraci, “Plug-and-play: An efficient post-training pruning method for large language models,” in The Twelfth International Conference on Learning Representations , 2024. [Online]. Available: https://openreview.net/forum?id=Tr0lPx9woF
2024
Later among the works it cites.
J. Zhao, M. Zhang, C. Zeng, M. Wang, X. Liu, and L. Nie, “LRQuant: Learnable and robust post-training quantization for large language models,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , L.-W. Ku, A. Martins, and V. Srikumar, Eds. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 2240–2255. [Online]. Available: https://aclanthology.org/2024.acl-long.122
2024
Later among the works it cites.
R. Jin, J. Du, W. Huang, W. Liu, J. Luan, B. Wang, and D. Xiong, “A comprehensive evaluation of quantization strategies for large language models,” in Findings of the Association for Computational Linguistics: ACL 2024 , L.-W. Ku, A. Martins, and V. Srikumar, Eds. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 12 186–12 215. [Online]. Available: https://aclanthology.org/2024.findings-acl.726
2024
Later among the works it cites.
R. Liu, H. Bai, H. Lin, Y. Li, H. Gao, Z. Xu, L. Hou, J. Yao, and C. Yuan, “IntactKV: Improving large language model quantization by keeping pivot tokens intact,” in Findings of the Association for Computational Linguistics: ACL 2024 , L.-W. Ku, A. Martins, and V. Srikumar, Eds. Bangkok, Thailand: Association for Computational Linguistics, Aug. 2024, pp. 7716–7741. [Online]. Available: https://aclanthology.org/2024.findings-acl.460
2024
Later among the works it cites.
O. Shliazhko, A. Fenogenova, M. Tikhonova, A. Kozlova, V. Mikhailov, and T. Shavrina, “mGPT: Few-shot learners go multilingual,” Transactions of the Association for Computational Linguistics , vol. 12, pp. 58–79, 2024. [Online]. Available: https://aclanthology.org/2024.tacl-1.4
2024
Later among the works it cites.
X. Zhuang, X. Cheng, and Y. Zou, “Towards explainable joint models via information theory for multiple intent detection and slot filling,” Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, pp. 19 786–19 794, Mar. 2024. [Online]. Available: https://ojs.aaai.org/index.php/AAAI/article/view/29953
2024
Later among the works it cites.
X. Zhuang, H. Li, X. Cheng, Z. Zhu, Y. Xie, and Y. Zou, “Kdpror: A knowledge-decoupling probabilistic framework for video-text retrieval,” in European Conference on Computer Vision . Springer, 2024, pp. 313–331
2024
Later among the works it cites.
Y. Xie, Z. Zhu, X. Zhuang, L. Liang, Z. Wang, and Y. Zou, “Gpa: global and prototype alignment for audio-text retrieval,” in Proc. Interspeech 2024 , 2024, pp. 5078–5082
2024
Later among the works it cites.
Y. Pan, Y. Yuan, Y. Yin, J. Shi, Z. Xu, M. Zhang, L. Shang, X. Jiang, and Q. Liu, “Preparing lessons for progressive training on language models,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 38, no. 17, 2024, pp. 18 860–18 868
2024
Later among the works it cites.
Q. Fan, H. Huang, M. Chen, H. Liu, and R. He, “Rmt: Retentive networks meet vision transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp. 5641–5651
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Zhuang, X. Cheng, L. Liang, Y. Xie, Z. Wang, Z. Huang, and Y. Zou, “Pcad: Towards asr-robust spoken language understanding via prototype calibration and asymmetric decoupling,” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , 2024, pp. 5235–5246
2024
Later among the works it cites.
X. Zhuang, Z. Wang, X. Cheng, Y. Xie, L. Liang, and Y. Zou, “Macsc: Towards multimodal-augmented pre-trained language models via conceptual prototypes and self-balancing calibration,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , 2024, pp. 8070–8083
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
Y. Chai, Q. Liu, S. Wang, Y. Sun, Q. Peng, and H. Wu, “On training data influence of GPT models,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, Eds. Miami, Florida, USA: Association for Computational Linguistics, Nov. 2024, pp. 3126–3150. [Online]. Available: https://aclanthology.org/2024.emnlp-main.183
2024
Later among the works it cites.
M. Li, Y. Zhang, Z. Li, J. Chen, L. Chen, N. Cheng, J. Wang, T. Zhou, and J. Xiao, “From quantity to quality: Boosting LLM performance with self-guided data selection for instruction tuning,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) , K. Duh, H. Gomez, and S. Bethard, Eds. Mexico City, Mexico: Association for Computational Linguistics, Jun. 2024, pp. 7602–7635. [Online]. Available: https://aclanthology.org/2024.naacl-long.421
2024
Later among the works it cites.
2024
Later among the works it cites.
J. Li, Z. Yu, Z. Du, L. Zhu, and H. T. Shen, “A comprehensive survey on source-free domain adaptation,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 8, pp. 5743–5762, 2024
2024
Later among the works it cites.
K. Wu, J. Li, L. Meng, F. Li, and K. Lu, “Online adaptive fault diagnosis with test-time domain adaptation,” IEEE Transactions on Industrial Informatics , pp. 1–11, 2024
2024
Later among the works it cites.
J. Ru, J. Tian, C. Xiao, J. Li, and H. T. Shen, “Imbalanced open set domain adaptation via moving-threshold estimation and gradual alignment,” IEEE Transactions on Multimedia , vol. 26, pp. 2504–2514, 2024
2024
Later among the works it cites.
2024
Later among the works it cites.
2024
Later among the works it cites.
X. Zhuang, Z. Zhu, Z. Chen, Y. Xie, L. Liang, and Y. Zou, “Game on tree: Visual hallucination mitigation via coarse-to-fine view tree and game theory,” in Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , 2024, pp. 17 984–18 003
2024
Later among the works it cites.
L. Tang, P. Yi, M. Chen, M. Yang, and D. Liang, “Not all texts are the same: Dynamically querying texts for scene text detection,” in Chinese Conference on Pattern Recognition and Computer Vision (PRCV) . Springer, 2024, pp. 363–377
2024
Later among the works it cites.
X. Zhuang, X. Cheng, Z. Zhu, Z. Chen, H. Li, and Y. Zou, “Towards multimodal-augmented pre-trained language models via self-balanced expectation-maximization iteration,” in Proceedings of the 32nd ACM International Conference on Multimedia , 2024, pp. 4670–4679
2024
Later among the works it cites.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
2025
Closest in time.
X. Zhuang, H. Wang, X. He, S. Fu, and H. Hu, “Semigmmpoint: Semi-supervised point cloud segmentation based on gaussian mixture models,” Pattern Recognition , vol. 158, p. 111045, 2025
2025
Closest in time.