Fetching the paper…
Reading the bibliography…
This technical report aims to fill a deficiency in the assessment of large multimodal models (LMMs) by specifically examining the self-consistency of their outputs when subjected to common corruptions.
Text processing like humans do: Visually attacking and shielding NLP systems
Eger, S., Sahin, G. G., Rücklé, A., Lee, J., Schulz, C., Mesgar, M., Swarnkar, K., Simpson, E., and Gurevych, I · 1903
Earlier work this paper cites.
Sentence-bert: Sentence embeddings using siamese bert-networks
Reimers, N. and Gurevych, I · 1908
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L · 2014
Earlier work this paper cites.
Esc: Dataset for environmental sound classification
Piczak, K. J · 2015
Earlier work this paper cites.
Benchmarking neural network robustness to common corruptions and perturbations
Hendrycks, D. and Dietterich, T · 2018
Earlier work this paper cites.
Natural tts synthesis by conditioning wavenet on mel spectrogram predictions
Shen, J., Pang, R., Weiss, R. J., Schuster, M., Jaitly, N., Yang, Z., Chen, Z., Zhang, Y., Wang, Y., Skerrv-Ryan, R., et al · 2018
Earlier work this paper cites.
ESPnet: End-to-end speech processing toolkit
Watanabe, S., Hori, T., Karita, S., Hayashi, T., Nishitoba, J., Unno, Y., Enrique Yalta Soplin, N., Heymann, J., Wiesner, M., Chen, N., Renduchintala, A., and Ochiai, T · 2018
Earlier work this paper cites.
Nemo: a toolkit for building ai applications using neural modules
Kuchaiev, O., Li, J., Nguyen, H., Hrinchuk, O., Leary, R., Ginsburg, B., Kriman, S., Beliaev, S., Lavrukhin, V., Cook, J., et al · 2019
Earlier work this paper cites.
Nlp augmentation
Ma, E · 2019
Earlier work this paper cites.
Benchmarking robustness in object detection: Autonomous driving when winter is coming
Michaelis, C., Mitzkus, B., Geirhos, R., Rusak, E., Bringmann, O., Ecker, A. S., Bethge, M., and Brendel, W · 2019
Earlier work this paper cites.
Facebook fair’s wmt19 news translation task submission
Ng, N., Yee, K., Baevski, A., Ott, M., Auli, M., and Edunov, S · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Data augmentation using back-translation for context-aware neural machine translation
Sugiyama, A. and Yoshinaga, N · 2019
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Ardila, R., Branson, M., Davis, K., Kohler, M., Meyer, J., Henretty, M., Morais, R., Saunders, L., Tyers, F., and Weber, G · 2020
Earlier work this paper cites.
wav2vec 2.0: A framework for self-supervised learning of speech representations
Baevski, A., Zhou, Y., Mohamed, A., and Auli, M · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Unsupervised cross-lingual representation learning for speech recognition
Conneau, A., Baevski, A., Collobert, R., Mohamed, A., and Auli, M · 2020
Earlier work this paper cites.
Benchmarking adversarial robustness on image classification
Dong, Y., Fu, Q.-A., Yang, X., Pang, T., Su, H., Xiao, Z., and Zhu, J · 2020
Earlier work this paper cites.
Conformer: Convolution-augmented transformer for speech recognition
Gulati, A., Qin, J., Chiu, C.-C., Parmar, N., Zhang, Y., Yu, J., Han, W., Wang, S., Zhang, Z., Wu, Y., et al · 2020
Earlier work this paper cites.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
Jin, D., Jin, Z., Zhou, J. T., and Szolovits, P · 2020
Earlier work this paper cites.
Reformulating unsupervised style transfer as paraphrase generation
Krishna, K., Wieting, J., and Iyyer, M · 2020
Earlier work this paper cites.
Bert-attack: Adversarial attack against bert using bert
Li, L., Ma, R., Guo, Q., Xue, X., and Qiu, X · 2020
Earlier work this paper cites.
Improving short text classification through global augmentation methods
Marivate, V. and Sefara, T · 2020
Earlier work this paper cites.
Textattack: A framework for adversarial attacks, data augmentation, and adversarial training in nlp
Morris, J. X., Lifland, E., Yoo, J. Y., Grigsby, J., Jin, D., and Qi, Y · 2020
Earlier work this paper cites.
Fastspeech 2: Fast and high-quality end-to-end text to speech
Ren, Y., Hu, C., Tan, X., Qin, T., Zhao, S., Zhao, Z., and Liu, T.-Y · 2020
Earlier work this paper cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Ribeiro, M. T., Wu, T., Guestrin, C., and Singh, S · 2020
Earlier work this paper cites.
Fairseq s2t: Fast speech-to-text modeling with fairseq
Wang, C., Tang, Y., Ma, X., Wu, A., Popuri, S., Okhonko, D., and Pino, J · 2020
Earlier work this paper cites.
Robustbench: a standardized adversarial robustness benchmark
Croce, F., Andriushchenko, M., Sehwag, V., Debenedetti, E., Flammarion, N., Chiang, M., Mittal, P., and Hein, M · 2021
Earlier work this paper cites.
Nl-augmenter: A framework for task-sensitive natural language augmentation
Dhole, K. D., Gangal, V., Gehrmann, S., Gupta, A., Li, Z., Mahamood, S., Mahendiran, A., Mille, S., Shrivastava, A., Tan, S., et al · 2021
Earlier work this paper cites.
Cogview: Mastering text-to-image generation via transformers
Ding, M., Yang, Z., Hong, W., Zheng, W., Zhou, C., Yin, D., Lin, J., Zou, X., Shao, Z., Yang, H., et al · 2021
Earlier work this paper cites.
Robustness gym: Unifying the nlp evaluation landscape
Goel, K., Rajani, N. F., Vig, J., Taschdjian, Z., Bansal, M., and Ré, C · 2021
Earlier work this paper cites.
Fine-tuned XLSR-53 large model for speech recognition in English
Grosman, J · 2021
Earlier work this paper cites.
Natural adversarial examples
Hendrycks, D., Zhao, K., Basart, S., Steinhardt, J., and Song, D · 2021
Earlier work this paper cites.
Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech
Kim, J., Kong, J., and Son, J · 2021
Cited alongside, same era.
Fastpitch: Parallel text-to-speech with pitch prediction
Łańcucki, A · 2021
Cited alongside, same era.
A token-level reference-free hallucination detection benchmark for free-form text generation
Liu, T., Zhang, Y., Brockett, C., Mao, Y., Sui, Z., Chen, W., and Dolan, B · 2021
Cited alongside, same era.
Learning transferable visual models from natural language supervision
Radford, A., Kim, J. W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al · 2021
Cited alongside, same era.
SpeechBrain: A general-purpose speech toolkit, 2021
Ravanelli, M., Parcollet, T., Plantinga, P., Rouhe, A., Cornell, S., Lugosch, L., Subakan, C., Dawalatabad, N., Heba, A., Zhong, J., Chou, J.-C., Yeh, S.-L., Fu, S.-W., Liao, C.-F., Rastorgueva, E., Grondin, F., Aris, W., Na, H., Gao, Y., Mori, R. D., and Bengio, Y · 2021
Holistic analysis of hallucination in gpt-4v (ision): Bias and interference challenges
Cui, C., Zhou, Y., Yang, X., Wu, S., Zhang, L., Zou, J., and Yao, H · 2023
Later among the works it cites.
Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Dai, W., Li, J., Li, D., Tiong, A. M. H., Zhao, J., Wang, W., Li, B., Fung, P., and Hoi, S · 2023
Later among the works it cites.
How robust is google’s bard to adversarial image attacks?
Dong, Y., Chen, H., Chen, J., Fang, Z., Yang, X., Zhang, Y., Tian, Y., Su, H., and Zhu, J · 2023
Later among the works it cites.
Dreamlike photoreal 2.0
Dreamlike Art · 2023
Later among the works it cites.
Clap learning audio concepts from natural language supervision
Elizalde, B., Deshmukh, S., Al Ismail, M., and Wang, H · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Laion-400m: Open dataset of clip-filtered 400 million image-text pairs
Schuhmann, C., Vencu, R., Beaumont, R., Kaczmarczyk, R., Mullis, C., Katta, A., Coombes, T., Jitsev, J., and Komatsuzaki, A · 2021
Cited alongside, same era.
Self-training and pre-training are complementary for speech recognition
Xu, Q., Baevski, A., Likhomanenko, T., Tomasello, P., Conneau, A., Collobert, R., Synnaeve, G., and Auli, M · 2021
Cited alongside, same era.
Lafite: Towards language-free training for text-to-image generation
Zhou, Y., Zhang, R., Chen, C., Li, C., Tensmeyer, C., Yu, T., Gu, J., Xu, J., and Sun, T · 2021
Cited alongside, same era.
Flamingo: a visual language model for few-shot learning
Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., Ring, R., Rutherford, E., Cabi, S., Han, T., Gong, Z., Samangooei, S., Monteiro, M., Menick, J., Borgeaud, S., Brock, A., Nematzadeh, A., Sharifzadeh, S., Binkowski, M., Barreira, R., Vinyals, O., Zisserman, A., and Simonyan, K · 2022
Cited alongside, same era.
Speecht5: Unified-modal encoder-decoder pre-training for spoken language processing
Ao, J., Wang, R., Zhou, L., Wang, C., Ren, S., Wu, Y., Liu, S., Ko, T., Li, Q., Zhang, Y., et al · 2022
Cited alongside, same era.
Scaling instruction-finetuned language models
Chung, H. W., Hou, L., Longpre, S., Zoph, B., Tay, Y., Fedus, W., Li, Y., Wang, X., Dehghani, M., Brahma, S., et al · 2022
Cited alongside, same era.
Glm: General language model pretraining with autoregressive blank infilling
Du, Z., Qian, Y., Liu, X., Ding, M., Qiu, J., Yang, Z., and Tang, J · 2022
Cited alongside, same era.
Fu, C., Chen, P., Shen, Y., Qin, Y., Zhang, M., Lin, X., Yang, J., Zheng, X., Li, K., Sun, X., et al · 2023
Later among the works it cites.
Distil-whisper: Robust knowledge distillation via large-scale pseudo labelling, 2023
Gandhi, S., von Platen, P., and Rush, A. M · 2023
Later among the works it cites.
Llama-adapter v2: Parameter-efficient visual instruction model
Gao, P., Han, J., Zhang, R., Lin, Z., Geng, S., Zhou, A., Zhang, W., Lu, P., He, C., Yue, X., Li, H., and Qiao, Y · 2023
Later among the works it cites.
Multimodal-gpt: A vision and language model for dialogue with humans, 2023
Gong, T., Lyu, C., Zhang, S., Wang, Y., Zheng, M., Zhao, Q., Liu, K., Zhang, W., Luo, P., and Chen, K · 2023
Later among the works it cites.
Imagebind-llm: Multi-modality instruction tuning
Han, J., Zhang, R., Shao, W., Gao, P., Xu, P., Xiao, H., Zhang, K., Liu, C., Wen, S., Guo, Z., et al · 2023
Later among the works it cites.
Cogagent: A visual language model for gui agents, 2023
Hong, W., Wang, W., Lv, Q., Xu, J., Yu, W., Ji, J., Wang, Y., Wang, Z., Dong, Y., Ding, M., and Tang, J · 2023
Later among the works it cites.
On architectural compression of text-to-image diffusion models
Kim, B.-K., Song, H.-K., Castells, T., and Choi, S · 2023
Later among the works it cites.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Li, J., Cheng, X., Zhao, W. X., Nie, J.-Y., and Wen, J.-R · 2023
Later among the works it cites.
dreamshaper-7
Lykon · 2023
Later among the works it cites.
mpt-1b-redpajama-200b
mosaicml · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
open-chinese-llama-7b-patch
openlmlab · 2023
Later among the works it cites.
Imagenet-patch: A dataset for benchmarking machine learning robustness against adversarial patches
Pintor, M., Angioni, D., Sotgiu, A., Demetrio, L., Demontis, A., Biggio, B., and Roli, F · 2023
Later among the works it cites.
Sdxl: Improving latent diffusion models for high-resolution image synthesis
Podell, D., English, Z., Lacey, K., Blattmann, A., Dockhorn, T., Müller, J., Penna, J., and Rombach, R · 2023
Later among the works it cites.
Scaling speech technology to 1,000+ languages
Pratap, V., Tjandra, A., Shi, B., Tomasello, P., Babu, A., Kundu, S., Elkahky, A., Ni, Z., Vyas, A., Fazel-Zarandi, M., Baevski, A., Adi, Y., Zhang, X., Hsu, W.-N., Conneau, A., and Auli, M · 2023
Later among the works it cites.
openjourney-v4
prompthero · 2023
Later among the works it cites.
Fast conformer with linearly scalable attention for efficient speech recognition
Rekesh, D., Kriman, S., Majumdar, S., Noroozi, V., Juang, H., Hrinchuk, O., Kumar, A., and Ginsburg, B · 2023
Later among the works it cites.
Adversarial diffusion distillation
Sauer, A., Lorenz, D., Blattmann, A., and Rombach, R · 2023
Later among the works it cites.
kandinsky 2.2
Shakhmatov, A., Razzhigaev, A., Nikolich, A., Arkhipkin, V., Pavlov, I., Kuznetsov, A., and Dimitrov, D · 2023
Later among the works it cites.
Introducing mpt-7b: a new standard for open-source, commercially usable llms, 2023
Team, M. et al · 2023
Later among the works it cites.
Redpajama-incite-instruct-3b-v1
togethercomputer · 2023
Later among the works it cites.
How many unicorns are in this image? a safety evaluation benchmark for vision llms
Tu, H., Cui, C., Wang, Z., Zhou, Y., Zhao, B., Han, J., Zhou, W., Yao, H., and Xie, C · 2023
Later among the works it cites.
A paraphrasing model based on chatgpt paraphrases
Vorobev, M. K. V. and Kuznetsov, M · 2023
Later among the works it cites.
GLM-130b: An open bilingual pre-trained model
Zeng, A., Liu, X., Du, Z., Wang, Z., Lai, H., Ding, M., Yang, Z., Xu, Y., Zheng, W., Xia, X., Tam, W. L., Ma, Z., Xue, Y., Zhai, J., Chen, W., Liu, Z., Zhang, P., Dong, Y., and Tang, J · 2023
Later among the works it cites.
On evaluating adversarial robustness of large vision-language models
Zhao, Y., Pang, T., Du, C., Yang, X., Li, C., Cheung, N.-M., and Lin, M · 2023
Later among the works it cites.
Judging llm-as-a-judge with mt-bench and chatbot arena
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E., et al · 2023
Later among the works it cites.
Progressive knowledge distillation of stable diffusion xl using layer level loss, 2024
Gupta, Y., Jaddipal, V. V., Prabhala, H., Paul, S., and Platen, P. V · 2024
Closest in time.