Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have exploded a new heatwave of AI for their ability to engage end-users in human-level conversations with detailed and articulate answers across many knowledge domains.
An automata-theoretic approach to automatic program verification
M. Y. Vardi and P. Wolper · 1986
Earlier work this paper cites.
Powerful techniques for the automatic generation of invariants
S. Bensalem, Y. Lakhnech, and H. Saidi · 1996
Earlier work this paper cites.
Zonohedra and zonotopes
D. Eppstein · 1996
Earlier work this paper cites.
First-Order Logic and Automated Theorem Proving, Second Edition
M. Fitting · 1996
Earlier work this paper cites.
Invest: A tool for the verification of invariants
S. Bensalem, Y. Lakhnech, and S. Owre · 1998
Earlier work this paper cites.
Compilers: principles, techniques, and tools
M. Lam, R. Sethi, J. D. Ullman, and A. Aho · 2006
Earlier work this paper cites.
Large language models in machine translation
T. Brants, A. C. Popat, P. Xu, F. J. Och, and J. Dean · 2007
Earlier work this paper cites.
Z3: An efficient smt solver
L. De Moura and N. Bjørner · 2008
Earlier work this paper cites.
Exploiting machine learning to subvert your spam filter
B. Nelson, M. Barreno, F. J. Chi, A. D. Joseph, B. I. P. Rubinstein, U. Saini, C. Sutton, J. D. Tygar, and K. Xia · 2008
Earlier work this paper cites.
Kernel fusion: An effective method for better power efficiency on multithreaded gpu
G. Wang, Y. Lin, and W. Yi · 2010
Earlier work this paper cites.
Runtime verification for ltl and tltl
A. Bauer, M. Leucker, and C. Schallhart · 2011
Earlier work this paper cites.
Machine learning basics
X. Huang, G. Jin, and W. Ruan · 2012
Earlier work this paper cites.
The temporal logic of reactive and concurrent systems: Specification
Z. Manna and A. Pnueli · 2012
Earlier work this paper cites.
Automated theorem proving
W. Bibel · 2013
Earlier work this paper cites.
Intriguing properties of neural networks
C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus · 2013
Earlier work this paper cites.
Explaining and harnessing adversarial examples
I. J. Goodfellow, J. Shlens, and C. Szegedy · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Earlier work this paper cites.
Truenorth: Design and tool flow of a 65 mw 1 million neuron programmable neurosynaptic chip
F. Akopyan, J. Sawada, A. Cassidy, R. Alvarez-Icaza, J. Arthur, P. Merolla, N. Imam, Y. Nakamura, P. Datta, G.-J. Nam, et al · 2015
Earlier work this paper cites.
A comprehensive survey on safe reinforcement learning
J. Garcıa and F. Fernández · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton, O. Vinyals, and J. Dean · 2015
Earlier work this paper cites.
A baseline for detecting misclassified and out-of-distribution examples in neural networks
D. Hendrycks and K. Gimpel · 2016
Earlier work this paper cites.
"why should I trust you?": Explaining the predictions of any classifier
M. T. Ribeiro, S. Singh, and C. Guestrin · 2016
Earlier work this paper cites.
Discrete variational autoencoders
J. T. Rolfe · 2016
Earlier work this paper cites.
Towards verified artificial intelligence
S. A. Seshia, D. Sadigh, and S. S. Sastry · 2016
Earlier work this paper cites.
Natural language processing for aviation safety reports: From classification to interactive analysis
L. Tanguy, N. Tulechki, A. Urieli, E. Hermann, and C. Raynal · 2016
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Y. Belinkov and Y. Bisk · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Hotflip: White-box adversarial examples for text classification
J. Ebrahimi, A. Rao, D. Lowd, and D. Dou · 2017
Earlier work this paper cites.
Detecting adversarial samples from artifacts
R. Feinman, R. R. Curtin, S. Shintre, and A. B. Gardner · 2017
Earlier work this paper cites.
The challenge of verification and testing of machine learning
I. Goodfellow and N. Papernot · 2017
Earlier work this paper cites.
Deceiving google’s perspective api built for detecting toxic comments
H. Hosseini, S. Kannan, B. Zhang, and R. Poovendran · 2017
Earlier work this paper cites.
Toward controlled generation of text
Z. Hu, Z. Yang, X. Liang, R. Salakhutdinov, and E. P. Xing · 2017
Earlier work this paper cites.
Safety verification of deep neural networks
X. Huang, M. Kwiatkowska, S. Wang, and M. Wu · 2017
Earlier work this paper cites.
Adversarial examples for evaluating reading comprehension systems
R. Jia and P. Liang · 2017
Earlier work this paper cites.
Deep text classification can be fooled
B. Liang, H. Li, M. Su, P. Bian, X. Li, and W. Shi · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu · 2017
Earlier work this paper cites.
Learning to generate reviews and discovering sentiment
A. Radford, R. Jozefowicz, and I. Sutskever · 2017
Earlier work this paper cites.
Conversion of continuous-valued deep networks to efficient event-driven networks for image classification
B. Rueckauer, I.-A. Lungu, Y. Hu, M. Pfeiffer, and S.-C. Liu · 2017
Earlier work this paper cites.
Towards crafting text adversarial samples
S. Samanta and S. Mehta · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Certifying some distributional robustness with principled adversarial training
A. Sinha, H. Namkoong, R. Volpi, and J. Duchi · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, ℒ \mathcal{L} . Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
Generating natural adversarial examples
Z. Zhao, D. Dua, and S. Singh · 2017
Earlier work this paper cites.
Safe reinforcement learning via shielding
M. Alshiekh, R. Bloem, R. Ehlers, B. Könighofer, S. Niekum, and U. Topcu · 2018
Earlier work this paper cites.
Generating natural language adversarial examples
M. Alzantot, Y. Sharma, A. Elgohary, B.-J. Ho, M. Srivastava, and K.-W. Chang · 2018
Earlier work this paper cites.
Lectures on runtime verification
E. Bartocci and Y. Falcone · 2018
Earlier work this paper cites.
Loihi: A neuromorphic manycore processor with on-chip learning
M. Davies, N. Srinivasa, T.-H. Lin, G. Chinya, Y. Cao, S. H. Choday, G. Dimou, P. Joshi, N. Imam, S. Jain, et al · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Earlier work this paper cites.
Learning confidence for out-of-distribution detection in neural networks
T. DeVries and G. W. Taylor · 2018
Earlier work this paper cites.
A review of user interface design for interactive machine learning
J. J. Dudley and P. O. Kristensson · 2018
Earlier work this paper cites.
Black-box generation of adversarial text sequences to evade deep learning classifiers
J. Gao, J. Lanchantin, M. L. Soffa, and Y. Qi · 2018
Earlier work this paper cites.
Symbolic execution for deep neural networks
D. Gopinath, K. Wang, M. Zhang, C. S. Pasareanu, and S. Khurshid · 2018
Earlier work this paper cites.
On the effectiveness of interval bound propagation for training verifiably robust models
S. Gowal, K. Dvijotham, R. Stanforth, R. Bunel, C. Qin, J. Uesato, R. Arandjelovic, T. Mann, and P. Kohli · 2018
Earlier work this paper cites.
Adversarial example generation with syntactically controlled paraphrase networks
M. Iyyer, J. Wieting, K. Gimpel, and L. Zettlemoyer · 2018
Earlier work this paper cites.
Shielded decision-making in mdps
N. Jansen, B. Könighofer, S. Junges, and R. Bloem · 2018
Earlier work this paper cites.
Adversarial examples for natural language classification problems
V. Kuleshov, S. Thakoor, T. Lau, and S. Ermon · 2018
Earlier work this paper cites.
Textbugger: Generating adversarial text against real-world applications
J. Li, S. Ji, T. Du, B. Li, and T. Wang · 2018
Earlier work this paper cites.
On the decision boundary of deep neural networks
Y. Li, L. Ding, and X. Gao · 2018
Earlier work this paper cites.
Enhancing the reliability of out-of-distribution image detection in neural networks
S. Liang, Y. Li, and R. Srikant · 2018
Earlier work this paper cites.
Differentiable abstract interpretation for provably robust neural networks
M. Mirman, T. Gehr, and M. Vechev · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
A. Radford, K. Narasimhan, T. Salimans, I. Sutskever, et al · 2018
Earlier work this paper cites.
Reachability analysis of deep neural networks with provable guarantees
W. Ruan, X. Huang, and M. Kwiatkowska · 2018
Earlier work this paper cites.
Understanding measures of uncertainty for adversarial example detection
L. Smith and Y. Gal · 2018
Earlier work this paper cites.
Y. Sun, X. Huang, D. Kroening, J. Sharp, M. Hill, and R. Ashmore · 2018
Earlier work this paper cites.
Concolic testing for deep neural networks
Y. Sun, M. Wu, W. Ruan, X. Huang, M. Kwiatkowska, and D. Kroening · 2018
Earlier work this paper cites.
Fever: A large-scale dataset for fact extraction and verification
J. Thorne, A. Vlachos, C. Christodoulopoulos, and A. Mittal · 2018
Earlier work this paper cites.
Robust machine comprehension models via adversarial training
Y. Wang and M. Bansal · 2018
Earlier work this paper cites.
Evaluating the robustness of neural networks: An extreme value theory approach
T.-W. Weng, H. Zhang, P.-Y. Chen, J. Yi, D. Su, Y. Gao, C.-J. Hsieh, and L. Daniel · 2018
Earlier work this paper cites.
Feature-guided black-box safety testing of deep neural networks
M. Wicker, X. Huang, and M. Kwiatkowska · 2018
Earlier work this paper cites.
Specifying and evaluating quality metrics for vision-based perception systems
A. Balakrishnan, A. G. Puranic, X. Qin, A. Dokhanchi, J. V. Deshmukh, H. Ben Amor, and G. Fainekos · 2019
Earlier work this paper cites.
Detecting backdoor attacks on deep neural networks by activation clustering
B. Chen, W. Carvalho, N. Baracaldo, H. Ludwig, B. Edwards, T. Lee, I. Molloy, and B. Srivastava · 2019
Earlier work this paper cites.
Runtime monitoring neuron activation patterns
C. Cheng, G. Nührenberg, and H. Yasuoka · 2019
Earlier work this paper cites.
Robust neural machine translation with doubly adversarial inputs
Y. Cheng, L. Jiang, and W. Macherey · 2019
Earlier work this paper cites.
Certified adversarial robustness via randomized smoothing
J. Cohen, E. Rosenfeld, and Z. Kolter · 2019
Earlier work this paper cites.
A backdoor attack against lstm-based text classification systems
J. Dai, C. Chen, and Y. Li · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
Badnets: Evaluating backdooring attacks on deep neural networks
T. Gu, K. Liu, B. Dolan-Gavitt, and S. Garg · 2019
Earlier work this paper cites.
Xai—explainable artificial intelligence
D. Gunning, M. Stefik, J. Choi, T. Miller, S. Stumpf, and G.-Z. Yang · 2019
Earlier work this paper cites.
Parameter-efficient transfer learning for nlp
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. De Laroussilhe, A. Gesmundo, M. Attariyan, and S. Gelly · 2019
Earlier work this paper cites.
Achieving verified robustness to symbol substitutions via interval bound propagation
P.-S. Huang, R. Stanforth, J. Welbl, C. Dyer, D. Yogatama, S. Gowal, K. Dvijotham, and P. Kohli · 2019
Earlier work this paper cites.
Neuroninspect: Detecting backdoors in neural networks via output explanations
X. Huang, M. Alzantot, and M. Srivastava · 2019
Earlier work this paper cites.
Adversarial examples are not bugs, they are features
A. Ilyas, S. Santurkar, D. Tsipras, L. Engstrom, B. Tran, and A. Madry · 2019
Earlier work this paper cites.
Certified robustness to adversarial word substitutions
R. Jia, A. Raghunathan, K. Göksel, and P. Liang · 2019
Earlier work this paper cites.
Popqorn: Quantifying robustness of recurrent neural networks
C.-Y. Ko, Z. Lyu, L. Weng, L. Daniel, N. Wong, and D. Lin · 2019
Earlier work this paper cites.
Albert: A lite bert for self-supervised learning of language representations
Z. Lan, M. Chen, S. Goodman, K. Gimpel, P. Sharma, and R. Soricut · 2019
Earlier work this paper cites.
Mixout: Effective regularization to finetune large-scale pretrained language models
C. Lee, K. Cho, and W. Kang · 2019
Earlier work this paper cites.
Caire: An empathetic neural chatbot
Z. Lin, P. Xu, G. I. Winata, F. B. Siddique, Z. Liu, J. Shin, and P. Fung · 2019
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov · 2019
Earlier work this paper cites.
Generalizable adversarial examples detection based on bi-model decision mismatch
J. Monteiro, I. Albuquerque, Z. Akhtar, and T. H. Falk · 2019
Earlier work this paper cites.
Adversarial nli: A new benchmark for natural language understanding
Y. Nie, A. Williams, E. Dinan, M. Bansal, J. Weston, and D. Kiela · 2019
Earlier work this paper cites.
Likelihood ratios for out-of-distribution detection
J. Ren, P. J. Liu, E. Fertig, J. Snoek, R. Poplin, M. Depristo, J. Dillon, and B. Lakshminarayanan · 2019
Earlier work this paper cites.
Generating natural language adversarial examples through probability weighted word saliency
S. Ren, Y. Deng, K. He, and W. Che · 2019
Earlier work this paper cites.
Global robustness evaluation of deep neural networks with provable guarantees for the hamming distance
W. Ruan, M. Wu, Y. Sun, X. Huang, D. Kroening, and M. Kwiatkowska · 2019
Earlier work this paper cites.
Transfer learning in natural language processing
S. Ruder, M. E. Peters, S. Swayamdipta, and T. Wolf · 2019
Earlier work this paper cites.
Robustness verification for transformers
Z. Shi, H. Zhang, K.-W. Chang, M. Huang, and C.-J. Hsieh · 2019
Earlier work this paper cites.
Structural test coverage criteria for deep neural networks
Y. Sun, X. Huang, D. Kroening, J. Sharp, M. Hill, and R. Ashmore · 2019
Earlier work this paper cites.
Survey on virtual assistant: Google assistant, siri, cortana, alexa
A. S. Tulshan and S. N. Dhage · 2019
Earlier work this paper cites.
Explainable ai: A brief survey on history, research areas, approaches and challenges
F. Xu, H. Uszkoreit, Y. Du, W. Fan, D. Zhao, and J. Zhu · 2019
Earlier work this paper cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Z. Yang, Z. Dai, Y. Yang, J. Carbonell, R. R. Salakhutdinov, and Q. V. Le · 2019
Earlier work this paper cites.
Fine-tuning language models from human preferences
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Seq2sick: Evaluating the robustness of sequence-to-sequence models with adversarial examples
M. Cheng, J. Yi, P.-Y. Chen, H. Zhang, and C.-J. Hsieh · 2020
Earlier work this paper cites.
Electra: Pre-training text encoders as discriminators rather than generators
K. Clark, M.-T. Luong, Q. V. Le, and C. D. Manning · 2020
Earlier work this paper cites.
Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks
F. Croce and M. Hein · 2020
Earlier work this paper cites.
Calibration of pre-trained transformers
S. Desai and G. Durrett · 2020
Earlier work this paper cites.
Fine-tuning pretrained language models: Weight initializations, data orders, and early stopping
J. Dodge, G. Ilharco, R. Schwartz, A. Farhadi, H. Hajishirzi, and N. Smith · 2020
Earlier work this paper cites.
Likelihood ratios and generative classifiers for unsupervised out-of-domain detection in task oriented dialog
V. Gangal, A. Arora, A. Einolghozati, and S. Gupta · 2020
Earlier work this paper cites.
Generative adversarial networks
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2020
Earlier work this paper cites.
Speaker-aware bert for multi-turn response selection in retrieval-based chatbots
J.-C. Gu, T. Li, Q. Liu, Z.-H. Ling, Z. Su, S. Wei, and X. Zhu · 2020
Earlier work this paper cites.
Pretrained transformers improve out-of-distribution robustness
D. Hendrycks, X. Liu, E. Wallace, A. Dziedzic, R. Krishnan, and D. Song · 2020
Earlier work this paper cites.
Correction of automatic speech recognition with transformer sequence-to-sequence model
O. Hrinchuk, M. Popova, and B. Ginsburg · 2020
Earlier work this paper cites.
Feature space singularity for out-of-distribution detection
H. Huang, Z. Li, L. Wang, S. Chen, B. Dong, and X. Zhou · 2020
Earlier work this paper cites.
A survey of safety and trustworthiness of deep neural networks: Verification, testing, adversarial attack and defence, and interpretability
X. Huang, D. Kroening, W. Ruan, J. Sharp, Y. Sun, E. Thamo, M. Wu, and X. Yi · 2020
Earlier work this paper cites.
A survey on contrastive self-supervised learning
A. Jaiswal, A. R. Babu, M. Z. Zadeh, D. Banerjee, and F. Makedon · 2020
Earlier work this paper cites.
Safe reinforcement learning using probabilistic shields
N. Jansen, B. Könighofer, J. Junges, A. Serban, and R. Bloem · 2020
Earlier work this paper cites.
Is bert really robust? a strong baseline for natural language attack on text classification and entailment
D. Jin, Z. Jin, J. T. Zhou, and P. Szolovits · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Syntax-guided controlled generation of paraphrases
A. Kumar, K. Ahuja, R. Vadapalli, and P. Talukdar · 2020
Earlier work this paper cites.
Weight poisoning attacks on pretrained models
K. Kurita, P. Michel, and G. Neubig · 2020
Earlier work this paper cites.
Assessing robustness of text classification through maximal safe radius computation
E. La Malfa, M. Wu, L. Laurenti, B. Wang, A. Hartshorn, and M. Kwiatkowska · 2020
Earlier work this paper cites.
Misinformation has high perplexity
N. Lee, Y. Bang, A. Madotto, and P. Fung · 2020
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2020
Cited alongside, same era.
BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer · 2020
Cited alongside, same era.
Energy-based out-of-distribution detection
W. Liu, X. Wang, J. Owens, and Y. Li · 2020
Cited alongside, same era.
Up or down? adaptive rounding for post-training quantization
M. Nagel, R. A. Amjad, M. Van Baalen, C. Louizos, and T. Blankevoort · 2020
Cited alongside, same era.
Accessed: 2023-08-20
The data protection act. https://www.legislation.gov.uk/ukpga/2018/12/contents/enacted , 2018 · 2023
Closest in time.
https://ec.europa.eu/futurium/en/ai-alliance-consultation.1.html , 2018
Ethics guidelines for trustworthy ai · 2023
Closest in time.
Accessed: 2023-08-20
China’s regulations on the administration of deep synthesis internet information services. https://www.chinalawtranslate.com/en/deep-synthesis/ , 2021 · 2023
Closest in time.
Accessed: 2023-08-20
Ai risk management framework. https://www.nist.gov/itl/ai-risk-management-framework , 2022 · 2023
Closest in time.
Accessed: 2023-08-20
China’s regulations on recommendation algorithms. http://www.cac.gov.cn/2022-01/04/c_1642894606258238.htm , 2022 · 2023
Closest in time.
Accessed: 2023-08-20
Content at scale, https://contentatscale.ai/ai-content-detector/ , 2022 · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Cited alongside, same era.
Fast is better than free: Revisiting adversarial training
E. Wong, L. Rice, and J. Z. Kolter · 2020
Cited alongside, same era.
A game-based approximate verification of deep neural networks with provable guarantees
M. Wu, M. Wicker, W. Ruan, X. Huang, and M. Kwiatkowska · 2020
Cited alongside, same era.
A deep generative distance-based classifier for out-of-domain detection with mahalanobis space
H. Xu, K. He, Y. Yan, S. Liu, Z. Liu, and W. Xu · 2020
Cited alongside, same era.
Adversarial attacks and defenses in images, graphs and text: A review
H. Xu, Y. Ma, H.-C. Liu, D. Deb, H. Liu, J.-L. Tang, and A. K. Jain · 2020
Cited alongside, same era.
Safer: A structure-free approach for certified robustness to adversarial word substitutions
M. Ye, C. Gong, and Q. Liu · 2020
Cited alongside, same era.
PEGASUS: Pre-training with extracted gap-sentences for abstractive summarization
J. Zhang, Y. Zhao, M. Saleh, and P. Liu · 2020
Cited alongside, same era.
Copyleaks, https://copyleaks.com/ai-content-detector , 2022 · 2023
Closest in time.
Accessed: 2023-08-20
New meta ai demo writes racist and inaccurate scientific literature, gets pulled. https://arstechnica.com/information-technology/2022/11/after-controversy-meta-pulls-demo-of-ai-model-that-writes-scientific-papers/ , 2022 · 2023
Closest in time.
Accessed: 2023-08-20
Originality ai, https://originality.ai , 2022 · 2023
Closest in time.
Accessed: 2023-08-20
Prompt injection attacks against gpt-3. https://simonwillison.net/2022/Sep/12/prompt-injection/ , 2022 · 2023
Closest in time.
Accessed: 2023-08-20
Blueprint for an ai bill of rights. https://www.whitehouse.gov/ostp/ai-bill-of-rights/ , 2023 · 2023
Closest in time.
Accessed: 2023-08-20
Chatgpt: get instant answers, find creative inspiration, and learn something new. https://openai.com/chatgpt , 2023 · 2023
Closest in time.
Accessed: 2023-08-23
Chatgpt: Us lawyer admits using ai for case research, 2023 · 2023
Closest in time.
Accessed: 2023-08-20
China’s algorithm registry. https://beian.cac.gov.cn/#/index , 2023 · 2023
Closest in time.
Accessed: 2023-08-20
Eu ai act. https://artificialintelligenceact.eu , 2023 · 2023
Closest in time.
Accessed: 2023-08-20
Eu data act. https://ec.europa.eu/commission/presscorner/detail/en/ip_22_1113 , 2023 · 2023
Closest in time.
Accessed: 2023-08-23
’he would still be here’: Man dies by suicide after talking with ai chatbot, widow says, 2023 · 2023
Closest in time.
Accessed: 2023-08-20
A pro-innovation approach to ai regulation. https://assets.publishing.service.gov.uk/government/uploads/system/uploads/attachment_data/file/1146542/a_pro-innovation_approach_to_AI_regulation.pdf , 2023 · 2023
Closest in time.
https://learnprompting.org/docs/prompt_hacking/leaking , 2023
Prompt leaking · 2023
Closest in time.
https://www.microsoft.com/en-us/ai/responsible-ai , 2023
Responsible ai priciples from microsoft · 2023
Closest in time.
Accessed: 2023-08-20
Three samsung employees reportedly leaked sensitive data to chatgpt. https://www.engadget.com/three-samsung-employees-reportedly-leaked-sensitive-data-to-chatgpt-190221114.html , 2023 · 2023
Closest in time.
Accessed: 2023-08-20
Understanding artificial intelligence ethics and safety: A guide for the responsible design and implementation of ai systems in the public sector., 2023 · 2023
Closest in time.
Trojanpuzzle: Covertly poisoning code-suggestion models
H. Aghakhani, W. Dai, A. Manoel, X. Fernandes, A. Kharkar, C. Kruegel, G. Vigna, D. Evans, B. Zorn, and R. Sim · 2023
Closest in time.
Can we trust the evaluation on chatgpt?
R. Aiyappa, J. An, H. Kwak, and Y.-Y. Ahn · 2023
Closest in time.
Y. Bang, S. Cahyawijaya, N. Lee, W. Dai, D. Su, B. Wilie, H. Lovenia, Z. Ji, T. Yu, W. Chung, et al · 2023
Closest in time.
What, indeed, is an achievable provable guarantee for learning-enabled safety critical systems
S. Bensalem, C.-H. Cheng, W. Huang, X. Huang, C. Wu, and X. Zhao · 2023
Closest in time.
A categorical archive of chatgpt failures
A. Borji · 2023
Closest in time.
Gpthreats-3: Is automatic malware generation a threat?
M. Botacin · 2023
Closest in time.
Introduction to red teaming large language models (llms)
M. Bullwinkle and E. Urban · 2023
Closest in time.
Attacks against machine learning — an overview
E. Bursztein · 2023
Closest in time.
Scamming the scammers: Using chatgpt to reply mails for wasting time and resources
E. Cambiaso and L. Caviglione · 2023
Closest in time.
Poisoning web-scale training datasets is practical
N. Carlini, M. Jagielski, C. A. Choquette-Choo, D. Paleka, W. Pearce, H. Anderson, A. Terzis, K. Thomas, and F. Tramèr · 2023
Closest in time.
How is chatgpt’s behavior changing over time?
L. Chen, M. Zaharia, and J. Zou · 2023
Closest in time.
The utility of chatgpt for cancer treatment information
S. Chen, B. H. Kann, M. B. Foote, H. J. Aerts, G. K. Savova, R. H. Mak, and D. S. Bitterman · 2023
Closest in time.
Fine-tuning deteriorates general textual out-of-distribution detection by distorting task-agnostic features
S. Chen, W. Yang, X. Bi, and X. Sun · 2023
Closest in time.
Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
W.-L. Chiang, Z. Li, Z. Lin, Y. Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y. Zhuang, J. E. Gonzalez, et al · 2023
Closest in time.
Toxicity in chatgpt: Analyzing persona-assigned language models
A. Deshpande, V. Murahari, T. Rajpurohit, A. Kalyan, and K. Narasimhan · 2023
Closest in time.
Gpt: A family of open, compute-efficient, large language models
N. Dey · 2023
Closest in time.
Are diffusion models vulnerable to membership inference attacks?
J. Duan, F. Kong, S. Wang, X. Shi, and K. Xu · 2023
Closest in time.
Gpt-4: Everything you want to know about openai’s new ai model
E2Analyst · 2023
Closest in time.
Study claims chatgpt is losing capability, but some experts aren’t convinced
B. EDWARDS · 2023
Closest in time.
How trustworthy is chatgpt? the case of bibliometric analyses
F. Farhat, S. Sohail, and D. Madsen · 2023
Closest in time.
Gptq: Accurate quantization for generative pre-trained transformers
E. Frantar, S. Ashkboos, T. Hoefler, and D. Alistarh · 2023
Closest in time.
Mathematical capabilities of chatgpt
S. Frieder, L. Pinchetti, R.-R. Griffiths, T. Salvatori, T. Lukasiewicz, P. C. Petersen, A. Chevalier, and J. Berner · 2023
Closest in time.
The capacity for moral self-correction in large language models
D. Ganguli, A. Askell, N. Schiefer, T. Liao, K. Lukošiūtė, A. Chen, A. Goldie, A. Mirhoseini, C. Olsson, D. Hernandez, et al · 2023
Closest in time.
Pal: Program-aided language models, 2023
L. Gao, A. Madaan, S. Zhou, U. Alon, P. Liu, Y. Yang, J. Callan, and G. Neubig · 2023
Closest in time.
Hackers are selling a service that bypasses chatgpt restrictions on malware
D. GOODIN · 2023
Closest in time.
K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz · 2023
Closest in time.
A human-centered safe robot reinforcement learning framework with interactive behaviors
S. Gu, A. Kshirsagar, Y. Du, G. Chen, Y. Yang, J. Peters, and A. Knoll · 2023
Closest in time.
Knowledge distillation of large language models
Y. Gu, L. Dong, F. Wei, and M. Huang · 2023
Closest in time.
How close is chatgpt to human experts? comparison corpus, evaluation, and detection
B. Guo, X. Zhang, Z. Wang, M. Jiang, J. Nie, Y. Ding, J. Yue, and Y. Wu · 2023
Closest in time.
Chatgpt believes it is conscious
A. Hintze · 2023
Closest in time.
Evaluating large language models on a highly-specialized topic, radiation oncology physics
J. Holmes, Z. Liu, L. Zhang, Y. Ding, T. T. Sio, L. A. McGee, J. B. Ashman, X. Li, T. Liu, J. Shen, et al · 2023
Closest in time.
Consistency analysis of chatgpt
M. Jang and T. Lukasiewicz · 2023
Closest in time.
Exploring chatgpt’s ability to rank content: A preliminary study on consistency with human preferences, 2023
Y. Ji, Y. Gong, Y. Peng, C. Ni, P. Sun, D. Pan, B. Ma, and X. Li · 2023
Closest in time.
Is chatgpt a good translator? a preliminary study
W. Jiao, W. Wang, J.-t. Huang, X. Wang, and Z. Tu · 2023
Closest in time.
Llm-assisted generation of hardware assertions
R. Kande, H. Pearce, B. Tan, B. Dolan-Gavitt, S. Thakur, R. Karri, and J. Rajendran · 2023
Closest in time.
Exploiting programmatic behavior of llms: Dual-use through standard security attacks
D. Kang, X. Li, I. Stoica, C. Guestrin, M. Zaharia, and T. Hashimoto · 2023
Closest in time.
The ethics of ai-generated maps: A study of dalle 2 and implications for cartography
Y. Kang, Q. Zhang, and R. Roth · 2023
Closest in time.
Gpt-4 passes the bar exam
D. M. Katz, M. J. Bommarito, S. Gao, and P. Arredondo · 2023
Closest in time.
How secure is code generated by chatgpt?
R. Khoury, A. R. Avila, J. Brunelle, and B. M. Camara · 2023
Closest in time.
Data and fair use
Y.-M. Kim · 2023
Closest in time.
Generating images with multimodal language models
J. Y. Koh, D. Fried, and R. Salakhutdinov · 2023
Closest in time.
Can an artificial intelligence chatbot be the author of a scholarly article?
J. Y. Lee · 2023
Closest in time.
Aligning text-to-image models using human feedback
K. Lee, H. Liu, M. Ryu, O. Watkins, Y. Du, C. Boutilier, P. Abbeel, M. Ghavamzadeh, and S. S. Gu · 2023
Closest in time.
Learning from tay’s introduction
P. Lee · 2023
Closest in time.
Multi-step jailbreaking privacy attacks on chatgpt
H. Li, D. Guo, W. Fan, M. Xu, and Y. Song · 2023
Closest in time.
Halueval: A large-scale hallucination evaluation benchmark for large language models
J. Li, X. Cheng, W. X. Zhao, J.-Y. Nie, and J.-R. Wen · 2023
Closest in time.
Evaluating the logical reasoning ability of chatgpt and gpt-4
H. Liu, R. Ning, Z. Teng, J. Liu, Q. Zhou, and Y. Zhang · 2023
Closest in time.
J. Liu, C. S. Xia, Y. Wang, and L. Zhang · 2023
Closest in time.
Summary of chatgpt/gpt-4 research and perspective towards the future of large language models
Y. Liu, T. Han, S. Ma, J. Zhang, Y. Yang, J. Tian, H. He, A. Li, M. He, Z. Liu, et al · 2023
Closest in time.
Deid-gpt: Zero-shot medical text de-identification by gpt-4
Z. Liu, X. Yu, L. Zhang, Z. Wu, C. Cao, H. Dai, L. Zhao, W. Liu, D. Shen, Q. Li, et al · 2023
Closest in time.
Is prompt all you need? no. a comprehensive and broader view of instruction learning
R. Lou, K. Zhang, and W. Yin · 2023
Closest in time.
On the educational impact of chatgpt: Is artificial intelligence ready to obtain a university degree?
K. Malinka, M. Peresíni, A. Firc, O. Hujnák, and F. Janus · 2023
Closest in time.
Adversarial prompting for black box foundation models
N. Maus, P. Chao, E. Wong, and J. Gardner · 2023
Closest in time.
Prover9 and mace4
W. McCune · 2023
Closest in time.
Announcing the next wave of ai innovation with microsoft bing and edge, May 2023
Y. Mehdi · 2023
Closest in time.
Chatgpt or human? detect and explain. explaining decisions of machine learning model for detecting short chatgpt-generated text, 2023
S. Mitrović, D. Andreoletti, and O. Ayoub · 2023
Closest in time.
Wormgpt: New ai tool allows cybercriminals to launch sophisticated cyber attacks
T. H. News · 2023
Closest in time.
Lever: Learning to verify language-to-code generation with execution
A. Ni, S. Iyer, D. Radev, V. Stoyanov, W.-t. Yih, S. Wang, and X. V. Lin · 2023
Closest in time.
GPT-4 Technical Report
OpenAI · 2023
Closest in time.
Unifying large language models and knowledge graphs: A roadmap, 2023
S. Pan, L. Luo, Y. Wang, C. Chen, J. Wang, and X. Wu · 2023
Closest in time.
Examining zero-shot vulnerability repair with large language models
H. Pearce, B. Tan, B. Ahmad, R. Karri, and B. Dolan-Gavitt · 2023
Closest in time.
To chatgpt, or not to chatgpt: That is the question!
A. Pegoraro, K. Kumari, H. Fereidooni, and A.-R. Sadeghi · 2023
Closest in time.
B. Peng, C. Li, P. He, M. Galley, and J. Gao · 2023
Closest in time.
safety analysis in the era of large language models: a case study of stpa using chatgpt
Y. Qi, X. Zhao, and X. Huang · 2023
Closest in time.
Testing the reliability of chatgpt for text annotation and classification: A cautionary remark
M. V. Reiss · 2023
Closest in time.
Pangu- σ \sigma : Towards trillion parameter language model with sparse heterogeneous computing
X. Ren, P. Zhou, X. Meng, X. Huang, Y. Wang, W. Wang, P. Li, X. Zhang, A. Podolskiy, G. Arshinov, et al · 2023
Closest in time.
The self-perception and political biases of chatgpt
J. Rutinowski, S. Franke, J. Endendyk, I. Dormuth, and M. Pauly · 2023
Closest in time.
Lost at c: A user study on the security implications of large language model code assistants
G. Sandoval, H. Pearce, T. Nys, R. Karri, S. Garg, and B. Dolan-Gavitt · 2023
Closest in time.
Senate judiciary subcommittee hearing on oversight of ai
U. Senate · 2023
Closest in time.
In chatgpt we trust? measuring and characterizing the reliability of chatgpt
X. Shen, Z. Chen, M. Backes, and Y. Zhang · 2023
Closest in time.
An analysis of the automatic bug fixing performance of chatgpt
D. Sobania, M. Briesch, C. Hanna, and J. Petke · 2023
Closest in time.
Safety assessment of chinese large language models
H. Sun, Z. Zhang, J. Deng, J. Cheng, and M. Huang · 2023
Closest in time.
Stanford alpaca: An instruction-following llama model, 2023
R. Taori, I. Gulrajani, T. Zhang, Y. Dubois, X. Li, C. Guestrin, P. Liang, and T. B. Hashimoto · 2023
Closest in time.
Defending against patch-based backdoor attacks on self-supervised learning
A. Tejankar, M. Sanjabi, Q. Wang, S. Wang, H. Firooz, H. Pirsiavash, and L. Tan · 2023
Closest in time.
Benchmarking large language models for automated verilog rtl code generation
S. Thakur, B. Ahmad, Z. Fan, H. Pearce, B. Tan, R. Karri, B. Dolan-Gavitt, and S. Garg · 2023
Closest in time.
Llama: Open and efficient foundation language models
H. Touvron, T. Lavril, G. Izacard, X. Martinet, M.-A. Lachaux, T. Lacroix, B. Rozière, N. Goyal, E. Hambro, F. Azhar, et al · 2023
Closest in time.
Understanding individual and team-based human factors in detecting deepfake texts
A. Uchendu, J. Lee, H. Shen, T. Le, T. K. Huang, and D. Lee · 2023
Closest in time.
Towards verifying the geometric robustness of large-scale neural networks
F. Wang, P. Xu, W. Ruan, and X. Huang · 2023
Closest in time.
On the robustness of chatgpt: An adversarial and out-of-distribution perspective
J. Wang, X. Hu, W. Hou, H. Chen, R. Zheng, Y. Wang, L. Yang, H. Huang, W. Ye, X. Geng, et al · 2023
Closest in time.
On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective
J. Wang, X. Hu, W. Hou, H. Chen, R. Zheng, Y. Wang, L. Yang, H. Huang, W. Ye, X. Geng, B. Jiao, Y. Zhang, and X. Xie · 2023
Closest in time.
Self-consistency improves chain of thought reasoning in language models
X. Wang, J. Wei, D. Schuurmans, Q. V. Le, E. H. Chi, S. Narang, A. Chowdhery, and D. Zhou · 2023
Closest in time.
Leveraging large language models to power chatbots for collecting user self-reported data
J. Wei, S. Kim, H. Jung, and Y.-H. Kim · 2023
Closest in time.
Neural comprehension: Language models with compiled neural networks
Y. Weng, M. Zhu, F. Xia, B. Li, S. He, K. Liu, and J. Zhao · 2023
Closest in time.
Fundamental limitations of alignment in large language models
Y. Wolf, N. Wies, Y. Levine, and A. Shashua · 2023
Closest in time.
Optimising event-driven spiking neural network with regularisation and cutoff
D. Wu, G. Jin, H. Yu, X. Yi, and X. Huang · 2023
Closest in time.
Chatgpt or grammarly? evaluating chatgpt on grammatical error correction benchmark
H. Wu, W. Wang, Y. Wan, W. Jiao, and M. Lyu · 2023
Closest in time.
Lamini-lm: A diverse herd of distilled models from large-scale instructions
M. Wu, A. Waheed, C. Zhang, M. Abdul-Mageed, and A. F. Aji · 2023
Closest in time.
Bloomberggpt: A large language model for finance
S. Wu, O. Irsoy, S. Lu, V. Dabravolski, M. Dredze, S. Gehrmann, P. Kambadur, D. Rosenberg, and G. Mann · 2023
Closest in time.
Better aligning text-to-image models with human preference
X. Wu, K. Sun, F. Zhu, R. Zhao, and H. Li · 2023
Closest in time.
Imagereward: Learning and evaluating human preferences for text-to-image generation
J. Xu, X. Liu, Y. Wu, Y. Tong, Q. Li, M. Ding, J. Tang, and Y. Dong · 2023
Closest in time.
Yandex/yalm-100b: Pretrained language model with 100b parameters
Yandex · 2023
Closest in time.
Harnessing the power of llms in practice: A survey on chatgpt and beyond
J. Yang, H. Jin, R. Tang, X. Han, Q. Feng, H. Jiang, B. Yin, and X. Hu · 2023
Closest in time.
Chinese tech giant baidu just released its answer to chatgpt, Mar 2023
Z. Yang · 2023
Closest in time.
React: Synergizing reasoning and acting in language models
S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao · 2023
Closest in time.
Model-agnostic reachability analysis on deep neural networks
C. Zhang, W. Ruan, F. Wang, P. Xu, G. Min, and X. Huang · 2023
Closest in time.
Reachability analysis of neural network control systems
C. Zhang, W. Ruan, and P. Xu · 2023
Closest in time.
Benchmarking large language models for news summarization
T. Zhang, F. Ladhak, E. Durmus, P. Liang, K. McKeown, and T. B. Hashimoto · 2023
Closest in time.
R. Zhao, X. Li, Y. K. Chia, B. Ding, and L. Bing · 2023
Closest in time.
A survey of large language models
W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al · 2023
Closest in time.
Can chatgpt understand too? a comparative study on chatgpt and fine-tuned bert
Q. Zhong, L. Ding, J. Liu, B. Du, and D. Tao · 2023
Closest in time.
A comprehensive survey on pretrained foundation models: A history from bert to chatgpt
C. Zhou, Q. Li, C. Li, J. Yu, Y. Liu, G. Wang, K. Zhang, C. Ji, Q. Yan, L. He, et al · 2023
Closest in time.
Spikegpt: Generative pre-trained language model with spiking neural networks
R.-J. Zhu, Q. Zhao, and J. K. Eshraghian · 2023
Closest in time.