Fetching the paper…
Reading the bibliography…
Progress in AI is driven largely by the scale and quality of training data.
“AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection”
Joseph Roth et al · 1901
Earlier work this paper cites.
“AVA-ActiveSpeaker: An Audio-Visual Dataset for Active Speaker Detection”
Joseph Roth et al · 1901
Earlier work this paper cites.
“The Second Conversational Intelligence Challenge (ConvAI2)”
Emily Dinan et al · 1902
Earlier work this paper cites.
“DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion”
Mor Geva, Eric Malmi, Idan Szpektor and Jonathan Berant · 1902
Earlier work this paper cites.
“The Second Conversational Intelligence Challenge (ConvAI2)”
Emily Dinan et al · 1902
Earlier work this paper cites.
“DiscoFuse: A Large-Scale Dataset for Discourse-Based Sentence Fusion”
Mor Geva, Eric Malmi, Idan Szpektor and Jonathan Berant · 1902
Earlier work this paper cites.
“Benchmarking Natural Language Understanding Services for building Conversational Agents”
Xingkun Liu, Arash Eshghi, Pawel Swietojanski and Verena Rieser · 1903
Earlier work this paper cites.
“COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis”
Yansong Tang et al · 1903
Earlier work this paper cites.
“Cross-task weakly supervised learning from instructional videos”
Dimitri Zhukov et al · 1903
Earlier work this paper cites.
“Benchmarking Natural Language Understanding Services for building Conversational Agents”
Xingkun Liu, Arash Eshghi, Pawel Swietojanski and Verena Rieser · 1903
Earlier work this paper cites.
“COIN: A Large-scale Dataset for Comprehensive Instructional Video Analysis”
Yansong Tang et al · 1903
Earlier work this paper cites.
“Cross-task weakly supervised learning from instructional videos”
Dimitri Zhukov et al · 1903
Earlier work this paper cites.
“SocialIQA: Commonsense Reasoning about Social Interactions”
Maarten Sap et al · 1904
Earlier work this paper cites.
“PAWS: Paraphrase Adversaries from Word Scrambling”
Yuan Zhang, Jason Baldridge and Luheng He · 1904
Earlier work this paper cites.
“Large Scale Holistic Video Understanding”
Ali Diba et al · 1904
Earlier work this paper cites.
“VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research”
Xin Wang et al · 1904
Earlier work this paper cites.
“SocialIQA: Commonsense Reasoning about Social Interactions”
Maarten Sap et al · 1904
Earlier work this paper cites.
“PAWS: Paraphrase Adversaries from Word Scrambling”
Yuan Zhang, Jason Baldridge and Luheng He · 1904
Earlier work this paper cites.
“Large Scale Holistic Video Understanding”
Ali Diba et al · 1904
Earlier work this paper cites.
“VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research”
Xin Wang et al · 1904
Earlier work this paper cites.
“MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms”
Aida Amini et al · 1905
Earlier work this paper cites.
“Semantically Conditioned Dialog Response Generation via Hierarchical Disentangled Self-Attention”
Wenhu Chen et al · 1905
Earlier work this paper cites.
“BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions”
Christopher Clark et al · 1905
Earlier work this paper cites.
“HellaSwag: Can a Machine Really Finish Your Sentence?”
Rowan Zellers et al · 1905
Earlier work this paper cites.
“MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms”
Aida Amini et al · 1905
Earlier work this paper cites.
“Semantically Conditioned Dialog Response Generation via Hierarchical Disentangled Self-Attention”
Wenhu Chen et al · 1905
Earlier work this paper cites.
“BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions”
Christopher Clark et al · 1905
Earlier work this paper cites.
“HellaSwag: Can a Machine Really Finish Your Sentence?”
Rowan Zellers et al · 1905
Earlier work this paper cites.
“Constrained Decoding for Neural NLG from Compositional Representations in Task-Oriented Dialogue”
Anusha Balakrishnan et al · 1906
Earlier work this paper cites.
“Neural Legal Judgment Prediction in English”
Ilias Chalkidis, Ion Androutsopoulos and Nikolaos Aletras · 1906
Earlier work this paper cites.
“Large-Scale Multi-Label Text Classification on EU Legislation”
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis and Ion Androutsopoulos · 1906
Earlier work this paper cites.
“Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model”
Alexander. Fabbri et al · 1906
Earlier work this paper cites.
“HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips”
Antoine Miech et al · 1906
Earlier work this paper cites.
“Neural Arabic Question Answering”
Hussein Mozannar, Karl Hajal, Elie Maamary and Hazem Hajj · 1906
Earlier work this paper cites.
“Generating Summaries with Topic Templates and Structured Convolutional Decoders”
Laura Perez-Beltrachini, Yang Liu and Mirella Lapata · 1906
Earlier work this paper cites.
“Explain Yourself! Leveraging Language Models for Commonsense Reasoning”
Nazneen Rajani, Bryan McCann, Caiming Xiong and Richard Socher · 1906
Earlier work this paper cites.
“BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization”
Eva Sharma, Chen Li and Lu Wang · 1906
Earlier work this paper cites.
“Constrained Decoding for Neural NLG from Compositional Representations in Task-Oriented Dialogue”
Anusha Balakrishnan et al · 1906
Earlier work this paper cites.
“Neural Legal Judgment Prediction in English”
Ilias Chalkidis, Ion Androutsopoulos and Nikolaos Aletras · 1906
Earlier work this paper cites.
“Large-Scale Multi-Label Text Classification on EU Legislation”
Ilias Chalkidis, Manos Fergadiotis, Prodromos Malakasiotis and Ion Androutsopoulos · 1906
Earlier work this paper cites.
“Multi-News: a Large-Scale Multi-Document Summarization Dataset and Abstractive Hierarchical Model”
Alexander. Fabbri et al · 1906
Earlier work this paper cites.
“HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips”
Antoine Miech et al · 1906
Earlier work this paper cites.
“Neural Arabic Question Answering”
Hussein Mozannar, Karl Hajal, Elie Maamary and Hazem Hajj · 1906
Earlier work this paper cites.
“Generating Summaries with Topic Templates and Structured Convolutional Decoders”
Laura Perez-Beltrachini, Yang Liu and Mirella Lapata · 1906
Earlier work this paper cites.
“Explain Yourself! Leveraging Language Models for Commonsense Reasoning”
Nazneen Rajani, Bryan McCann, Caiming Xiong and Richard Socher · 1906
Earlier work this paper cites.
“BIGPATENT: A Large-Scale Dataset for Abstractive and Coherent Summarization”
Eva Sharma, Chen Li and Lu Wang · 1906
Earlier work this paper cites.
Mihail Eric et al · 1907
Earlier work this paper cites.
“WinoGrande: An Adversarial Winograd Schema Challenge at Scale”
Keisuke Sakaguchi, Ronan Bras, Chandra Bhagavatula and Yejin Choi · 1907
Earlier work this paper cites.
“TWEETQA: A Social Media Focused Question Answering Dataset”
Wenhan Xiong et al · 1907
Earlier work this paper cites.
Marcely Boito et al · 1907
Earlier work this paper cites.
Mihail Eric et al · 1907
Earlier work this paper cites.
“WinoGrande: An Adversarial Winograd Schema Challenge at Scale”
Keisuke Sakaguchi, Ronan Bras, Chandra Bhagavatula and Yejin Choi · 1907
Earlier work this paper cites.
“TWEETQA: A Social Media Focused Question Answering Dataset”
Wenhan Xiong et al · 1907
Earlier work this paper cites.
Marcely Boito et al · 1907
Earlier work this paper cites.
“Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning”
Pradeep Dasigi et al · 1908
Earlier work this paper cites.
“Neural Code Search Evaluation Dataset”
Hongyu Li, Seohyun Kim and Satish Chandra · 1908
Earlier work this paper cites.
“Reasoning Over Paragraph Effects in Situations”
Kevin Lin, Oyvind Tafjord, Peter Clark and Matt Gardner · 1908
Earlier work this paper cites.
“Few-Shot Dialogue Generation Without Annotated Data: A Transfer Learning Approach”
Igor Shalyminov, Sungjin Lee, Arash Eshghi and Oliver Lemon · 1908
Earlier work this paper cites.
“Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning”
Pradeep Dasigi et al · 1908
Earlier work this paper cites.
“Neural Code Search Evaluation Dataset”
Hongyu Li, Seohyun Kim and Satish Chandra · 1908
Earlier work this paper cites.
“Reasoning Over Paragraph Effects in Situations”
Kevin Lin, Oyvind Tafjord, Peter Clark and Matt Gardner · 1908
Earlier work this paper cites.
“Few-Shot Dialogue Generation Without Annotated Data: A Transfer Learning Approach”
Igor Shalyminov, Sungjin Lee, Arash Eshghi and Oliver Lemon · 1908
Earlier work this paper cites.
“Taskmaster-1: Toward a Realistic and Diverse Dialog Dataset”
Bill Byrne et al · 1909
Earlier work this paper cites.
“Cosmos QA: Machine Reading Comprehension with Contextual Commonsense Reasoning”
Lifu Huang, Ronan Bras, Chandra Bhagavatula and Yejin Choi · 1909
Earlier work this paper cites.
“PubMedQA: A Dataset for Biomedical Research Question Answering”
Qiao Jin et al · 1909
Earlier work this paper cites.
“Language Models as Knowledge Bases?”
Fabio Petroni et al · 1909
Earlier work this paper cites.
“QuaRTz: An Open-Domain Dataset of Qualitative Relationship Questions”
Oyvind Tafjord, Matt Gardner, Kevin Lin and Peter Clark · 1909
Earlier work this paper cites.
“WIQA: A dataset for "What if…" reasoning over procedural text”
Niket Tandon et al · 1909
Earlier work this paper cites.
Tao Yu et al · 1909
Earlier work this paper cites.
Yaqin Zhou et al · 1909
Earlier work this paper cites.
“Learning the Difference that Makes a Difference with Counterfactually-Augmented Data”
Divyansh Kaushik, Eduard Hovy and Zachary. Lipton · 1909
Earlier work this paper cites.
“Towards Scalable Multi-domain Conversational Agents: The Schema-Guided Dialogue Dataset”
Abhinav Rastogi et al · 1909
Earlier work this paper cites.
“Taskmaster-1: Toward a Realistic and Diverse Dialog Dataset”
Bill Byrne et al · 1909
Earlier work this paper cites.
“Cosmos QA: Machine Reading Comprehension with Contextual Commonsense Reasoning”
Lifu Huang, Ronan Bras, Chandra Bhagavatula and Yejin Choi · 1909
Earlier work this paper cites.
“PubMedQA: A Dataset for Biomedical Research Question Answering”
Qiao Jin et al · 1909
Earlier work this paper cites.
“Language Models as Knowledge Bases?”
Fabio Petroni et al · 1909
Earlier work this paper cites.
“QuaRTz: An Open-Domain Dataset of Qualitative Relationship Questions”
Oyvind Tafjord, Matt Gardner, Kevin Lin and Peter Clark · 1909
Earlier work this paper cites.
“WIQA: A dataset for "What if…" reasoning over procedural text”
Niket Tandon et al · 1909
Earlier work this paper cites.
Tao Yu et al · 1909
Earlier work this paper cites.
Yaqin Zhou et al · 1909
Earlier work this paper cites.
“Learning the Difference that Makes a Difference with Counterfactually-Augmented Data”
Divyansh Kaushik, Eduard Hovy and Zachary. Lipton · 1909
Earlier work this paper cites.
“Towards Scalable Multi-domain Conversational Agents: The Schema-Guided Dialogue Dataset”
Abhinav Rastogi et al · 1909
Earlier work this paper cites.
“ViGGO: A Video Game Corpus for Data-To-Text Generation in Open-Domain Conversation”
Juraj Juraska, Kevin. Bowden and Marilyn Walker · 1910
Earlier work this paper cites.
“A Graph-Based Framework to Bridge Movies and Synopses”
Yu Xiong et al · 1910
Earlier work this paper cites.
“QASC: A Dataset for Question Answering via Sentence Composition”
Tushar Khot et al · 1910
Earlier work this paper cites.
“Adversarial NLI: A New Benchmark for Natural Language Understanding”
Yixin Nie et al · 1910
Earlier work this paper cites.
“ViGGO: A Video Game Corpus for Data-To-Text Generation in Open-Domain Conversation”
Juraj Juraska, Kevin. Bowden and Marilyn Walker · 1910
Earlier work this paper cites.
“A Graph-Based Framework to Bridge Movies and Synopses”
Yu Xiong et al · 1910
Earlier work this paper cites.
“QASC: A Dataset for Question Answering via Sentence Composition”
Tushar Khot et al · 1910
Earlier work this paper cites.
“Adversarial NLI: A New Benchmark for Natural Language Understanding”
Yixin Nie et al · 1910
Earlier work this paper cites.
“PIQA: Reasoning about Physical Commonsense in Natural Language”
Yonatan Bisk et al · 1911
Earlier work this paper cites.
“Conversational implicatures in English dialogue: Annotated dataset”
Elizabeth George and Radhika Mamidi · 1911
Earlier work this paper cites.
“End-to-End Trainable Non-Collaborative Dialog System”
Yu Li, Kun Qian, Weiyan Shi and Zhou Yu · 1911
Earlier work this paper cites.
“CommonGen: A Constrained Text Generation Challenge for Generative Commonsense Reasoning”
Bill Lin et al · 1911
Earlier work this paper cites.
“Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding”
Mathew Monfort et al · 1911
Earlier work this paper cites.
“PIQA: Reasoning about Physical Commonsense in Natural Language”
Yonatan Bisk et al · 1911
Earlier work this paper cites.
“Conversational implicatures in English dialogue: Annotated dataset”
Elizabeth George and Radhika Mamidi · 1911
Earlier work this paper cites.
“End-to-End Trainable Non-Collaborative Dialog System”
Yu Li, Kun Qian, Weiyan Shi and Zhou Yu · 1911
Earlier work this paper cites.
“CommonGen: A Constrained Text Generation Challenge for Generative Commonsense Reasoning”
Bill Lin et al · 1911
Earlier work this paper cites.
“Multi-Moments in Time: Learning and Interpreting Models for Multi-Action Video Understanding”
Mathew Monfort et al · 1911
Earlier work this paper cites.
“Common Voice: A Massively-Multilingual Speech Corpus”
Rosana Ardila et al · 1912
Earlier work this paper cites.
“Mimetics: Towards Understanding Human Actions Out of Context”
Philippe Weinzaepfel and Grégory Rogez · 1912
Earlier work this paper cites.
“BLiMP: The Benchmark of Linguistic Minimal Pairs for English”
Alex Warstadt et al · 1912
Earlier work this paper cites.
“Common Voice: A Massively-Multilingual Speech Corpus”
Rosana Ardila et al · 1912
Earlier work this paper cites.
“Mimetics: Towards Understanding Human Actions Out of Context”
Philippe Weinzaepfel and Grégory Rogez · 1912
Earlier work this paper cites.
“BLiMP: The Benchmark of Linguistic Minimal Pairs for English”
Alex Warstadt et al · 1912
Earlier work this paper cites.
“Untitled review”
E.. Wilson · 1914
Earlier work this paper cites.
“Untitled review”
E.. Wilson · 1914
Earlier work this paper cites.
“On the measurement of inequality”
Anthony Atkinson · 1970
Earlier work this paper cites.
“On the measurement of inequality”
Anthony Atkinson · 1970
Earlier work this paper cites.
“The ATIS Spoken Language Systems Pilot Corpus”
Charles. Hemphill, John. Godfrey and George. Doddington · 1990
Earlier work this paper cites.
“The ATIS Spoken Language Systems Pilot Corpus”
Charles. Hemphill, John. Godfrey and George. Doddington · 1990
Earlier work this paper cites.
“SWITCHBOARD: telephone speech corpus for research and development”
J.J. Godfrey, E.C. Holliman and J. McDaniel · 1992
Earlier work this paper cites.
“SWITCHBOARD: telephone speech corpus for research and development”
J.J. Godfrey, E.C. Holliman and J. McDaniel · 1992
Earlier work this paper cites.
“TIMIT Acoustic-Phonetic Continuous Speech Corpus” Artwork Size: 715776 KB Pages: 715776 KB
Garofolo, John S. et al · 1993
Earlier work this paper cites.
“TIMIT Acoustic-Phonetic Continuous Speech Corpus” Artwork Size: 715776 KB Pages: 715776 KB
Garofolo, John S. et al · 1993
Earlier work this paper cites.
“Can prosody aid the automatic classification of dialog acts in conversational speech?”
E. Shriberg et al · 1998
Earlier work this paper cites.
“Can prosody aid the automatic classification of dialog acts in conversational speech?”
E. Shriberg et al · 1998
Earlier work this paper cites.
“Dialogue Act Modeling for Automatic Tagging and Recognition of Conversational Speech”
A. Stolcke et al · 2000
Earlier work this paper cites.
“Dialogue Act Modeling for Automatic Tagging and Recognition of Conversational Speech”
A. Stolcke et al · 2000
Earlier work this paper cites.
“Learning and Evaluating Contextual Embedding of Source Code”
Aditya Kanade, Petros Maniatis, Gogul Balakrishnan and Kensen Shi · 2001
Earlier work this paper cites.
“EEV: A Large-Scale Dataset for Studying Evoked Expressions from Video”
Jennifer. Sun et al · 2001
Earlier work this paper cites.
“Learning and Evaluating Contextual Embedding of Source Code”
Aditya Kanade, Petros Maniatis, Gogul Balakrishnan and Kensen Shi · 2001
Earlier work this paper cites.
“EEV: A Large-Scale Dataset for Studying Evoked Expressions from Video”
Jennifer. Sun et al · 2001
Earlier work this paper cites.
“Detecting Code Clones with Graph Neural Networkand Flow-Augmented Abstract Syntax Tree”
Wenhan Wang et al · 2002
Earlier work this paper cites.
“ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning”
Weihao Yu, Zihang Jiang, Yanfei Dong and Jiashi Feng · 2002
Earlier work this paper cites.
“Detecting Code Clones with Graph Neural Networkand Flow-Augmented Abstract Syntax Tree”
Wenhan Wang et al · 2002
Earlier work this paper cites.
“ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning”
Weihao Yu, Zihang Jiang, Yanfei Dong and Jiashi Feng · 2002
Earlier work this paper cites.
“Corpus of spontaneous Japanese: its design and evaluation”
Kikuo Maekawa · 2003
Earlier work this paper cites.
“Omni-sourced Webly-supervised Learning for Video Recognition”
Haodong Duan et al · 2003
Earlier work this paper cites.
“XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization”
Junjie Hu et al · 2003
Earlier work this paper cites.
“VIOLIN: A Large-Scale Dataset for Video-and-Language Inference”
Jingzhou Liu et al · 2003
Earlier work this paper cites.
“Corpus of spontaneous Japanese: its design and evaluation”
Kikuo Maekawa · 2003
Earlier work this paper cites.
“Omni-sourced Webly-supervised Learning for Video Recognition”
Haodong Duan et al · 2003
Earlier work this paper cites.
“XTREME: A Massively Multilingual Multi-task Benchmark for Evaluating Cross-lingual Generalization”
Junjie Hu et al · 2003
Earlier work this paper cites.
“VIOLIN: A Large-Scale Dataset for Video-and-Language Inference”
Jingzhou Liu et al · 2003
Earlier work this paper cites.
“The Fisher Corpus: a Resource for the Next Generations of Speech-to-Text”
Christopher Cieri, David Miller and Kevin Walker · 2004
Earlier work this paper cites.
“A Local-to-Global Approach to Multi-modal Movie Scene Segmentation”
Anyi Rao et al · 2004
Earlier work this paper cites.
“FineGym: A Hierarchical Video Dataset for Fine-grained Action Understanding”
Dian Shao, Yue Zhao, Bo Dai and Dahua Lin · 2004
Earlier work this paper cites.
“Deep Multimodal Feature Encoding for Video Ordering”
Vivek Sharma, Makarand Tapaswi and Rainer Stiefelhagen · 2004
Earlier work this paper cites.
“Rapidly Bootstrapping a Question Answering Dataset for COVID-19”
Raphael Tang et al · 2004
Earlier work this paper cites.
“CORD-19: The COVID-19 Open Research Dataset”
Lucy Wang et al · 2004
Earlier work this paper cites.
“The Fisher Corpus: a Resource for the Next Generations of Speech-to-Text”
Christopher Cieri, David Miller and Kevin Walker · 2004
Earlier work this paper cites.
“A Local-to-Global Approach to Multi-modal Movie Scene Segmentation”
Anyi Rao et al · 2004
Earlier work this paper cites.
“FineGym: A Hierarchical Video Dataset for Fine-grained Action Understanding”
Dian Shao, Yue Zhao, Bo Dai and Dahua Lin · 2004
Earlier work this paper cites.
“Deep Multimodal Feature Encoding for Video Ordering”
Vivek Sharma, Makarand Tapaswi and Rainer Stiefelhagen · 2004
Earlier work this paper cites.
“Rapidly Bootstrapping a Question Answering Dataset for COVID-19”
Raphael Tang et al · 2004
Earlier work this paper cites.
“CORD-19: The COVID-19 Open Research Dataset”
Lucy Wang et al · 2004
Earlier work this paper cites.
“CSLU: 22 Languages Corpus”
Lander, T · 2005
Earlier work this paper cites.
“Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales”
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
“Living Machines: A study of atypical animacy”
Mariona Ardanuy et al · 2005
Earlier work this paper cites.
“Condensed Movies: Story Based Retrieval with Contextual Embeddings”
Max Bain, Arsha Nagrani, Andrew Brown and Andrew Zisserman · 2005
Earlier work this paper cites.
“ProtoQA: A Question Answering Dataset for Prototypical Common-Sense Reasoning”
Michael Boratko et al · 2005
Earlier work this paper cites.
“Language Models are Few-Shot Learners”
Tom. Brown et al · 2005
Earlier work this paper cites.
“Span-ConveRT: Few-shot Span Extraction for Dialog with Pretrained Conversational Representations”
Sam Coope et al · 2005
Earlier work this paper cites.
Bill Lin, Seyeon Lee, Rahul Khanna and Xiang Ren · 2005
Earlier work this paper cites.
“Question-Driven Summarization of Answers to Consumer Health Questions”
Max Savery, Asma Abacha, Soumya Gayen and Dina Demner-Fushman · 2005
Earlier work this paper cites.
“Neural CRF Model for Sentence Alignment in Text Simplification”
Chao Jiang et al · 2005
Earlier work this paper cites.
“CSLU: 22 Languages Corpus”
Lander, T · 2005
Earlier work this paper cites.
“Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales”
Bo Pang and Lillian Lee · 2005
Earlier work this paper cites.
“Living Machines: A study of atypical animacy”
Mariona Ardanuy et al · 2005
Earlier work this paper cites.
“Condensed Movies: Story Based Retrieval with Contextual Embeddings”
Max Bain, Arsha Nagrani, Andrew Brown and Andrew Zisserman · 2005
Earlier work this paper cites.
“ProtoQA: A Question Answering Dataset for Prototypical Common-Sense Reasoning”
Michael Boratko et al · 2005
Earlier work this paper cites.
“Language Models are Few-Shot Learners”
Tom. Brown et al · 2005
Earlier work this paper cites.
“Span-ConveRT: Few-shot Span Extraction for Dialog with Pretrained Conversational Representations”
Sam Coope et al · 2005
Earlier work this paper cites.
Bill Lin, Seyeon Lee, Rahul Khanna and Xiang Ren · 2005
Earlier work this paper cites.
“Question-Driven Summarization of Answers to Consumer Health Questions”
Max Savery, Asma Abacha, Soumya Gayen and Dina Demner-Fushman · 2005
Earlier work this paper cites.
“Neural CRF Model for Sentence Alignment in Text Simplification”
Chao Jiang et al · 2005
Earlier work this paper cites.
“The AMI Meeting Corpus: A Pre-announcement” Series Title: Lecture Notes in Computer Science
Jean Carletta et al · 2006
Earlier work this paper cites.
“Understanding Human Hands in Contact at Internet Scale”
Dandan Shan, Jiaqi Geng, Michelle Shu and David. Fouhey · 2006
Earlier work this paper cites.
“The AMI Meeting Corpus: A Pre-announcement” Series Title: Lecture Notes in Computer Science
Jean Carletta et al · 2006
Earlier work this paper cites.
“Understanding Human Hands in Contact at Internet Scale”
Dandan Shan, Jiaqi Geng, Michelle Shu and David. Fouhey · 2006
Earlier work this paper cites.
“CSLU: Foreign Accented English Release 1.2” Artwork Size: 1468006 KB Pages: 1468006 KB
Lander, T · 2007
Earlier work this paper cites.
“Measures of semantic similarity and relatedness in the biomedical domain”
Ted Pedersen, Serguei.. Pakhomov, Siddharth Patwardhan and Christopher. Chute · 2007
Earlier work this paper cites.
“TinyVIRAT: Low-resolution Video Action Recognition”
Ugur Demir, Yogesh. Rawat and Mubarak Shah · 2007
Earlier work this paper cites.
“MovieNet: A Holistic Dataset for Movie Understanding”
Qingqiu Huang et al · 2007
Earlier work this paper cites.
“LEMMA: A Multi-view Dataset for Learning Multi-agent Multi-task Activities”
Baoxiong Jia et al · 2007
Earlier work this paper cites.
“CoVoST 2 and Massively Multilingual Speech-to-Text Translation”
Changhan Wang, Anne Wu and Juan Pino · 2007
Earlier work this paper cites.
Xiaoxue Zang et al · 2007
Earlier work this paper cites.
Shulin Cao et al · 2007
Earlier work this paper cites.
“CSLU: Foreign Accented English Release 1.2” Artwork Size: 1468006 KB Pages: 1468006 KB
Lander, T · 2007
Earlier work this paper cites.
“Measures of semantic similarity and relatedness in the biomedical domain”
Ted Pedersen, Serguei.. Pakhomov, Siddharth Patwardhan and Christopher. Chute · 2007
Earlier work this paper cites.
“TinyVIRAT: Low-resolution Video Action Recognition”
Ugur Demir, Yogesh. Rawat and Mubarak Shah · 2007
Earlier work this paper cites.
“MovieNet: A Holistic Dataset for Movie Understanding”
Qingqiu Huang et al · 2007
Earlier work this paper cites.
“LEMMA: A Multi-view Dataset for Learning Multi-agent Multi-task Activities”
Baoxiong Jia et al · 2007
Earlier work this paper cites.
“CoVoST 2 and Massively Multilingual Speech-to-Text Translation”
Changhan Wang, Anne Wu and Juan Pino · 2007
Earlier work this paper cites.
Xiaoxue Zang et al · 2007
Earlier work this paper cites.
Shulin Cao et al · 2007
Earlier work this paper cites.
“RareAct: A video dataset of unusual interactions”
Antoine Miech et al · 2008
Earlier work this paper cites.
“RareAct: A video dataset of unusual interactions”
Antoine Miech et al · 2008
Earlier work this paper cites.
“Actions in context”
Marcin Marszalek, Ivan Laptev and Cordelia Schmid · 2009
Earlier work this paper cites.
“What are they doing? : Collective activity classification using spatio-temporal relationship among people”
Wongun Choi, Khuram Shahid and Silvio Savarese · 2009
Earlier work this paper cites.
George Awad et al · 2009
Earlier work this paper cites.
Di Jin et al · 2009
Earlier work this paper cites.
“GLUCOSE: GeneraLized and COntextualized Story Explanations”
Nasrin Mostafazadeh et al · 2009
Earlier work this paper cites.
“Hierarchical Pre-training for Sequence Labelling in Spoken Dialog”
Emile Chapuis et al · 2009
Earlier work this paper cites.
“HAA500: Human-Centric Atomic Action Dataset with Curated Videos”
Jihoon Chung et al · 2009
Earlier work this paper cites.
“A Crowdsourced Open-Source Kazakh Speech Corpus and Initial Speech Recognition Baseline”
Yerbolat Khassanov et al · 2009
Earlier work this paper cites.
“Actions in context”
Marcin Marszalek, Ivan Laptev and Cordelia Schmid · 2009
Earlier work this paper cites.
“What are they doing? : Collective activity classification using spatio-temporal relationship among people”
Wongun Choi, Khuram Shahid and Silvio Savarese · 2009
Earlier work this paper cites.
George Awad et al · 2009
Earlier work this paper cites.
Di Jin et al · 2009
Earlier work this paper cites.
“GLUCOSE: GeneraLized and COntextualized Story Explanations”
Nasrin Mostafazadeh et al · 2009
Earlier work this paper cites.
“Hierarchical Pre-training for Sequence Labelling in Spoken Dialog”
Emile Chapuis et al · 2009
Earlier work this paper cites.
“HAA500: Human-Centric Atomic Action Dataset with Curated Videos”
Jihoon Chung et al · 2009
Earlier work this paper cites.
“A Crowdsourced Open-Source Kazakh Speech Corpus and Initial Speech Recognition Baseline”
Yerbolat Khassanov et al · 2009
Earlier work this paper cites.
“Scaling Laws for Autoregressive Generative Modeling”, 2020
Tom Henighan et al · 2010
Earlier work this paper cites.
“An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”, 2021
Alexey Dosovitskiy et al · 2010
Earlier work this paper cites.
“ALLSSTAR: Archive of L1 and L2 Scripted and Spontaneous Transcripts And Recordings”, 2010
A.R. Bradlow · 2010
Earlier work this paper cites.
“Opinosis: A Graph Based Approach to Abstractive Summarization of Highly Redundant Opinions”
Kavita Ganesan, ChengXiang Zhai and Jiawei Han · 2010
Earlier work this paper cites.
“Semantic Similarity and Relatedness between Clinical Terms: An Experimental Study”
Serguei Pakhomov et al · 2010
Earlier work this paper cites.
“"I’d rather just go to bed": Understanding Indirect Answers”
Annie Louis, Dan Roth and Filip Radlinski · 2010
Earlier work this paper cites.
“STAR: A Schema-Guided Dialog Dataset for Transfer Learning”
Johannes.. Mosig, Shikib Mehri and Thomas Kober · 2010
Earlier work this paper cites.
“Biomedical Concept Relatedness – A large EHR-based benchmark”
Claudia Schulz et al · 2010
Earlier work this paper cites.
“A Short Note on the Kinetics-700-2020 Human Action Dataset”
Lucas Smaira et al · 2010
Earlier work this paper cites.
“ANLIzing the Adversarial Natural Language Inference Dataset”
Adina Williams, Tristan Thrush and Douwe Kiela · 2010
Earlier work this paper cites.
Oshin Agarwal, Heming Ge, Siamak Shakeri and Rami Al-Rfou · 2010
Earlier work this paper cites.
“DiDiSpeech: A Large Scale Mandarin Speech Corpus”
Tingwei Guo et al · 2010
Earlier work this paper cites.
“ALFWorld: Aligning Text and Embodied Environments for Interactive Learning”
Mohit Shridhar et al · 2010
Earlier work this paper cites.
“Scaling Laws for Autoregressive Generative Modeling”, 2020
Tom Henighan et al · 2010
Earlier work this paper cites.
“An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale”, 2021
Alexey Dosovitskiy et al · 2010
Earlier work this paper cites.
“ALLSSTAR: Archive of L1 and L2 Scripted and Spontaneous Transcripts And Recordings”, 2010
A.R. Bradlow · 2010
Earlier work this paper cites.
“Opinosis: A Graph Based Approach to Abstractive Summarization of Highly Redundant Opinions”
Kavita Ganesan, ChengXiang Zhai and Jiawei Han · 2010
Earlier work this paper cites.
“Semantic Similarity and Relatedness between Clinical Terms: An Experimental Study”
Serguei Pakhomov et al · 2010
Earlier work this paper cites.
“"I’d rather just go to bed": Understanding Indirect Answers”
Annie Louis, Dan Roth and Filip Radlinski · 2010
Earlier work this paper cites.
“STAR: A Schema-Guided Dialog Dataset for Transfer Learning”
Johannes.. Mosig, Shikib Mehri and Thomas Kober · 2010
Earlier work this paper cites.
“Biomedical Concept Relatedness – A large EHR-based benchmark”
Claudia Schulz et al · 2010
Earlier work this paper cites.
“A Short Note on the Kinetics-700-2020 Human Action Dataset”
Lucas Smaira et al · 2010
Earlier work this paper cites.
“ANLIzing the Adversarial Natural Language Inference Dataset”
Adina Williams, Tristan Thrush and Douwe Kiela · 2010
Earlier work this paper cites.
Oshin Agarwal, Heming Ge, Siamak Shakeri and Rami Al-Rfou · 2010
Earlier work this paper cites.
“DiDiSpeech: A Large Scale Mandarin Speech Corpus”
Tingwei Guo et al · 2010
Earlier work this paper cites.
“ALFWorld: Aligning Text and Embodied Environments for Interactive Learning”
Mohit Shridhar et al · 2010
Earlier work this paper cites.
“Evaluation of Topic Identification Methods on Arabic Corpora”
Mourad Abbas, Kamel Smaïli and D. Berkani · 2011
Earlier work this paper cites.
“HMDB: A large video database for human motion recognition”
H. Kuehne et al · 2011
Earlier work this paper cites.
Hannah Chen, Yangfeng Ji and David Evans · 2011
Earlier work this paper cites.
“HoVer: A Dataset for Many-Hop Fact Extraction And Claim Verification”
Yichen Jiang et al · 2011
Earlier work this paper cites.
“Learning from Task Descriptions”
Orion Weller, Nicholas Lourie, Matt Gardner and Matthew. Peters · 2011
Earlier work this paper cites.
“QuerYD: A video dataset with high-quality text and audio narrations”
Andreea-Maria Oncescu et al · 2011
Earlier work this paper cites.
“Evaluation of Topic Identification Methods on Arabic Corpora”
Mourad Abbas, Kamel Smaïli and D. Berkani · 2011
Earlier work this paper cites.
“HMDB: A large video database for human motion recognition”
H. Kuehne et al · 2011
Earlier work this paper cites.
Hannah Chen, Yangfeng Ji and David Evans · 2011
Earlier work this paper cites.
“HoVer: A Dataset for Many-Hop Fact Extraction And Claim Verification”
Yichen Jiang et al · 2011
Earlier work this paper cites.
“Learning from Task Descriptions”
Orion Weller, Nicholas Lourie, Matt Gardner and Matthew. Peters · 2011
Earlier work this paper cites.
“QuerYD: A video dataset with high-quality text and audio narrations”
Andreea-Maria Oncescu et al · 2011
Earlier work this paper cites.
“A Comprehensive Study of Deep Video Action Recognition”, 2020
Yi Zhu et al · 2012
Earlier work this paper cites.
“UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild”
Khurram Soomro, Amir Zamir and Mubarak Shah · 2012
Earlier work this paper cites.
“Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences”
Denis Emelin et al · 2012
Earlier work this paper cites.
“CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims”
Thomas Diggelmann et al · 2012
Earlier work this paper cites.
“A Comprehensive Study of Deep Video Action Recognition”, 2020
Yi Zhu et al · 2012
Earlier work this paper cites.
“UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild”
Khurram Soomro, Amir Zamir and Mubarak Shah · 2012
Earlier work this paper cites.
“Moral Stories: Situated Reasoning about Norms, Intents, Actions, and their Consequences”
Denis Emelin et al · 2012
Earlier work this paper cites.
“CLIMATE-FEVER: A Dataset for Verification of Real-World Climate Claims”
Thomas Diggelmann et al · 2012
Earlier work this paper cites.
“A survey of video datasets for human action and activity recognition”
Jose. Chaquet, Enrique. Carmona and Antonio Fernández-Caballero · 2013
Earlier work this paper cites.
“A Thousand Frames in Just a Few Words: Lingual Description of Videos through Latent Topics and Sparse Object Stitching”
Pradipto Das, Chenliang Xu, Richard. Doell and Jason. Corso · 2013
Earlier work this paper cites.
“Asgard: A portable architecture for multilingual dialogue systems”
Jingjing Liu, Panupong Pasupat, Scott Cyphers and Jim Glass · 2013
Earlier work this paper cites.
“Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts”
Pekka Malo et al · 2013
Earlier work this paper cites.
“Combining embedded accelerometers with computer vision for recognizing food preparation activities”
Sebastian Stein and Stephen. McKenna · 2013
Earlier work this paper cites.
“Spatial Pattern Templates for Recognition of Objects with Regular Structure”
Radim Tyleček and Radim Šára · 2013
Earlier work this paper cites.
“A survey of video datasets for human action and activity recognition”
Jose. Chaquet, Enrique. Carmona and Antonio Fernández-Caballero · 2013
Earlier work this paper cites.
“A Thousand Frames in Just a Few Words: Lingual Description of Videos through Latent Topics and Sparse Object Stitching”
Pradipto Das, Chenliang Xu, Richard. Doell and Jason. Corso · 2013
Earlier work this paper cites.
“Asgard: A portable architecture for multilingual dialogue systems”
Jingjing Liu, Panupong Pasupat, Scott Cyphers and Jim Glass · 2013
Earlier work this paper cites.
“Good Debt or Bad Debt: Detecting Semantic Orientations in Economic Texts”
Pekka Malo et al · 2013
Earlier work this paper cites.
“Combining embedded accelerometers with computer vision for recognizing food preparation activities”
Sebastian Stein and Stephen. McKenna · 2013
Earlier work this paper cites.
“Spatial Pattern Templates for Recognition of Objects with Regular Structure”
Radim Tyleček and Radim Šára · 2013
Earlier work this paper cites.
“Weakly Supervised Action Labeling in Videos Under Ordering Constraints”
Piotr Bojanowski et al · 2014
Earlier work this paper cites.
“Creating Summaries from User Videos” Series Title: Lecture Notes in Computer Science
Michael Gygli, Helmut Grabner, Hayko Riemenschneider and Luc Van · 2014
Earlier work this paper cites.
“VideoStory: A New Multimedia Embedding for Few-Example Recognition and Translation of Events”
Amirhossein Habibian, Thomas Mensink and Cees.M. Snoek · 2014
Earlier work this paper cites.
“Large-Scale Video Classification with Convolutional Neural Networks”
Andrej Karpathy et al · 2014
Earlier work this paper cites.
“Free English and Czech telephone speech corpus shared under the CC-BY-SA 3.0 license”
Matěj Korvas et al · 2014
Earlier work this paper cites.
“The Language of Actions: Recovering the Syntax and Semantics of Goal-Directed Human Activities”
Hilde Kuehne, Ali Arslan and Thomas Serre · 2014
Earlier work this paper cites.
“StoryGraphs: Visualizing Character Interactions as a Timeline”
Makarand Tapaswi, Martin Bauml and Rainer Stiefelhagen · 2014
Earlier work this paper cites.
“Weakly Supervised Action Labeling in Videos Under Ordering Constraints”
Piotr Bojanowski et al · 2014
Earlier work this paper cites.
“Creating Summaries from User Videos” Series Title: Lecture Notes in Computer Science
Michael Gygli, Helmut Grabner, Hayko Riemenschneider and Luc Van · 2014
Earlier work this paper cites.
“VideoStory: A New Multimedia Embedding for Few-Example Recognition and Translation of Events”
Amirhossein Habibian, Thomas Mensink and Cees.M. Snoek · 2014
Earlier work this paper cites.
“Large-Scale Video Classification with Convolutional Neural Networks”
Andrej Karpathy et al · 2014
Earlier work this paper cites.
“Free English and Czech telephone speech corpus shared under the CC-BY-SA 3.0 license”
Matěj Korvas et al · 2014
Earlier work this paper cites.
“The Language of Actions: Recovering the Syntax and Semantics of Goal-Directed Human Activities”
Hilde Kuehne, Ali Arslan and Thomas Serre · 2014
Earlier work this paper cites.
“StoryGraphs: Visualizing Character Interactions as a Timeline”
Makarand Tapaswi, Martin Bauml and Rainer Stiefelhagen · 2014
Earlier work this paper cites.
“Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research”
Alex Bravo et al · 2015
Earlier work this paper cites.
“Building Large Arabic Multi-domain Resources for Sentiment Analysis”
Hady ElSahar and Samhaa. El-Beltagy · 2015
Earlier work this paper cites.
“ActivityNet: A large-scale video benchmark for human activity understanding”
Fabian Heilbron, Victor Escorcia, Bernard Ghanem and Juan Niebles · 2015
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books”
Vassil Panayotov, Guoguo Chen, Daniel Povey and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
“Compositional Semantic Parsing on Semi-Structured Tables”
Panupong Pasupat and Percy Liang · 2015
Earlier work this paper cites.
“A Dataset for Movie Description”
Anna Rohrbach, Marcus Rohrbach, Niket Tandon and Bernt Schiele · 2015
Earlier work this paper cites.
“An open/free database and Benchmark for Uyghur speaker recognition”
Askar Rozi, Dong Wang, Zhiyong Zhang and Thomas Zheng · 2015
Earlier work this paper cites.
“A Neural Attention Model for Abstractive Sentence Summarization”
Alexander. Rush, Sumit Chopra and Jason Weston · 2015
Earlier work this paper cites.
“ImageNet Large Scale Visual Recognition Challenge”
Olga Russakovsky et al · 2015
Earlier work this paper cites.
“THCHS-30 : A Free Chinese Speech Corpus”
Dong Wang and Xuewei Zhang · 2015
Earlier work this paper cites.
“Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks”
Jason Weston et al · 2015
Earlier work this paper cites.
“TVSum: Summarizing web videos using titles”
Yale Song, Jordi Vallmitjana, Amanda Stent and Alejandro Jaimes · 2015
Earlier work this paper cites.
“Extraction of relations between genes and diseases from text and large-scale data analysis: implications for translational research”
Alex Bravo et al · 2015
Earlier work this paper cites.
“Building Large Arabic Multi-domain Resources for Sentiment Analysis”
Hady ElSahar and Samhaa. El-Beltagy · 2015
Earlier work this paper cites.
“ActivityNet: A large-scale video benchmark for human activity understanding”
Fabian Heilbron, Victor Escorcia, Bernard Ghanem and Juan Niebles · 2015
Earlier work this paper cites.
“Librispeech: An ASR corpus based on public domain audio books”
Vassil Panayotov, Guoguo Chen, Daniel Povey and Sanjeev Khudanpur · 2015
Earlier work this paper cites.
“Compositional Semantic Parsing on Semi-Structured Tables”
Panupong Pasupat and Percy Liang · 2015
Earlier work this paper cites.
“A Dataset for Movie Description”
Anna Rohrbach, Marcus Rohrbach, Niket Tandon and Bernt Schiele · 2015
Earlier work this paper cites.
“An open/free database and Benchmark for Uyghur speaker recognition”
Askar Rozi, Dong Wang, Zhiyong Zhang and Thomas Zheng · 2015
Earlier work this paper cites.
“A Neural Attention Model for Abstractive Sentence Summarization”
Alexander. Rush, Sumit Chopra and Jason Weston · 2015
Earlier work this paper cites.
“ImageNet Large Scale Visual Recognition Challenge”
Olga Russakovsky et al · 2015
Earlier work this paper cites.
“THCHS-30 : A Free Chinese Speech Corpus”
Dong Wang and Xuewei Zhang · 2015
Earlier work this paper cites.
“Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks”
Jason Weston et al · 2015
Earlier work this paper cites.
“TVSum: Summarizing web videos using titles”
Yale Song, Jordi Vallmitjana, Amanda Stent and Alejandro Jaimes · 2015
Earlier work this paper cites.
“Youtube-8m: A large-scale video classification benchmark”
Sami Abu-El-Haija et al · 2016
Earlier work this paper cites.
“Squad: 100,000+ questions for machine comprehension of text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
“Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding”, 2016
Gunnar. Sigurdsson et al · 2016
Earlier work this paper cites.
“YouTube-8M: A Large-Scale Video Classification Benchmark”
Sami Abu-El-Haija et al · 2016
Earlier work this paper cites.
“Unsupervised Learning from Narrated Instruction Videos”
Jean-Baptiste Alayrac et al · 2016
Earlier work this paper cites.
“BRAD 1.0: Book reviews in Arabic dataset”
Ashraf Elnagar and Omar Einea · 2016
Earlier work this paper cites.
“A Hierarchical Deep Temporal Model for Group Activity Recognition”
Moustafa Ibrahim et al · 2016
Earlier work this paper cites.
“SelQA: A New Benchmark for Selection-based Question Answering”
Tomasz Jurczyk, Michael Zhai and Jinho. Choi · 2016
Earlier work this paper cites.
“1.5 billion words Arabic Corpus”
Ibrahim El-khair · 2016
Earlier work this paper cites.
“Neural Text Generation from Structured Data with Application to the Biography Domain”
Remi Lebret, David Grangier and Michael Auli · 2016
Earlier work this paper cites.
“TGIF: A New Dataset and Benchmark on Animated GIF Description”
Yuncheng Li et al · 2016
Earlier work this paper cites.
Ryan Lowe, Nissan Pow, Iulian Serban and Joelle Pineau · 2016
Earlier work this paper cites.
“Pointer Sentinel Mixture Models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2016
Earlier work this paper cites.
“A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation” ISSN: 1063-6919
F. Perazzi et al · 2016
Earlier work this paper cites.
“SQuAD: 100,000+ Questions for Machine Comprehension of Text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
Anna Rohrbach et al · 2016
Earlier work this paper cites.
“Recognizing Fine-Grained and Composite Activities using Hand-Centric Features and Script Data”
Marcus Rohrbach et al · 2016
Earlier work this paper cites.
“NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis”
Amir Shahroudy, Jun Liu, Tian-Tsong Ng and Gang Wang · 2016
Earlier work this paper cites.
“Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding”
Gunnar. Sigurdsson et al · 2016
Earlier work this paper cites.
“MovieQA: Understanding Stories in Movies through Question-Answering”
Makarand Tapaswi et al · 2016
Earlier work this paper cites.
“MSR-VTT: A Large Video Description Dataset for Bridging Video and Language” ISSN: 1063-6919
Jun Xu, Tao Mei, Ting Yao and Yong Rui · 2016
Earlier work this paper cites.
“The Value of Semantic Parse Labeling for Knowledge Base Question Answering”
Wen-tau Yih et al · 2016
Earlier work this paper cites.
“Title Generation for User Generated Videos”
Kuo-Hao Zeng, Tseng-Hung Chen, Juan Niebles and Min Sun · 2016
Earlier work this paper cites.
“Character-level Convolutional Networks for Text Classification”
Xiang Zhang, Junbo Zhao and Yann LeCun · 2016
Earlier work this paper cites.
“MARS: A Video Benchmark for Large-Scale Person Re-Identification”
Liang Zheng et al · 2016
Earlier work this paper cites.
“Youtube-8m: A large-scale video classification benchmark”
Sami Abu-El-Haija et al · 2016
Earlier work this paper cites.
“Squad: 100,000+ questions for machine comprehension of text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
“Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding”, 2016
Gunnar. Sigurdsson et al · 2016
Earlier work this paper cites.
“YouTube-8M: A Large-Scale Video Classification Benchmark”
Sami Abu-El-Haija et al · 2016
Earlier work this paper cites.
“Unsupervised Learning from Narrated Instruction Videos”
Jean-Baptiste Alayrac et al · 2016
Earlier work this paper cites.
“BRAD 1.0: Book reviews in Arabic dataset”
Ashraf Elnagar and Omar Einea · 2016
Earlier work this paper cites.
“A Hierarchical Deep Temporal Model for Group Activity Recognition”
Moustafa Ibrahim et al · 2016
Earlier work this paper cites.
“SelQA: A New Benchmark for Selection-based Question Answering”
Tomasz Jurczyk, Michael Zhai and Jinho. Choi · 2016
Earlier work this paper cites.
“1.5 billion words Arabic Corpus”
Ibrahim El-khair · 2016
Earlier work this paper cites.
“Neural Text Generation from Structured Data with Application to the Biography Domain”
Remi Lebret, David Grangier and Michael Auli · 2016
Earlier work this paper cites.
“TGIF: A New Dataset and Benchmark on Animated GIF Description”
Yuncheng Li et al · 2016
Earlier work this paper cites.
Ryan Lowe, Nissan Pow, Iulian Serban and Joelle Pineau · 2016
Earlier work this paper cites.
“Pointer Sentinel Mixture Models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2016
Earlier work this paper cites.
“A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation” ISSN: 1063-6919
F. Perazzi et al · 2016
Earlier work this paper cites.
“SQuAD: 100,000+ Questions for Machine Comprehension of Text”
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev and Percy Liang · 2016
Earlier work this paper cites.
Anna Rohrbach et al · 2016
Earlier work this paper cites.
“Recognizing Fine-Grained and Composite Activities using Hand-Centric Features and Script Data”
Marcus Rohrbach et al · 2016
Earlier work this paper cites.
“NTU RGB+D: A Large Scale Dataset for 3D Human Activity Analysis”
Amir Shahroudy, Jun Liu, Tian-Tsong Ng and Gang Wang · 2016
Earlier work this paper cites.
“Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding”
Gunnar. Sigurdsson et al · 2016
Earlier work this paper cites.
“MovieQA: Understanding Stories in Movies through Question-Answering”
Makarand Tapaswi et al · 2016
Earlier work this paper cites.
“MSR-VTT: A Large Video Description Dataset for Bridging Video and Language” ISSN: 1063-6919
Jun Xu, Tao Mei, Ting Yao and Yong Rui · 2016
Earlier work this paper cites.
“The Value of Semantic Parse Labeling for Knowledge Base Question Answering”
Wen-tau Yih et al · 2016
Earlier work this paper cites.
“Title Generation for User Generated Videos”
Kuo-Hao Zeng, Tseng-Hung Chen, Juan Niebles and Min Sun · 2016
Earlier work this paper cites.
“Character-level Convolutional Networks for Text Classification”
Xiang Zhang, Junbo Zhao and Yann LeCun · 2016
Earlier work this paper cites.
“MARS: A Video Benchmark for Large-Scale Person Re-Identification”
Liang Zheng et al · 2016
Earlier work this paper cites.
“The LJ Speech Dataset”, 2017
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
Shreya Shankar et al · 2017
Earlier work this paper cites.
“AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline”
Hui Bu et al · 2017
Earlier work this paper cites.
“Fair prediction with disparate impact: A study of bias in recidivism prediction instruments”
Alexandra Chouldechova · 2017
Earlier work this paper cites.
“Frames: a corpus for adding memory to goal-oriented dialogue systems”
Layla El et al · 2017
Earlier work this paper cites.
“Key-Value Retrieval Networks for Task-Oriented Dialogue”
Mihail Eric and Christopher. Manning · 2017
Earlier work this paper cites.
“From Lifestyle Vlogs to Everyday Interactions”
David. Fouhey, Wei-cheng Kuo, Alexei. Efros and Jitendra Malik · 2017
Earlier work this paper cites.
“The "something something" video database for learning and evaluating visual common sense”
Raghav Goyal et al · 2017
Earlier work this paper cites.
“A Neural Representation of Sketch Drawings”
David Ha and Douglas Eck · 2017
Earlier work this paper cites.
“The THUMOS Challenge on Action Recognition for Videos "in the Wild"”
Haroon Idrees et al · 2017
Earlier work this paper cites.
“The LJ Speech Dataset”, 2017
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
“Learning a Neural Semantic Parser from User Feedback”
Srinivasan Iyer et al · 2017
Earlier work this paper cites.
“Search-based Neural Structured Learning for Sequential Question Answering”
Mohit Iyyer, Wen-tau Yih and Ming-Wei Chang · 2017
Earlier work this paper cites.
“TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension”
Mandar Joshi, Eunsol Choi, Daniel. Weld and Luke Zettlemoyer · 2017
Earlier work this paper cites.
“The Kinetics Human Action Video Dataset”
Will Kay et al · 2017
Earlier work this paper cites.
“Polish Read Speech Corpus for Speech Tools and Services”
Danijel Korzinek, Krzysztof Marasek, Lukasz Brocki and Krzysztof Wolk · 2017
Earlier work this paper cites.
“RACE: Large-scale ReAding Comprehension Dataset From Examinations”
Guokun Lai et al · 2017
Earlier work this paper cites.
“Deal or No Deal? End-to-End Learning for Negotiation Dialogues”
Mike Lewis et al · 2017
Earlier work this paper cites.
“Free linguistic and speech resources for Tibetan”
Guanyu Li et al · 2017
Earlier work this paper cites.
“Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems”
Wang Ling, Dani Yogatama, Chris Dyer and Phil Blunsom · 2017
Earlier work this paper cites.
“PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding”
Chunhui Liu et al · 2017
Earlier work this paper cites.
“Neural Belief Tracker: Data-Driven Dialogue State Tracking”
Nikola Mrksic et al · 2017
Earlier work this paper cites.
“The E2E Dataset: New Challenges For End-to-End Generation”
Jekaterina Novikova, Ondřej Dušek and Verena Rieser · 2017
Earlier work this paper cites.
“Get To The Point: Summarization with Pointer-Generator Networks”
Abigail See, Peter. Liu and Christopher. Manning · 2017
Earlier work this paper cites.
“Query-Focused Video Summarization: Dataset, Evaluation, and A Memory Network Based Approach”
Aidean Sharghi, Jacob. Laurel and Boqing Gong · 2017
Earlier work this paper cites.
“A free Kazakh speech database and a speech recognition baseline”
Ying Shi et al · 2017
Earlier work this paper cites.
“Crowdsourcing Multiple Choice Science Questions”
Johannes Welbl, Nelson. Liu and Matt Gardner · 2017
Earlier work this paper cites.
“Every Moment Counts: Dense Detailed Labeling of Actions in Complex Videos”
Serena Yeung et al · 2017
Earlier work this paper cites.
“Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning”
Victor Zhong, Caiming Xiong and Richard Socher · 2017
Earlier work this paper cites.
“Towards Automatic Learning of Procedures from Web Instructional Videos”
Luowei Zhou, Chenliang Xu and Jason. Corso · 2017
Earlier work this paper cites.
“VoxCeleb: a large-scale speaker identification dataset”, 2018
Arsha Nagrani, Joon Chung and Andrew Zisserman · 2017
Earlier work this paper cites.
“MuST-C: a Multilingual Speech Translation Corpus”
Mattia. Di et al · 2017
Earlier work this paper cites.
“The LJ Speech Dataset”, 2017
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
Shreya Shankar et al · 2017
Earlier work this paper cites.
“AISHELL-1: An Open-Source Mandarin Speech Corpus and A Speech Recognition Baseline”
Hui Bu et al · 2017
Earlier work this paper cites.
“Fair prediction with disparate impact: A study of bias in recidivism prediction instruments”
Alexandra Chouldechova · 2017
Earlier work this paper cites.
“Frames: a corpus for adding memory to goal-oriented dialogue systems”
Layla El et al · 2017
Earlier work this paper cites.
“Key-Value Retrieval Networks for Task-Oriented Dialogue”
Mihail Eric and Christopher. Manning · 2017
Earlier work this paper cites.
“From Lifestyle Vlogs to Everyday Interactions”
David. Fouhey, Wei-cheng Kuo, Alexei. Efros and Jitendra Malik · 2017
Earlier work this paper cites.
“The "something something" video database for learning and evaluating visual common sense”
Raghav Goyal et al · 2017
Earlier work this paper cites.
“A Neural Representation of Sketch Drawings”
David Ha and Douglas Eck · 2017
Earlier work this paper cites.
“The THUMOS Challenge on Action Recognition for Videos "in the Wild"”
Haroon Idrees et al · 2017
Earlier work this paper cites.
“The LJ Speech Dataset”, 2017
Keith Ito and Linda Johnson · 2017
Earlier work this paper cites.
“Learning a Neural Semantic Parser from User Feedback”
Srinivasan Iyer et al · 2017
Earlier work this paper cites.
“Search-based Neural Structured Learning for Sequential Question Answering”
Mohit Iyyer, Wen-tau Yih and Ming-Wei Chang · 2017
Earlier work this paper cites.
“TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension”
Mandar Joshi, Eunsol Choi, Daniel. Weld and Luke Zettlemoyer · 2017
Cited alongside, same era.
“The Kinetics Human Action Video Dataset”
Will Kay et al · 2017
Cited alongside, same era.
“Polish Read Speech Corpus for Speech Tools and Services”
Danijel Korzinek, Krzysztof Marasek, Lukasz Brocki and Krzysztof Wolk · 2017
Cited alongside, same era.
“RACE: Large-scale ReAding Comprehension Dataset From Examinations”
Guokun Lai et al · 2017
Cited alongside, same era.
“Deal or No Deal? End-to-End Learning for Negotiation Dialogues”
“TopiOCQA: Open-domain Conversational Question Answering with Topic Switching”
Vaibhav Adlakha et al · 2022
Later among the works it cites.
“Towards Tracing Factual Knowledge in Language Models Back to the Training Data”
Ekin Akyürek et al · 2022
Later among the works it cites.
“ArMATH: a Dataset for Solving Arabic Math Word Problems”
Reem Alghamdi, Zhenwen Liang and Xiangliang Zhang · 2022
Later among the works it cites.
“Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback”
Yuntao Bai et al · 2022
Later among the works it cites.
“Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval”
Max Bain, Arsha Nagrani, Gül Varol and Andrew Zisserman · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mike Lewis et al · 2017
Cited alongside, same era.
“Free linguistic and speech resources for Tibetan”
Guanyu Li et al · 2017
Cited alongside, same era.
“Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems”
Wang Ling, Dani Yogatama, Chris Dyer and Phil Blunsom · 2017
Cited alongside, same era.
“PKU-MMD: A Large Scale Benchmark for Continuous Multi-Modal Human Action Understanding”
Chunhui Liu et al · 2017
Cited alongside, same era.
“Neural Belief Tracker: Data-Driven Dialogue State Tracking”
Nikola Mrksic et al · 2017
Cited alongside, same era.
“The E2E Dataset: New Challenges For End-to-End Generation”
Jekaterina Novikova, Ondřej Dušek and Verena Rieser · 2017
Cited alongside, same era.
“Get To The Point: Summarization with Pointer-Generator Networks”
Abigail See, Peter. Liu and Christopher. Manning · 2017
Cited alongside, same era.
“Query-Focused Video Summarization: Dataset, Evaluation, and A Memory Network Based Approach”
Aidean Sharghi, Jacob. Laurel and Boqing Gong · 2017
Cited alongside, same era.
“Gram Vaani ASR Challenge on spontaneous telephone speech recordings in regional variations of Hindi”
Anish Bhanushali et al · 2022
Later among the works it cites.
Kaushal Bhogale et al · 2022
Later among the works it cites.
“Few-shot Adaptation Works with UnpredicTable Data”
Jun Chan et al · 2022
Later among the works it cites.
“SummScreen: A Dataset for Abstractive Screenplay Summarization”
Mingda Chen, Zewei Chu, Sam Wiseman and Kevin Gimpel · 2022
Later among the works it cites.
“KETOD: Knowledge-Enriched Task-Oriented Dialogue”
Zhiyu Chen et al · 2022
Later among the works it cites.
“HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language Generation”
Zhoujun Cheng et al · 2022
Later among the works it cites.
“SalesBot: Transitioning from Chit-Chat to Task-Oriented Dialogues”
Ssu Chiu, Maolin Li, Yen-Ting Lin and Yun-Nung Chen · 2022
Later among the works it cites.
“FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech” version: 1
Alexis Conneau et al · 2022
Later among the works it cites.
“Can Transformer be Too Compositional? Analysing Idiom Processing in Neural Machine Translation”
Verna Dankers, Christopher Lucas and Ivan Titov · 2022
Later among the works it cites.
“Earnings-22: A Practical Benchmark for Accents in the Wild”
Miguel Del et al · 2022
Later among the works it cites.
“Ego4D: Around the World in 3,000 Hours of Egocentric Video”
Kristen Grauman et al · 2022
Later among the works it cites.
“Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing”
Yu Gu et al · 2022
Later among the works it cites.
Peter Henderson et al · 2022
Later among the works it cites.
“Samrómur Children: An Icelandic Speech Corpus”
Carlos Hernandez, David Mollberg, Michal Borsky and Jon Gudnason · 2022
Later among the works it cites.
“’Tis but Thy Name: Semantic Question Answering Evaluation with 11M Names for 1M Entities”
Albert Huang · 2022
Later among the works it cites.
“Efficient Long-Text Understanding with Short-Text Models”
Maor Ivgi, Uri Shaham and Jonathan Berant · 2022
Later among the works it cites.
“IndicSUPERB: A Speech Processing Universal Performance Benchmark for Indian languages”
Tahir Javed et al · 2022
Later among the works it cites.
“ProsocialDialog: A Prosocial Backbone for Conversational Agents”
Hyunwoo Kim et al · 2022
Later among the works it cites.
“Internet-Augmented Dialogue Generation”
Mojtaba Komeili, Kurt Shuster and Jason Weston · 2022
Later among the works it cites.
“Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks”
Colin Leong et al · 2022
Later among the works it cites.
“Competition-Level Code Generation with AlphaCode”
Yujia Li et al · 2022
Later among the works it cites.
“TruthfulQA: Measuring How Models Mimic Human Falsehoods”
Stephanie Lin, Jacob Hilton and Owain Evans · 2022
Later among the works it cites.
“Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering”
Pan Lu et al · 2022
Later among the works it cites.
“ECTSum: A New Benchmark Dataset For Bullet Point Summarization of Long Earnings Call Transcripts”
Rajdeep Mukherjee et al · 2022
Later among the works it cites.
“FeTaQA: Free-form Table Question Answering” Place: Cambridge, MA Publisher: MIT Press
Linyong Nan et al · 2022
Later among the works it cites.
“A High-Quality and Large-Scale Dataset for English-Vietnamese Speech Translation”
Linh Nguyen et al · 2022
Later among the works it cites.
“SDS-200: A Swiss German Speech to Standard German Text Corpus”
Michel Plüss et al · 2022
Later among the works it cites.
“Multilingual Event Linking to Wikidata”
Adithya Pratapa, Rishubh Gupta and Teruko Mitamura · 2022
Later among the works it cites.
“Database Search Results Disambiguation for Task-Oriented Dialog Systems”
Kun Qian et al · 2022
Later among the works it cites.
“Multitask Prompted Training Enables Zero-Shot Task Generalization”
Victor Sanh et al · 2022
Later among the works it cites.
“The Norwegian Parliamentary Speech Corpus”
Per Solberg and Pablo Ortiz · 2022
Later among the works it cites.
“MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions”
Mattia Soldan et al · 2022
Later among the works it cites.
“Finnish Parliament ASR corpus - Analysis, benchmarks and statistics”
Anja Virkkunen, Aku Rouhe, Nhan Phan and Mikko Kurimo · 2022
Later among the works it cites.
“Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models”
Boxin Wang et al · 2022
Later among the works it cites.
“FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos”
Yan Wang et al · 2022
Later among the works it cites.
“CDAD: A Common Daily Action Dataset with Collected Hard Negative Samples”
Wangmeng Xiang et al · 2022
Later among the works it cites.
“Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions”
Hongwei Xue et al · 2022
Later among the works it cites.
“Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset”
Zehui Yang et al · 2022
Later among the works it cites.
“WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition”
Binbin Zhang et al · 2022
Later among the works it cites.
Jianguo Zhang et al · 2022
Later among the works it cites.
“XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence”
Ming Zhu et al · 2022
Later among the works it cites.
“Masader: Metadata Sourcing for Arabic Text and Speech Data Resources”
Zaid Alyafeai, Maraim Masoud, Mustafa Ghaleb and Maged Al-shaibani · 2022
Later among the works it cites.
“Quantifying memorization across neural language models”, 2022
Nicholas Carlini et al · 2022
Later among the works it cites.
“CLAP: Learning Audio Concepts From Natural Language Supervision”, 2022
Benjamin Elizalde, Soham Deshmukh, Mahmoud Ismail and Huaming Wang · 2022
Later among the works it cites.
“Dataset Geography: Mapping Language Data to Language Users”
Fahim Faisal, Yinkai Wang and Antonios Anastasopoulos · 2022
Later among the works it cites.
“The flores-101 evaluation benchmark for low-resource and multilingual machine translation”
Naman Goyal et al · 2022
Later among the works it cites.
“Training compute-optimal large language models”
Jordan Hoffmann et al · 2022
Later among the works it cites.
“Leakage and the reproducibility crisis in ML-based science”
Sayash Kapoor and Arvind Narayanan · 2022
Later among the works it cites.
“Quality at a glance: An audit of web-crawled multilingual datasets”
Julia Kreutzer et al · 2022
Later among the works it cites.
“The BigScience ROOTS Corpus: A 1.6TB Composite Multilingual Dataset”
Hugo Laurençon et al · 2022
Later among the works it cites.
Angelina McMillan-Major et al · 2022
Later among the works it cites.
Angelina McMillan-Major et al · 2022
Later among the works it cites.
“Video Captioning: a comparative review of where we are and which could be the route”, 2022
Daniela Moctezuma, Tania Ramírez-delReal, Guillermo Ruiz and Othón González-Chávez · 2022
Later among the works it cites.
“Training language models to follow instructions with human feedback”
Long Ouyang et al · 2022
Later among the works it cites.
“Hierarchical Text-Conditional Image Generation with CLIP Latents” arXiv: arXiv:2204.06125, 2022
Aditya Ramesh et al · 2022
Later among the works it cites.
“Red-Teaming the Stable Diffusion Safety Filter”, 2022
Javier Rando et al · 2022
Later among the works it cites.
“Make-A-Video: Text-to-Video Generation without Text-Video Data” arXiv: arXiv:2209.14792, 2022
Uriel Singer et al · 2022
Later among the works it cites.
“BigSSL: Exploring the Frontier of Large-Scale Semi-Supervised Learning for Automatic Speech Recognition”
Yu Zhang et al · 2022
Later among the works it cites.
“Survey of video object detection algorithms based on deep learning”
Lu Zheng, Tongtong Zhou, Rongqi Jiang and Yueping Peng · 2022
Later among the works it cites.
“TopiOCQA: Open-domain Conversational Question Answering with Topic Switching”
Vaibhav Adlakha et al · 2022
Later among the works it cites.
“Towards Tracing Factual Knowledge in Language Models Back to the Training Data”
Ekin Akyürek et al · 2022
Later among the works it cites.
“ArMATH: a Dataset for Solving Arabic Math Word Problems”
Reem Alghamdi, Zhenwen Liang and Xiangliang Zhang · 2022
Later among the works it cites.
“Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback”
Yuntao Bai et al · 2022
Later among the works it cites.
“Frozen in Time: A Joint Video and Image Encoder for End-to-End Retrieval”
Max Bain, Arsha Nagrani, Gül Varol and Andrew Zisserman · 2022
Later among the works it cites.
“Gram Vaani ASR Challenge on spontaneous telephone speech recordings in regional variations of Hindi”
Anish Bhanushali et al · 2022
Later among the works it cites.
Kaushal Bhogale et al · 2022
Later among the works it cites.
“Few-shot Adaptation Works with UnpredicTable Data”
Jun Chan et al · 2022
Later among the works it cites.
“SummScreen: A Dataset for Abstractive Screenplay Summarization”
Mingda Chen, Zewei Chu, Sam Wiseman and Kevin Gimpel · 2022
Later among the works it cites.
“KETOD: Knowledge-Enriched Task-Oriented Dialogue”
Zhiyu Chen et al · 2022
Later among the works it cites.
“HiTab: A Hierarchical Table Dataset for Question Answering and Natural Language Generation”
Zhoujun Cheng et al · 2022
Later among the works it cites.
“SalesBot: Transitioning from Chit-Chat to Task-Oriented Dialogues”
Ssu Chiu, Maolin Li, Yen-Ting Lin and Yun-Nung Chen · 2022
Later among the works it cites.
“FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech” version: 1
Alexis Conneau et al · 2022
Later among the works it cites.
“Can Transformer be Too Compositional? Analysing Idiom Processing in Neural Machine Translation”
Verna Dankers, Christopher Lucas and Ivan Titov · 2022
Later among the works it cites.
“Earnings-22: A Practical Benchmark for Accents in the Wild”
Miguel Del et al · 2022
Later among the works it cites.
“Ego4D: Around the World in 3,000 Hours of Egocentric Video”
Kristen Grauman et al · 2022
Later among the works it cites.
“Domain-Specific Language Model Pretraining for Biomedical Natural Language Processing”
Yu Gu et al · 2022
Later among the works it cites.
Peter Henderson et al · 2022
Later among the works it cites.
“Samrómur Children: An Icelandic Speech Corpus”
Carlos Hernandez, David Mollberg, Michal Borsky and Jon Gudnason · 2022
Later among the works it cites.
“’Tis but Thy Name: Semantic Question Answering Evaluation with 11M Names for 1M Entities”
Albert Huang · 2022
Later among the works it cites.
“Efficient Long-Text Understanding with Short-Text Models”
Maor Ivgi, Uri Shaham and Jonathan Berant · 2022
Later among the works it cites.
“IndicSUPERB: A Speech Processing Universal Performance Benchmark for Indian languages”
Tahir Javed et al · 2022
Later among the works it cites.
“ProsocialDialog: A Prosocial Backbone for Conversational Agents”
Hyunwoo Kim et al · 2022
Later among the works it cites.
“Internet-Augmented Dialogue Generation”
Mojtaba Komeili, Kurt Shuster and Jason Weston · 2022
Later among the works it cites.
“Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks”
Colin Leong et al · 2022
Later among the works it cites.
“Competition-Level Code Generation with AlphaCode”
Yujia Li et al · 2022
Later among the works it cites.
“TruthfulQA: Measuring How Models Mimic Human Falsehoods”
Stephanie Lin, Jacob Hilton and Owain Evans · 2022
Later among the works it cites.
“Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering”
Pan Lu et al · 2022
Later among the works it cites.
“ECTSum: A New Benchmark Dataset For Bullet Point Summarization of Long Earnings Call Transcripts”
Rajdeep Mukherjee et al · 2022
Later among the works it cites.
“FeTaQA: Free-form Table Question Answering” Place: Cambridge, MA Publisher: MIT Press
Linyong Nan et al · 2022
Later among the works it cites.
“A High-Quality and Large-Scale Dataset for English-Vietnamese Speech Translation”
Linh Nguyen et al · 2022
Later among the works it cites.
“SDS-200: A Swiss German Speech to Standard German Text Corpus”
Michel Plüss et al · 2022
Later among the works it cites.
“Multilingual Event Linking to Wikidata”
Adithya Pratapa, Rishubh Gupta and Teruko Mitamura · 2022
Later among the works it cites.
“Database Search Results Disambiguation for Task-Oriented Dialog Systems”
Kun Qian et al · 2022
Later among the works it cites.
“Multitask Prompted Training Enables Zero-Shot Task Generalization”
Victor Sanh et al · 2022
Later among the works it cites.
“The Norwegian Parliamentary Speech Corpus”
Per Solberg and Pablo Ortiz · 2022
Later among the works it cites.
“MAD: A Scalable Dataset for Language Grounding in Videos from Movie Audio Descriptions”
Mattia Soldan et al · 2022
Later among the works it cites.
“Finnish Parliament ASR corpus - Analysis, benchmarks and statistics”
Anja Virkkunen, Aku Rouhe, Nhan Phan and Mikko Kurimo · 2022
Later among the works it cites.
“Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models”
Boxin Wang et al · 2022
Later among the works it cites.
“FERV39k: A Large-Scale Multi-Scene Dataset for Facial Expression Recognition in Videos”
Yan Wang et al · 2022
Later among the works it cites.
“CDAD: A Common Daily Action Dataset with Collected Hard Negative Samples”
Wangmeng Xiang et al · 2022
Later among the works it cites.
“Advancing High-Resolution Video-Language Representation with Large-Scale Video Transcriptions”
Hongwei Xue et al · 2022
Later among the works it cites.
“Open Source MagicData-RAMC: A Rich Annotated Mandarin Conversational(RAMC) Speech Dataset”
Zehui Yang et al · 2022
Later among the works it cites.
“WenetSpeech: A 10000+ Hours Multi-domain Mandarin Corpus for Speech Recognition”
Binbin Zhang et al · 2022
Later among the works it cites.
Jianguo Zhang et al · 2022
Later among the works it cites.
“XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence”
Ming Zhu et al · 2022
Later among the works it cites.
“Scaling laws for generative mixed-modal language models”
Armen Aghajanyan et al · 2023
Later among the works it cites.
“Into the LAIONs Den: Investigating Hate in Multimodal Datasets”
Abeba Birhane et al · 2023
Later among the works it cites.
“Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets”, 2023
Andreas Blattmann et al · 2023
Later among the works it cites.
“The Foundation Model Transparency Index”, 2023
Rishi Bommasani et al · 2023
Later among the works it cites.
“Quantifying Memorization Across Neural Language Models”
Nicholas Carlini et al · 2023
Later among the works it cites.
“Extracting Training Data from Diffusion Models”
Nicolas Carlini et al · 2023
Later among the works it cites.
“AI Supply Chains”, 2023
Sarah Cen et al · 2023
Later among the works it cites.
“Gender Bias in Hiring: An Analysis of the Impact of Amazon’s Recruiting Algorithm”
Xinyu Chang · 2023
Later among the works it cites.
“Can Language Models be Instructed to Protect Personal Information?”, 2023
Yang Chen et al · 2023
Later among the works it cites.
“Dialect corpora from YouTube”
Steven Coats · 2023
Later among the works it cites.
“AI image training dataset found to include child sexual abuse imagery” 7:57 AM PST
Emilia David · 2023
Later among the works it cites.
“What’s In My Big Data?”
Yanai Elazar et al · 2023
Later among the works it cites.
“Structure and Content-Guided Video Synthesis with Diffusion Models”, 2023
Patrick Esser et al · 2023
Later among the works it cites.
“DataComp: In search of the next generation of multimodal datasets”
Samir Gadre et al · 2023
Later among the works it cites.
“Foundation models and fair use”
Peter Henderson et al · 2023
Later among the works it cites.
“Understanding Catastrophic Forgetting in Language Models via Implicit Inference”
Suhas Kotha, Jacob Springer and Aditi Raghunathan · 2023
Later among the works it cites.
“Harnessing large-language models to generate private synthetic text”, 2023
Alexey Kurakin et al · 2023
Later among the works it cites.
“Platypus: Quick, Cheap, and Powerful Refinement of LLMs”
Ariel. Lee, Cole. Hunter and Nataniel Ruiz · 2023
Later among the works it cites.
“Talkin”Bout AI Generation: Copyright and the Generative-AI Supply Chain”
Katherine Lee, A Cooper and James Grimmelmann · 2023
Later among the works it cites.
“Yodas: Youtube-Oriented Dataset for Audio and Speech”
Xinjian Li et al · 2023
Later among the works it cites.
“Improved Baselines with Visual Instruction Tuning”
Haotian Liu, Chunyuan Li, Yuheng Li and Yong Lee · 2023
Later among the works it cites.
“Visual Instruction Tuning”
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Later among the works it cites.
Shayne Longpre et al · 2023
Later among the works it cites.
“Discit Ergo Est: Training Data Provenance and Fair Use”
Robert Mahari and Shayne Longpre · 2023
Later among the works it cites.
“Comment to US Copyright Office on Data Provenance and Copyright”
Robert Mahari et al · 2023
Later among the works it cites.
“When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale”, 2023
Max Marion et al · 2023
Later among the works it cites.
“SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore”
Sewon Min et al · 2023
Later among the works it cites.
“Licensed to Learn: Mitigating Copyright Infringement Liability of Generative AI Systems through Contracts”
Frank Morton-Park · 2023
Later among the works it cites.
“Crosslingual Generalization through Multitask Finetuning”
Niklas Muennighoff et al · 2023
Later among the works it cites.
Guilherme Penedo et al · 2023
Later among the works it cites.
“Reproducing whisper-style training using an open-source toolkit and publicly available data”
Yifan Peng et al · 2023
Later among the works it cites.
“The Casual Conversations v2 Dataset”, 2023
Bilal Porgali et al · 2023
Later among the works it cites.
“Goodtriever: Adaptive Toxicity Mitigation with Retrieval-augmented Models”, 2023
Luiza Pozzobon, Beyza Ermis, Patrick Lewis and Sara Hooker · 2023
Later among the works it cites.
“End-to-end speech recognition: A survey”
Rohit Prabhavalkar et al · 2023
Later among the works it cites.
“Robust speech recognition via large-scale weak supervision”
Alec Radford et al · 2023
Later among the works it cites.
“Direct preference optimization: Your language model is secretly a reward model”
Rafael Rafailov et al · 2023
Later among the works it cites.
“Self-Supervised Learning for Videos: A Survey”
Madeline. Schiappa, Yogesh. Rawat and Mubarak Shah · 2023
Later among the works it cites.
“Detecting Personal Information in Training Corpora: an Analysis”
Nishant Subramani, Sasha Luccioni, Jesse Dodge and Margaret Mitchell · 2023
Later among the works it cites.
“Gemini: a family of highly capable multimodal models”
Gemini Team et al · 2023
Later among the works it cites.
“YouTube-ASL: A Large-Scale, Open-Domain American Sign Language-English Parallel Corpus”
Dave Uthus, Garrett Tanzer and Manfred Georg · 2023
Later among the works it cites.
Dawen Zhang et al · 2023
Later among the works it cites.
“Deep Learning for Video-Text Retrieval: a Review”, 2023
Cunjuan Zhu et al · 2023
Later among the works it cites.
“ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics”
Zhangir Azerbayev et al · 2023
Later among the works it cites.
“MegaWika: Millions of reports and their sources across 50 diverse languages”
Samuel Barham et al · 2023
Later among the works it cites.
“PLACES: Prompting Language Models for Social Conversation Synthesis”
Maximillian Chen et al · 2023
Later among the works it cites.
“TheoremQA: A Theorem-driven Question Answering dataset”
Wenhu Chen et al · 2023
Later among the works it cites.
“Mind2Web: Towards a Generalist Agent for the Web”
Xiang Deng et al · 2023
Later among the works it cites.
“QLoRA: Efficient Finetuning of Quantized LLMs”
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer · 2023
Later among the works it cites.
“Enhancing Chat Language Models by Scaling High-quality Instructional Conversations”
Ning Ding et al · 2023
Later among the works it cites.
“MASC: Massive Arabic Speech Corpus”
Mohammad Al-Fetyani et al · 2023
Later among the works it cites.
“PIPPA: A Partially Synthetic Conversational Dataset”
Tear Gosling, Alpin Dale and Yinhe Zheng · 2023
Later among the works it cites.
“MedAlpaca – An Open-Source Collection of Medical Conversational AI Models and Training Data”
Tianyu Han et al · 2023
Later among the works it cites.
“OpenAssistant Conversations – Democratizing Large Language Model Alignment”
Andreas Köpf et al · 2023
Later among the works it cites.
“Bactrian-X: Multilingual Replicable Instruction-Following Models with Low-Rank Adaptation”
Haonan Li et al · 2023
Later among the works it cites.
“Yodas: Youtube-Oriented Dataset for Audio and Speech”
Xinjian Li et al · 2023
Later among the works it cites.
Yunxiang Li et al · 2023
Later among the works it cites.
Hunter Lightman et al · 2023
Later among the works it cites.
“AgentBench: Evaluating LLMs as Agents”
Xiao Liu et al · 2023
Later among the works it cites.
Jakub Macina et al · 2023
Later among the works it cites.
“M2ASR-KIRGHIZ: A Free Kirghiz Speech Database and Accompanied Baselines”
Ikram Mamtimin, Wenqiang Du and Askar Hamdulla · 2023
Later among the works it cites.
“Lila: A Unified Benchmark for Mathematical Reasoning”
Swaroop Mishra et al · 2023
Later among the works it cites.
“SeaLLMs – Large Language Models for Southeast Asia”
Xuan-Phi Nguyen et al · 2023
Later among the works it cites.
“AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR”
Tobi Olatunji et al · 2023
Later among the works it cites.
“Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception”
Xiaqing Pan et al · 2023
Later among the works it cites.
“OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset”
Jeongkyun Park et al · 2023
Later among the works it cites.
“PiC: A Phrase-in-Context Dataset for Phrase Understanding and Semantic Search”
Thang. Pham, Seunghyun Yoon, Trung Bui and Anh Nguyen · 2023
Later among the works it cites.
“Snow Mountain: Dataset of Audio Recordings of The Bible in Low Resource Languages”
Kavitha Raju, Anjaly V, Ryan Lish and Joel Mathew · 2023
Later among the works it cites.
“Exploring the Effectiveness of Instruction Tuning in Biomedical Language Processing”
Omid Rohanian, Mohammadmahdi Nouriborji and David. Clifton · 2023
Later among the works it cites.
“Language Modelling with Pixels”
Phillip Rust et al · 2023
Later among the works it cites.
“The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR”
Ramon Sanabria et al · 2023
Later among the works it cites.
“ARB: Advanced Reasoning Benchmark for Large Language Models”
Tomohiro Sawada et al · 2023
Later among the works it cites.
“Probing neural language models for understanding of words of estimative probability”
Damien Sileo and Marie-Francine Moens · 2023
Later among the works it cites.
“CsFEVER and CTKFacts: Acquiring Czech data for fact verification”
Herbert Ullrich et al · 2023
Later among the works it cites.
“Generative Language Models for Paragraph-Level Question Generation”
Asahi Ushio, Fernando Alva-Manchego and Jose Camacho-Collados · 2023
Later among the works it cites.
“Self-Instruct: Aligning Language Models with Self-Generated Instructions”
Yizhong Wang et al · 2023
Later among the works it cites.
“HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM”
Zhilin Wang et al · 2023
Later among the works it cites.
“PMC-LLaMA: Towards Building Open-source Language Models for Medicine”
Chaoyi Wu et al · 2023
Later among the works it cites.
“WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents”
Shunyu Yao, Howard Chen, John Yang and Karthik Narasimhan · 2023
Later among the works it cites.
“SelFee: Iterative Self-Revising LLM Empowered by Self-Feedback Generation”, 2023
Seonghyeon Ye et al · 2023
Later among the works it cites.
“ReazonSpeech: A Free and Massive Corpus for Japanese ASR”, 2023
Yue Yin, Daijiro Mori and Seiji Fujimoto · 2023
Later among the works it cites.
“MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models”
Longhui Yu et al · 2023
Later among the works it cites.
“MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning”
Xiang Yue et al · 2023
Later among the works it cites.
“AgentTuning: Enabling Generalized Agent Abilities for LLMs”
Aohan Zeng et al · 2023
Later among the works it cites.
“Chinese Open Instruction Generalist: A Preliminary Release”
Ge Zhang et al · 2023
Later among the works it cites.
“Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence”
Jiaxing Zhang et al · 2023
Later among the works it cites.
“WildChat: 1M ChatGPT Interaction Logs in the Wild”
Wenting Zhao et al · 2023
Later among the works it cites.
“Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”
Lianmin Zheng et al · 2023
Later among the works it cites.
“DocPrompting: Generating Code by Retrieving the Docs”
Shuyan Zhou et al · 2023
Later among the works it cites.
“COBRA Frames: Contextual Reasoning about Effects and Harms of Offensive Statements”
Xuhui Zhou et al · 2023
Later among the works it cites.
“Scaling laws for generative mixed-modal language models”
Armen Aghajanyan et al · 2023
Later among the works it cites.
“Into the LAIONs Den: Investigating Hate in Multimodal Datasets”
Abeba Birhane et al · 2023
Later among the works it cites.
“Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets”, 2023
Andreas Blattmann et al · 2023
Later among the works it cites.
“The Foundation Model Transparency Index”, 2023
Rishi Bommasani et al · 2023
Later among the works it cites.
“Quantifying Memorization Across Neural Language Models”
Nicholas Carlini et al · 2023
Later among the works it cites.
“Extracting Training Data from Diffusion Models”
Nicolas Carlini et al · 2023
Later among the works it cites.
“AI Supply Chains”, 2023
Sarah Cen et al · 2023
Later among the works it cites.
“Gender Bias in Hiring: An Analysis of the Impact of Amazon’s Recruiting Algorithm”
Xinyu Chang · 2023
Later among the works it cites.
“Can Language Models be Instructed to Protect Personal Information?”, 2023
Yang Chen et al · 2023
Later among the works it cites.
“Dialect corpora from YouTube”
Steven Coats · 2023
Later among the works it cites.
“AI image training dataset found to include child sexual abuse imagery” 7:57 AM PST
Emilia David · 2023
Later among the works it cites.
“What’s In My Big Data?”
Yanai Elazar et al · 2023
Later among the works it cites.
“Structure and Content-Guided Video Synthesis with Diffusion Models”, 2023
Patrick Esser et al · 2023
Later among the works it cites.
“DataComp: In search of the next generation of multimodal datasets”
Samir Gadre et al · 2023
Later among the works it cites.
“Foundation models and fair use”
Peter Henderson et al · 2023
Later among the works it cites.
“Understanding Catastrophic Forgetting in Language Models via Implicit Inference”
Suhas Kotha, Jacob Springer and Aditi Raghunathan · 2023
Later among the works it cites.
“Harnessing large-language models to generate private synthetic text”, 2023
Alexey Kurakin et al · 2023
Later among the works it cites.
“Platypus: Quick, Cheap, and Powerful Refinement of LLMs”
Ariel. Lee, Cole. Hunter and Nataniel Ruiz · 2023
Later among the works it cites.
“Talkin”Bout AI Generation: Copyright and the Generative-AI Supply Chain”
Katherine Lee, A Cooper and James Grimmelmann · 2023
Later among the works it cites.
“Yodas: Youtube-Oriented Dataset for Audio and Speech”
Xinjian Li et al · 2023
Later among the works it cites.
“Improved Baselines with Visual Instruction Tuning”
Haotian Liu, Chunyuan Li, Yuheng Li and Yong Lee · 2023
Later among the works it cites.
“Visual Instruction Tuning”
Haotian Liu, Chunyuan Li, Qingyang Wu and Yong Lee · 2023
Later among the works it cites.
Shayne Longpre et al · 2023
Later among the works it cites.
“Discit Ergo Est: Training Data Provenance and Fair Use”
Robert Mahari and Shayne Longpre · 2023
Later among the works it cites.
“Comment to US Copyright Office on Data Provenance and Copyright”
Robert Mahari et al · 2023
Later among the works it cites.
“When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale”, 2023
Max Marion et al · 2023
Later among the works it cites.
“SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore”
Sewon Min et al · 2023
Later among the works it cites.
“Licensed to Learn: Mitigating Copyright Infringement Liability of Generative AI Systems through Contracts”
Frank Morton-Park · 2023
Later among the works it cites.
“Crosslingual Generalization through Multitask Finetuning”
Niklas Muennighoff et al · 2023
Later among the works it cites.
Guilherme Penedo et al · 2023
Later among the works it cites.
“Reproducing whisper-style training using an open-source toolkit and publicly available data”
Yifan Peng et al · 2023
Later among the works it cites.
“The Casual Conversations v2 Dataset”, 2023
Bilal Porgali et al · 2023
Later among the works it cites.
“Goodtriever: Adaptive Toxicity Mitigation with Retrieval-augmented Models”, 2023
Luiza Pozzobon, Beyza Ermis, Patrick Lewis and Sara Hooker · 2023
Later among the works it cites.
“End-to-end speech recognition: A survey”
Rohit Prabhavalkar et al · 2023
Later among the works it cites.
“Robust speech recognition via large-scale weak supervision”
Alec Radford et al · 2023
Later among the works it cites.
“Direct preference optimization: Your language model is secretly a reward model”
Rafael Rafailov et al · 2023
Later among the works it cites.
“Self-Supervised Learning for Videos: A Survey”
Madeline. Schiappa, Yogesh. Rawat and Mubarak Shah · 2023
Later among the works it cites.
“Detecting Personal Information in Training Corpora: an Analysis”
Nishant Subramani, Sasha Luccioni, Jesse Dodge and Margaret Mitchell · 2023
Later among the works it cites.
“Gemini: a family of highly capable multimodal models”
Gemini Team et al · 2023
Later among the works it cites.
“YouTube-ASL: A Large-Scale, Open-Domain American Sign Language-English Parallel Corpus”
Dave Uthus, Garrett Tanzer and Manfred Georg · 2023
Later among the works it cites.
Dawen Zhang et al · 2023
Later among the works it cites.
“Deep Learning for Video-Text Retrieval: a Review”, 2023
Cunjuan Zhu et al · 2023
Later among the works it cites.
“ProofNet: Autoformalizing and Formally Proving Undergraduate-Level Mathematics”
Zhangir Azerbayev et al · 2023
Later among the works it cites.
“MegaWika: Millions of reports and their sources across 50 diverse languages”
Samuel Barham et al · 2023
Later among the works it cites.
“PLACES: Prompting Language Models for Social Conversation Synthesis”
Maximillian Chen et al · 2023
Later among the works it cites.
“TheoremQA: A Theorem-driven Question Answering dataset”
Wenhu Chen et al · 2023
Later among the works it cites.
“Mind2Web: Towards a Generalist Agent for the Web”
Xiang Deng et al · 2023
Later among the works it cites.
“QLoRA: Efficient Finetuning of Quantized LLMs”
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman and Luke Zettlemoyer · 2023
Later among the works it cites.
“Enhancing Chat Language Models by Scaling High-quality Instructional Conversations”
Ning Ding et al · 2023
Later among the works it cites.
“MASC: Massive Arabic Speech Corpus”
Mohammad Al-Fetyani et al · 2023
Later among the works it cites.
“PIPPA: A Partially Synthetic Conversational Dataset”
Tear Gosling, Alpin Dale and Yinhe Zheng · 2023
Later among the works it cites.
“MedAlpaca – An Open-Source Collection of Medical Conversational AI Models and Training Data”
Tianyu Han et al · 2023
Later among the works it cites.
“OpenAssistant Conversations – Democratizing Large Language Model Alignment”
Andreas Köpf et al · 2023
Later among the works it cites.
“Bactrian-X: Multilingual Replicable Instruction-Following Models with Low-Rank Adaptation”
Haonan Li et al · 2023
Later among the works it cites.
“Yodas: Youtube-Oriented Dataset for Audio and Speech”
Xinjian Li et al · 2023
Later among the works it cites.
Yunxiang Li et al · 2023
Later among the works it cites.
Hunter Lightman et al · 2023
Later among the works it cites.
“AgentBench: Evaluating LLMs as Agents”
Xiao Liu et al · 2023
Later among the works it cites.
Jakub Macina et al · 2023
Later among the works it cites.
“M2ASR-KIRGHIZ: A Free Kirghiz Speech Database and Accompanied Baselines”
Ikram Mamtimin, Wenqiang Du and Askar Hamdulla · 2023
Later among the works it cites.
“Lila: A Unified Benchmark for Mathematical Reasoning”
Swaroop Mishra et al · 2023
Later among the works it cites.
“SeaLLMs – Large Language Models for Southeast Asia”
Xuan-Phi Nguyen et al · 2023
Later among the works it cites.
“AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR”
Tobi Olatunji et al · 2023
Later among the works it cites.
“Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception”
Xiaqing Pan et al · 2023
Later among the works it cites.
“OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset”
Jeongkyun Park et al · 2023
Later among the works it cites.
“PiC: A Phrase-in-Context Dataset for Phrase Understanding and Semantic Search”
Thang. Pham, Seunghyun Yoon, Trung Bui and Anh Nguyen · 2023
Later among the works it cites.
“Snow Mountain: Dataset of Audio Recordings of The Bible in Low Resource Languages”
Kavitha Raju, Anjaly V, Ryan Lish and Joel Mathew · 2023
Later among the works it cites.
“Exploring the Effectiveness of Instruction Tuning in Biomedical Language Processing”
Omid Rohanian, Mohammadmahdi Nouriborji and David. Clifton · 2023
Later among the works it cites.
“Language Modelling with Pixels”
Phillip Rust et al · 2023
Later among the works it cites.
“The Edinburgh International Accents of English Corpus: Towards the Democratization of English ASR”
Ramon Sanabria et al · 2023
Later among the works it cites.
“ARB: Advanced Reasoning Benchmark for Large Language Models”
Tomohiro Sawada et al · 2023
Later among the works it cites.
“Probing neural language models for understanding of words of estimative probability”
Damien Sileo and Marie-Francine Moens · 2023
Later among the works it cites.
“CsFEVER and CTKFacts: Acquiring Czech data for fact verification”
Herbert Ullrich et al · 2023
Later among the works it cites.
“Generative Language Models for Paragraph-Level Question Generation”
Asahi Ushio, Fernando Alva-Manchego and Jose Camacho-Collados · 2023
Later among the works it cites.
“Self-Instruct: Aligning Language Models with Self-Generated Instructions”
Yizhong Wang et al · 2023
Later among the works it cites.
“HelpSteer: Multi-attribute Helpfulness Dataset for SteerLM”
Zhilin Wang et al · 2023
Later among the works it cites.
“PMC-LLaMA: Towards Building Open-source Language Models for Medicine”
Chaoyi Wu et al · 2023
Later among the works it cites.
“WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents”
Shunyu Yao, Howard Chen, John Yang and Karthik Narasimhan · 2023
Later among the works it cites.
“SelFee: Iterative Self-Revising LLM Empowered by Self-Feedback Generation”, 2023
Seonghyeon Ye et al · 2023
Later among the works it cites.
“ReazonSpeech: A Free and Massive Corpus for Japanese ASR”, 2023
Yue Yin, Daijiro Mori and Seiji Fujimoto · 2023
Later among the works it cites.
“MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models”
Longhui Yu et al · 2023
Later among the works it cites.
“MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning”
Xiang Yue et al · 2023
Later among the works it cites.
“AgentTuning: Enabling Generalized Agent Abilities for LLMs”
Aohan Zeng et al · 2023
Later among the works it cites.
“Chinese Open Instruction Generalist: A Preliminary Release”
Ge Zhang et al · 2023
Later among the works it cites.
“Fengshenbang 1.0: Being the Foundation of Chinese Cognitive Intelligence”
Jiaxing Zhang et al · 2023
Later among the works it cites.
“WildChat: 1M ChatGPT Interaction Logs in the Wild”
Wenting Zhao et al · 2023
Later among the works it cites.
“Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena”
Lianmin Zheng et al · 2023
Later among the works it cites.
“DocPrompting: Generating Code by Retrieving the Docs”
Shuyan Zhou et al · 2023
Later among the works it cites.
“COBRA Frames: Contextual Reasoning about Effects and Harms of Offensive Statements”
Xuhui Zhou et al · 2023
Later among the works it cites.
“The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm”, 2024
Aakanksha et al · 2024
Closest in time.
“IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models”, 2024
David Adelani et al · 2024
Closest in time.
“A Survey on Data Selection for Language Models”
Alon Albalak et al · 2024
Closest in time.
“Video generation models as world simulators”, 2024
Tim Brooks et al · 2024
Closest in time.
“Nvidia Sued for Scraping YouTube After 404 Media Investigation”
Samantha Cole · 2024
Closest in time.
“NVLM: Open Frontier-Class Multimodal LLMs”
Wenliang Dai et al · 2024
Closest in time.
“Datacomp: In search of the next generation of multimodal datasets”
Samir Gadre et al · 2024
Closest in time.
“Acceptable Use Policies for Foundation Models”, 2024
Kevin Klyman · 2024
Closest in time.
“Best practices and lessons learned on synthetic data”, 2024
Ruibo Liu et al · 2024
Closest in time.
“Datasets for large language models: A comprehensive survey”
Yang Liu et al · 2024
Closest in time.
“The responsible foundation model development cheatsheet: A review of tools & resources”
Shayne Longpre et al · 2024
Closest in time.
“A Large-Scale Audit of Dataset Licensing and Attribution in AI”
Shayne Longpre et al · 2024
Closest in time.
“Consent in Crisis: The Rapid Decline of the AI Data Commons”
Shayne Longpre et al · 2024
Closest in time.
“Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?”
Shayne Longpre et al · 2024
Closest in time.
“SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages”
Holy Lovenia et al · 2024
Closest in time.
“What was Sora trained on? Creatives demand answers.” [Accessed 28-09-2024], https://mashable.com/article/openai-sora-ai-video-generator-training-data , 2024
Cecily Mauran · 2024
Closest in time.
“Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers”
Rajiv Movva et al · 2024
Closest in time.
“Hello GPT-4o: We’re announcing GPT-4o, our new flagship model that can reason across audio, vision, and text in real time.”, 2024
OpenAI · 2024
Closest in time.
“Data, Data Everywhere: A Guide for Pretraining Dataset Construction”
Jupinder Parmar et al · 2024
Closest in time.
“Scaling speech technology to 1,000+ languages”
Vineel Pratap et al · 2024
Closest in time.
“Anatomy of Industrial Scale Multilingual ASR”
Francis Ramirez et al · 2024
Closest in time.
“INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge”, 2024
Angelika Romanou et al · 2024
Closest in time.
“Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning”, 2024
Shivalika Singh et al · 2024
Closest in time.
“OpenAI Sued Over Using YouTube Videos Without Creators’ Consent”
Sam Skolnik · 2024
Closest in time.
“Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research”
Luca Soldaini et al · 2024
Closest in time.
“Aya model: An instruction finetuned open-access multilingual language model”
Ahmet Üstün et al · 2024
Closest in time.
“Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models”
Wenhao Wang and Yi Yang · 2024
Closest in time.
Xinyu Yang, Weixin Liang and James Zou · 2024
Closest in time.
“Open-Sora: Democratizing Efficient Video Production for All”, 2024
Zangwei Zheng et al · 2024
Closest in time.
“ArabicaQA: A Comprehensive Dataset for Arabic Question Answering”
Abdelrahman Abdallah et al · 2024
Closest in time.
“101 Billion Arabic Words Dataset”
Manel Aloui et al · 2024
Closest in time.
“CIDAR: Culturally Relevant Instruction Dataset For Arabic”
Zaid Alyafeai et al · 2024
Closest in time.
“COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning” version: 1
Yuelin Bai et al · 2024
Closest in time.
“LongAlign: A Recipe for Long Context Alignment of Large Language Models”
Yushi Bai et al · 2024
Closest in time.
“ShareGPT4Video: Improving Video Understanding and Generation with Better Captions”
Lin Chen et al · 2024
Closest in time.
“GemmAr: Enhancing LLMs Through Arabic Instruction-Tuning”
Hasna Chouikhi et al · 2024
Closest in time.
“Airavata: Introducing Hindi Instruction-tuned LLM”
Jay Gala et al · 2024
Closest in time.
“Prometheus: Inducing Fine-grained Evaluation Capability in Language Models”
Seungone Kim et al · 2024
Closest in time.
“MVBench: A Comprehensive Multi-modal Video Understanding Benchmark”
Kunchang Li et al · 2024
Closest in time.
Wei Liu et al · 2024
Closest in time.
“Aria Everyday Activities Dataset”
Zhaoyang Lv et al · 2024
Closest in time.
“ExpertQA: Expert-Curated Questions and Attributed Answers”
Chaitanya Malaviya et al · 2024
Closest in time.
“Orca-Math: Unlocking the potential of SLMs in Grade School Math”
Arindam Mitra, Hamed Khanpour, Corby Rosset and Ahmed Awadallah · 2024
Closest in time.
“OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation”
Kepan Nan et al · 2024
Closest in time.
“BiMediX: Bilingual Medical Mixture of Experts LLM”
Sara Pieri et al · 2024
Closest in time.
“Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning”
Shivalika Singh et al · 2024
Closest in time.
“Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models”
Haoran Sun et al · 2024
Closest in time.
“OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset”
Shubham Toshniwal et al · 2024
Closest in time.
“VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models”
Wenhao Wang and Yi Yang · 2024
Closest in time.
“SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models”
Xiaoxuan Wang et al · 2024
Closest in time.
“KIWI: A Dataset of Knowledge-Intensive Writing Instructions for Answering Research Questions”
Fangyuan Xu et al · 2024
Closest in time.
“Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing”
Zhangchen Xu et al · 2024
Closest in time.
“AlpaCare:Instruction-tuned Large Language Models for Medical Application”
Xinlu Zhang et al · 2024
Closest in time.
“LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset”
Lianmin Zheng et al · 2024
Closest in time.
“The Multilingual Alignment Prism: Aligning Global and Local Preferences to Reduce Harm”, 2024
Aakanksha et al · 2024
Closest in time.
“IrokoBench: A New Benchmark for African Languages in the Age of Large Language Models”, 2024
David Adelani et al · 2024
Closest in time.
“A Survey on Data Selection for Language Models”
Alon Albalak et al · 2024
Closest in time.
“Video generation models as world simulators”, 2024
Tim Brooks et al · 2024
Closest in time.
“Nvidia Sued for Scraping YouTube After 404 Media Investigation”
Samantha Cole · 2024
Closest in time.
“NVLM: Open Frontier-Class Multimodal LLMs”
Wenliang Dai et al · 2024
Closest in time.
“Datacomp: In search of the next generation of multimodal datasets”
Samir Gadre et al · 2024
Closest in time.
“Acceptable Use Policies for Foundation Models”, 2024
Kevin Klyman · 2024
Closest in time.
“Best practices and lessons learned on synthetic data”, 2024
Ruibo Liu et al · 2024
Closest in time.
“Datasets for large language models: A comprehensive survey”
Yang Liu et al · 2024
Closest in time.
“The responsible foundation model development cheatsheet: A review of tools & resources”
Shayne Longpre et al · 2024
Closest in time.
“A Large-Scale Audit of Dataset Licensing and Attribution in AI”
Shayne Longpre et al · 2024
Closest in time.
“Consent in Crisis: The Rapid Decline of the AI Data Commons”
Shayne Longpre et al · 2024
Closest in time.
“Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?”
Shayne Longpre et al · 2024
Closest in time.
“SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages”
Holy Lovenia et al · 2024
Closest in time.
“What was Sora trained on? Creatives demand answers.” [Accessed 28-09-2024], https://mashable.com/article/openai-sora-ai-video-generator-training-data , 2024
Cecily Mauran · 2024
Closest in time.
“Topics, Authors, and Institutions in Large Language Model Research: Trends from 17K arXiv Papers”
Rajiv Movva et al · 2024
Closest in time.
“Hello GPT-4o: We’re announcing GPT-4o, our new flagship model that can reason across audio, vision, and text in real time.”, 2024
OpenAI · 2024
Closest in time.
“Data, Data Everywhere: A Guide for Pretraining Dataset Construction”
Jupinder Parmar et al · 2024
Closest in time.
“Scaling speech technology to 1,000+ languages”
Vineel Pratap et al · 2024
Closest in time.
“Anatomy of Industrial Scale Multilingual ASR”
Francis Ramirez et al · 2024
Closest in time.
“INCLUDE: Evaluating Multilingual Language Understanding with Regional Knowledge”, 2024
Angelika Romanou et al · 2024
Closest in time.
“Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning”, 2024
Shivalika Singh et al · 2024
Closest in time.
“OpenAI Sued Over Using YouTube Videos Without Creators’ Consent”
Sam Skolnik · 2024
Closest in time.
“Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research”
Luca Soldaini et al · 2024
Closest in time.
“Aya model: An instruction finetuned open-access multilingual language model”
Ahmet Üstün et al · 2024
Closest in time.
“Vidprom: A million-scale real prompt-gallery dataset for text-to-video diffusion models”
Wenhao Wang and Yi Yang · 2024
Closest in time.
Xinyu Yang, Weixin Liang and James Zou · 2024
Closest in time.
“Open-Sora: Democratizing Efficient Video Production for All”, 2024
Zangwei Zheng et al · 2024
Closest in time.
“ArabicaQA: A Comprehensive Dataset for Arabic Question Answering”
Abdelrahman Abdallah et al · 2024
Closest in time.
“101 Billion Arabic Words Dataset”
Manel Aloui et al · 2024
Closest in time.
“CIDAR: Culturally Relevant Instruction Dataset For Arabic”
Zaid Alyafeai et al · 2024
Closest in time.
“COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning” version: 1
Yuelin Bai et al · 2024
Closest in time.
“LongAlign: A Recipe for Long Context Alignment of Large Language Models”
Yushi Bai et al · 2024
Closest in time.
“ShareGPT4Video: Improving Video Understanding and Generation with Better Captions”
Lin Chen et al · 2024
Closest in time.
“GemmAr: Enhancing LLMs Through Arabic Instruction-Tuning”
Hasna Chouikhi et al · 2024
Closest in time.
“Airavata: Introducing Hindi Instruction-tuned LLM”
Jay Gala et al · 2024
Closest in time.
“Prometheus: Inducing Fine-grained Evaluation Capability in Language Models”
Seungone Kim et al · 2024
Closest in time.
“MVBench: A Comprehensive Multi-modal Video Understanding Benchmark”
Kunchang Li et al · 2024
Closest in time.
Wei Liu et al · 2024
Closest in time.
“Aria Everyday Activities Dataset”
Zhaoyang Lv et al · 2024
Closest in time.
“ExpertQA: Expert-Curated Questions and Attributed Answers”
Chaitanya Malaviya et al · 2024
Closest in time.
“Orca-Math: Unlocking the potential of SLMs in Grade School Math”
Arindam Mitra, Hamed Khanpour, Corby Rosset and Ahmed Awadallah · 2024
Closest in time.
“OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation”
Kepan Nan et al · 2024
Closest in time.
“BiMediX: Bilingual Medical Mixture of Experts LLM”
Sara Pieri et al · 2024
Closest in time.
“Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning”
Shivalika Singh et al · 2024
Closest in time.
“Conifer: Improving Complex Constrained Instruction-Following Ability of Large Language Models”
Haoran Sun et al · 2024
Closest in time.
“OpenMathInstruct-1: A 1.8 Million Math Instruction Tuning Dataset”
Shubham Toshniwal et al · 2024
Closest in time.
“VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion Models”
Wenhao Wang and Yi Yang · 2024
Closest in time.
“SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models”
Xiaoxuan Wang et al · 2024
Closest in time.
“KIWI: A Dataset of Knowledge-Intensive Writing Instructions for Answering Research Questions”
Fangyuan Xu et al · 2024
Closest in time.
“Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing”
Zhangchen Xu et al · 2024
Closest in time.
“AlpaCare:Instruction-tuned Large Language Models for Medical Application”
Xinlu Zhang et al · 2024
Closest in time.
“LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset”
Lianmin Zheng et al · 2024
Closest in time.