Fetching the paper…
Reading the bibliography…
We present Amazon Nova, a new generation of state-of-the-art foundation models that deliver frontier intelligence and industry-leading price performance.
Computational complexity: a modern approach
S. Arora and B. Barak · 2009
Earlier work this paper cites.
Auto-encoding variational Bayes
D. P. Kingma · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
CIDEr: Consensus-based Image Description Evaluation
R. Vedantam, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Proximal policy optimization algorithms, 2017
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Think you have solved question answering? try ARC, the AI2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
D. Dua, Y. Wang, P. Dasigi, G. Stanovsky, S. Singh, and M. Gardner · 2019
Earlier work this paper cites.
Two new evaluation datasets for low-resource machine translation: Nepali-english and sinhala-english
F. Guzmán, P.-J. Chen, M. Ott, J. Pino, G. Lample, P. Koehn, V. Chaudhary, and M. Ranzato · 2019
Earlier work this paper cites.
Towards VQA models that can read
A. Singh, V. Natarajan, M. Shah, Y. Jiang, X. Chen, D. Batra, D. Parikh, and M. Rohrbach · 2019
Earlier work this paper cites.
VATEX: A large-scale, high-quality multilingual dataset for video-and-language research
X. Wang, J. Wu, J. Chen, L. Li, Y.-F. Wang, and W. Y. Wang · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
RealToxicityPrompts: Evaluating neural toxic degeneration in language models
S. Gehman, S. Gururangan, M. Sap, Y. Choi, and N. A. Smith · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
BOLD: Dataset and metrics for measuring biases in open-ended language generation
J. Dhamala, T. Sun, V. Kumar, S. Krishna, Y. Pruksachatkun, K.-W. Chang, and R. Gupta · 2021
Earlier work this paper cites.
The FLORES-101 evaluation benchmark for low-resource and multilingual machine translation
N. Goyal, C. Gao, V. Chaudhary, P.-J. Chen, G. Wenzek, D. Ju, S. Krishnan, M. Ranzato, F. Guzmán, and A. Fan · 2021
Earlier work this paper cites.
DocVQA: A dataset for VQA on document images
M. Mathew, D. Karatzas, and C. Jawahar · 2021
Earlier work this paper cites.
ChartQA: A benchmark for question answering about charts with visual and logical reasoning
A. Masry, D. X. Long, J. Q. Tan, S. Joty, and E. Hoque · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, J. Schulman, J. Hilton, F. Kelton, L. Miller, M. Simens, A. Askell, P. Welinder, P. F. Christiano, J. Leike, and R. Lowe · 2022
Earlier work this paper cites.
COMET-22: Unbabel-IST 2022 submission for the metrics shared task
R. Rei, J. G. C. de Souza, D. Alves, C. Zerva, A. C. Farinha, T. Glushkova, A. Lavie, L. Coheur, and A. F. T. Martins · 2022
Earlier work this paper cites.
Photorealistic text-to-image diffusion models with deep language understanding
C. Saharia, W. Chan, S. Saxena, L. Li, J. Whang, E. L. Denton, K. Ghasemipour, R. Gontijo Lopes, B. Karagol Ayan, T. Salimans, et al · 2022
Earlier work this paper cites.
Challenging BIG-Bench tasks and whether chain-of-thought can solve them
M. Suzgun, N. Scales, N. Schärli, S. Gehrmann, Y. Tay, H. W. Chung, A. Chowdhery, Q. V. Le, E. H. Chi, D. Zhou, , and J. Wei · 2022
Earlier work this paper cites.
No language left behind: Scaling human-centered machine translation
N. Team, M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, A. Sun, S. Wang, G. Wenzek, A. Youngblood, B. Akula, L. Barrault, G. M. Gonzalez, P. Hansanti, J. Hoffman, S. Jarrett, K. R. Sadagopan, D. Rowe, S. Spruit, C. Tran, P. Andrews, N. F. Ayan, S. Bhosale, S. Edunov, A. Fan, C. Gao, V. Goswami, F. Guzmán, P. Koehn, A. Mourachko, C. Ropers, S. Saleem, H. Schwenk, and J. Wang · 2022
Earlier work this paper cites.
SQuALITY: Building a long-document summarization dataset the hard way
A. Wang, R. Y. Pang, A. Chen, J. Phang, and S. R. Bowman · 2022
Earlier work this paper cites.
Scaling autoregressive models for content-rich text-to-image generation
J. Yu, Y. Xu, J. Y. Koh, T. Luong, G. Baid, Z. Wang, V. Vasudevan, A. Ku, Y. Yang, B. K. Ayan, et al · 2022
Earlier work this paper cites.
The Claude 3 model family: Opus, Sonnet, Haiku
Anthropic · 2023
Earlier work this paper cites.
Improving image generation with better captions
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y. Guo, et al · 2023
Earlier work this paper cites.
DALL-eval: Probing the reasoning skills and social biases of text-to-image generation models
J. Cho, A. Zala, and M. Bansal · 2023
Cited alongside, same era.
Mind2Web: Towards a generalist agent for the web
X. Deng, Y. Gu, B. Zheng, S. Chen, S. Stevens, B. Wang, H. Sun, and Y. Su · 2023
Cited alongside, same era.
TIFA: Accurate and interpretable text-to-image faithfulness evaluation with question answering
Y. Hu, B. Liu, J. Kasai, Y. Wang, M. Ostendorf, R. Krishna, and N. A. Smith · 2023
Cited alongside, same era.
LLMTest NeedleInAHaystack, 2023
G. Kamradt · 2023
Cited alongside, same era.
Chameleon: Plug-and-play compositional reasoning with large language models
P. Lu, B. Peng, H. Cheng, M. Galley, K.-W. Chang, Y. N. Wu, S.-C. Zhu, and J. Gao · 2023
Cited alongside, same era.
EgoSchema: A diagnostic benchmark for very long-form video language understanding
Claude 3.5 Haiku and upgraded Claude 3.5 Sonnet, 2024
Anthropic AI Team · 2024
Later among the works it cites.
Flux models
Black Forest Labs · 2024
Later among the works it cites.
Amazon and Meta join the Frontier Model Forum to promote AI safety
Frontier Model Forum · 2024
Later among the works it cites.
Hiroshima process international code of conduct for organizations developing advanced AI systems
G7 Hiroshima Summit · 2024
Later among the works it cites.
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context, 2024
Gemini Team · 2024
Later among the works it cites.
Gemini Flash
Google Deepmind · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
K. Mangalam, R. Akshulakov, and J. Malik · 2023
Cited alongside, same era.
Gorilla: Large language model connected with massive APIs, 2023
S. G. Patil, T. Zhang, X. Wang, and J. E. Gonzalez · 2023
Cited alongside, same era.
Scalable diffusion models with transformers
W. Peebles and S. Xie · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2023
Cited alongside, same era.
GPQA: A graduate-level google-proof Q&A benchmark, 2023
D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman · 2023
Cited alongside, same era.
Toolformer: Language models can teach themselves to use tools
T. Schick, J. Dwivedi-Yu, R. Dessi, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom · 2023
Cited alongside, same era.
ZeroSCROLLS: A zero-shot benchmark for long text understanding
U. Shaham, M. Ivgi, A. Efrat, J. Berant, and O. Levy · 2023
Cited alongside, same era.
GPT-4o: The cutting-edge advancement in multimodal LLM
R. Islam and O. M. Moushi · 2024
Later among the works it cites.
VisualWebBench: How far have multimodal llms evolved in web page understanding and grounding?, 2024
J. Liu, Y. Song, B. Y. Lin, W. Lam, G. Neubig, Y. Li, and X. Yue · 2024
Later among the works it cites.
The Llama 3 herd of models, 2024
Llama Team, AI Meta · 2024
Later among the works it cites.
URL https://lumalabs.ai/dream-machine
Luma Labs, 2024 · 2024
Later among the works it cites.
Quantifying variance in evaluation benchmarks, 2024
L. Madaan, A. K. Singh, R. Schaeffer, A. Poulton, S. Koyejo, P. Stenetorp, S. Narang, and D. Hupkes · 2024
Later among the works it cites.
FLIRT: Feedback loop in-context red teaming
N. Mehrabi, P. Goyal, C. Dupuy, Q. Hu, S. Ghosh, R. Zemel, K.-W. Chang, A. Galstyan, and R. Gupta · 2024
Later among the works it cites.
Llama 3.2 Github model card vision
Meta · 2024
Later among the works it cites.
GPT 4o mini
OpenAI · 2024
Later among the works it cites.
Hello GPT 4o
OpenAI · 2024
Later among the works it cites.
URL https://runwayml.com/research/introducing-gen-3-alpha
Runway Research, 2024 · 2024
Later among the works it cites.
T2V-CompBench: A comprehensive benchmark for compositional text-to-video generation
K. Sun, K. Huang, X. Liu, Y. Wu, Z. Xu, Z. Li, and X. Liu · 2024
Later among the works it cites.
LVBench: An extreme long video understanding benchmark
W. Wang, Z. He, W. Hong, Y. Cheng, X. Zhang, J. Qi, X. Gu, S. Huang, B. Xu, Y. Dong, et al · 2024
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. H. Chi, Q. V. Le, and D. Zhou · 2024
Later among the works it cites.
ImageReward: Learning and evaluating human preferences for text-to-image generation
J. Xu, X. Liu, Y. Wu, Y. Tong, Q. Li, M. Ding, J. Tang, and Y. Dong · 2024
Later among the works it cites.
Berkeley function calling leaderboard
F. Yan, H. Mao, C. C.-J. Ji, T. Zhang, S. G. Patil, I. Stoica, and J. E. Gonzalez · 2024
Later among the works it cites.
Crag – comprehensive rag benchmark
X. Yang, K. Sun, H. Xin, Y. Sun, N. Bhalla, X. Chen, S. Choudhary, R. D. Gui, Z. W. Jiang, Z. Jiang, L. Kong, B. Moran, J. Wang, Y. E. Xu, A. Yan, C. Yang, E. Yuan, H. Zha, N. Tang, L. Chen, N. Scheffer, Y. Liu, N. Shah, R. Wanga, A. Kumar, W. tau Yih, and X. L. Dong · 2024
Later among the works it cites.
MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
X. Yue, Y. Ni, K. Zhang, T. Zheng, R. Liu, G. Zhang, S. Stevens, D. Jiang, W. Ren, Y. Sun, C. Wei, B. Yu, R. Yuan, R. Sun, M. Yin, B. Zheng, Z. Yang, Y. Liu, W. Huang, H. Sun, Y. Su, and W. Chen · 2024
Later among the works it cites.
Law of the weakest link: Cross capabilities of large language models
M. Zhong, A. Zhang, X. Wang, R. Hou, W. Xiong, C. Zhu, Z. Chen, L. Tan, C. Bi, M. Lewis, S. Popuri, S. Narang, M. Kambadur, D. Mahajan, S. Edunov, J. Han, and L. van der Maaten · 2024
Later among the works it cites.
MM-SafetyBench: A benchmark for safety evaluation of multimodal large language models
X. Liu, Y. Zhu, J. Gu, Y. Lan, C. Yang, and Y. Qiao · 2025
Closest in time.