Fetching the paper…
Reading the bibliography…
Developers often have questions about semantic aspects of code they are working on, e.g., "Is there a class whose parent classes declare a conflicting attribute?".
Text chunking using transformation-based learning
Ramshaw, L. A. and Marcus, M · 1995
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Robertson, S., Zaragoza, H., et al · 2009
Earlier work this paper cites.
QL: object-oriented queries on relational data
Avgustinov, P., de Moor, O., Jones, M. P., and Schäfer, M · 2016
Earlier work this paper cites.
Squad: 100, 000+ questions for machine comprehension of text
Rajpurkar, P., Zhang, J., Lopyrev, K., and Liang, P · 2016
Earlier work this paper cites.
Probabilistic model for code with decision trees
Raychev, V., Bielik, P., and Vechev, M · 2016
Earlier work this paper cites.
Learning a static analyzer from data
Bielik, P., Raychev, V., and Vechev, M. T · 2017
Earlier work this paper cites.
Simple and effective multi-paragraph reading comprehension
Clark, C. and Gardner, M · 2017
Earlier work this paper cites.
Fixing weight decay regularization in adam
Loshchilov, I. and Hutter, F · 2017
Earlier work this paper cites.
Making neural QA as simple as possible but not simpler
Weissenborn, D., Wiese, G., and Seiffe, L · 2017
Earlier work this paper cites.
Active learning of points-to specifications
Bastani, O., Sharma, R., Aiken, A., and Liang, P · 2018
Earlier work this paper cites.
Deep code search
Gu, X., Zhang, H., and Kim, S · 2018
Earlier work this paper cites.
Deep learning type inference
Hellendoorn, V. J., Bird, C., Barr, E. T., and Allamanis, M · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Rajpurkar, P., Jia, R., and Liang, P · 2018
Earlier work this paper cites.
Learning loop invariants for program verification
Si, X., Dai, H., Raghothaman, M., Naik, M., and Song, L · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Yang, Z., Qi, P., Zhang, S., Bengio, Y., Cohen, W. W., Salakhutdinov, R., and Manning, C. D · 2018
Earlier work this paper cites.
StaQC: A systematically mined question-code dataset from stack overflow
Yao, Z., Weld, D. S., Chen, W.-P., and Sun, H · 2018
Earlier work this paper cites.
When deep learning met code search
Cambronero, J., Li, H., Kim, S., Sen, K., and Chandra, S · 2019
Earlier work this paper cites.
Scalable taint specification inference with big code
Chibotaru, V., Bichsel, B., Raychev, V., and Vechev, M. T · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Codesearchnet challenge: Evaluating the state of semantic code search
Husain, H., Wu, H., Gazit, T., Allamanis, M., and Brockschmidt, M · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
Neural reverse engineering of stripped binaries using augmented control flow graphs
David, Y., Alon, U., and Yahav, E · 2020
Cited alongside, same era.
Cosqa: 20, 000+ web queries for code search and question answering
Huang, J., Tang, D., Shou, L., Gong, M., Xu, K., Jiang, D., Zhou, M., and Duan, N · 2021
Later among the works it cites.
Codeqa: A question answering dataset for source code comprehension
Liu, C. and Wan, X · 2021
Later among the works it cites.
Type4py: Deep similarity learning-based type inference for python
Mir, A. M., Latoskinas, E., Proksch, S., and Gousios, G · 2021
Later among the works it cites.
CodeTrek: Flexible Modeling of Code using an Extensible Relational Representation
Pashakhanloo, P., Naik, A., Wang, Y., Dai, H., Maniatis, P., and Naik, M · 2021
Later among the works it cites.
CS1QA: A dataset for assisting code-based question answering in an introductory programming course
Lee, C., Seonwoo, Y., and Oh, A · 2022
Closest in time.
Training language models to follow instructions with human feedback, 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Feng, Z., Guo, D., Tang, D., Duan, N., Feng, X., Gong, M., Shou, L., Qin, B., Liu, T., Jiang, D., and Zhou, M · 2020
Cited alongside, same era.
Graphcodebert: Pre-training code representations with data flow
Guo, D., Ren, S., Lu, S., Feng, Z., Tang, D., Liu, S., Zhou, L., Duan, N., Svyatkovskiy, A., Fu, S., et al · 2020
Cited alongside, same era.
Neural code search revisited: Enhancing code snippet retrieval through natural language intent
Heyman, G. and Cutsem, T. V · 2020
Cited alongside, same era.
Learning and evaluating contextual embedding of source code
Kanade, A., Maniatis, P., Balakrishnan, G., and Shi, K · 2020
Cited alongside, same era.
Opttyper: Probabilistic type inference by optimising logical and natural constraints
Pandi, I. V., Barr, E. T., Gordon, A. D., and Sutton, C · 2020
Cited alongside, same era.
Typewriter: neural type prediction with search-based validation
Pradel, M., Gousios, G., Liu, J., and Chandra, S · 2020
Cited alongside, same era.
Lambdanet: Probabilistic type inference using graph neural networks
Wei, J., Goyal, M., Durrett, G., and Dillig, I · 2020
Cited alongside, same era.
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., and Lowe, R · 2022
Closest in time.
Learning to walk over relational graphs of source code
Pashakhanloo, P., Naik, A., Dai, H., Maniatis, P., and Naik, M · 2022
Closest in time.
Static inference meets deep learning: a hybrid type inference approach for python
Peng, Y., Gao, C., Li, Z., Gao, B., Lo, D., Zhang, Q., and Lyu, M · 2022
Closest in time.
https://github.com/github/codeql/blob/main/python/ql/src/codeql-suites/python-lgtm.qls , 2022
Query Suite · 2022
Closest in time.
Google · 2023
Closest in time.
Evidence of meaning in language models trained on programs, 2023
Jin, C. and Rinard, M · 2023
Closest in time.
Starcoder: may the source be with you!
Li, R., Allal, L. B., Zi, Y., Muennighoff, N., Kocetkov, D., Mou, C., Marone, M., Akiki, C., Li, J., Chim, J., et al · 2023
Closest in time.
Codegen2: Lessons for training llms on programming and natural languages
Nijkamp, E., Hayashi, H., Xiong, C., Savarese, S., and Zhou, Y · 2023
Closest in time.
Can large language models reason about program invariants?
Sutton, C., Bieber, D., Shi, K., Pei, K., and Yin, P · 2023
Closest in time.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Closest in time.
https://github.com/tree-sitter/tree-sitter , 2021
tree-sitter project · 2023
Closest in time.