Recall Before Rerank: Benchmarking Deep Learning Models for Large-Scale Code-to-Code Retrieval

arXiv:2606.27401 · cs.SE, cs.CL, cs.IR, cs.LG · Submitted 2026-06-24 · Read on arXiv

cs.SE, cs.CL, cs.IR, cs.LG

Submitted: 2026-06-24

Updated: 2026-09-18

Comments: 15 pages, 4 figures. Accepted for publication in the Proceedings of the 27th International Conference on Web Information Systems Engineering (WISE 2026). Preliminary version (differs in formatting and minor revisions from the final camera-ready version). Source code and benchmark are available at https://github.com/leeeov4/code2code_benchmark

Code: https://github.com/leeeov4/code2code_benchmark

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Terminology

Related papers