Pith. sign in

REVIEW 2 cited by

CoSQA: 20,000+ Web Queries for Code Search and Question Answering

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.13239 v1 pith:UOG7JHPT submitted 2021-05-27 cs.CL cs.SE

classification cs.CLcs.SE
keywords codecosqatrainingansweringcoclrcodesfurtherintroduce
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Finding codes given natural language query isb eneficial to the productivity of software developers. Future progress towards better semantic matching between query and code requires richer supervised training resources. To remedy this, we introduce the CoSQA dataset.It includes 20,604 labels for pairs of natural language queries and codes, each annotated by at least 3 human annotators. We further introduce a contrastive learning method dubbed CoCLR to enhance query-code matching, which works as a data augmenter to bring more artificially generated training instances. We show that evaluated on CodeXGLUE with the same CodeBERT model, training on CoSQA improves the accuracy of code question answering by 5.1%, and incorporating CoCLR brings a further improvement of 10.5%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Comparative Study of Specialized LLMs as Dense Retrievers

    cs.IR 2025-07 conditional novelty 5.0 of 10

    Specialized Qwen2.5 7B models differ in dense retrieval quality: math and long-reasoning variants degrade performance, while coder and vision-language variants improve zero-shot text and code retrieval.

  2. SDLog: A Deep Learning Framework for Detecting Sensitive Information in Software Logs

    cs.SE 2025-05 conditional novelty 5.0 of 10

    SDLog, a CodeBERT-based sequence tagger, detects sensitive attributes in software logs with 92.9% F1 across 16 datasets and 98.4% F1 when fine-tuned on 100 target logs, outperforming regex baselines.

Pith tools