{"id":"21b398f1-2e77-479f-9ea4-2e4d2123952f","arxiv_id":"2501.12134","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In an empirical study of 437 code snippets, Bing CoPilot usually provided at least one link with code similar to its output, while Gemini rarely did, exposing a real provenance gap.","lead":"This study tested whether two AI code assistants, Bing CoPilot and Google Gemini, give users links to code that actually matches what they generate. It found that Bing's links were often relevant while Gemini's mostly were not, so users cannot rely on these tools to show where generated code came from.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's headline percentages (66%/28%) are not derivable from Tables IV or VI, and the paper does not define the aggregation, leaving the central quantitative claim unverifiable.","rationale":"The reader's weakest_assumption correctly identifies a real limitation, but it is acknowledged in the paper and is a known difficulty for any provenance study; accepting it would not change the qualitative conclusion that Gemini provides almost no relevant links. The deeper issue is that the abstract's quantitative headline is not supported by the reported tables. The paper presents no formula linking the 66% and 28% figures to the per-language measures in Tables IV and VI. Since the authors made the raw data available, the reader can settle this in minutes. The manual annotation process is solid (inter-rater agreement 0.77 for the key judgment, 1,520 links), and the automated similarity results are consistent with the manual assessment, which makes the qualitative finding credible but does not rescue the undefined aggregate. I therefore agree with the CONDITIONAL verdict, and I see this as the same condition: require the authors to reconcile the metrics before final publication.","tokens_in":21784,"tokens_out":3974,"duration_ms":40539,"concrete_test":"Download the replication package (Zenodo DOI 10.5281/zenodo.13151631) and recompute the proportion of links manually annotated as 'relevant to the generated snippet' for Bing CoPilot and Gemini separately, both at link level and query level. Check whether 66% and 28% match either computation, or whether they instead match the 'relevant to query' column of Table IV aggregated over languages. If the numbers cannot be reproduced from the raw annotations, the abstract should be corrected to report the table-derived metrics, or the missing aggregation must be specified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, quantified as '66% of the links from Bing CoPilot and 28% from Google Gemini are relevant', rests on a metric that is never defined at the level those numbers are reported. Table IV reports link relevance to the query, not to the generated snippet; Table VI reports the percentage of queries with at least one relevant link (Bing 47.8%–84%, Gemini 0%–14.3% depending on language). Neither table yields an aggregate of 66% or 28%, and the paper does not state whether these percentages are link-level, query-level, or weighted across languages. The reader therefore cannot check whether the flagship numbers support the conclusion. The direction-of-reuse threat is acknowledged in Section V, which notes 28 Bing sources and 1 Gemini source were updated after the model checkpoints, but that would only change the interpretation of positive links, not the fact that a precise aggregate is missing. Since the replication package is public, this is a checkable reporting gap rather than an unfalsifiable threat.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an empirical study of whether two LLM-based assistants, Bing CoPilot and Google Gemini, provide external links that contain code snippets similar to the code they generate. The authors use 99 CodeSearchNet tasks across six languages, issue prompts to the assistants, and analyze the returned links through manual annotation (with multiple annotators and reported Cohen's kappa) and automated clone detection/textual similarity. The main finding is that Bing CoPilot provides at least one snippet-relevant link for a majority of queries (about 48%–84% depending on language), while Gemini does so in only 0%–14% of queries. The paper concludes that LLM-provided links are noisy and suffer from 'provenance debt'.","tokens_in":21997,"tokens_out":8101,"duration_ms":77148,"significance":"If the results are reported accurately, this is a useful and timely empirical contribution: it is the first study, to the authors' knowledge, to evaluate the usefulness of LLM-provided links for code provenance. The study's strengths include a public replication package (Dataset DOI provided), a reasonably large manually annotated set of 1,520 links, the use of two independent annotators with reported inter-rater agreement (kappa values of 0.77–0.91), and the combination of manual assessment with automated clone detection and textual similarity analysis. The qualitative conclusion—that assistant-provided links often do not allow developers to verify the origin or licensing status of generated code—is practically relevant. However, the abstract's headline numbers (66% and 28%) are not consistently supported by the tables, and the aggregation method is not defined, which undermines the verifiability of the quantitative claim.","major_comments":[{"comment":"The abstract's statement that '66% of the links from Bing CoPilot and 28% from Google Gemini are relevant' cannot be verified from the reported data. Table VI reports the percentage of queries (per language) with at least one snippet-relevant link, not the percentage of links. Weighting those percentages by the number of queries per language in Table I yields about 66% for Bing, so 66% appears to be a query-level aggregate mislabeled as a link-level figure. For Gemini, the same weighted calculation yields about 6%, and no table or combination of tables in the paper yields 28%. The abstract's 28% also contradicts the RQ2 summary in Section III.B, which states that Gemini's relevant-link percentages are below 15%. Please correct the abstract, state explicitly whether the figures are link-level or query-level, and ensure every reported aggregate is derivable from the tables or the replication package.","section":"Abstract; Tables IV and VI"},{"comment":"The analysis methodology does not describe how the aggregate percentages in the abstract are obtained. Section II.E explains the per-language analysis for RQ2 but does not state whether overall figures are weighted by number of queries, weighted by number of links, or simple averages of the per-language percentages. Without this description, the reader cannot reproduce the headline numbers. Please provide the formula (or the exact computation script) used to aggregate the per-language results, and consider adding the overall row to Table VI.","section":"Section II.E"},{"comment":"The threat that the direction of reuse may be reversed (i.e., some web sources updated after the LLM checkpoints may have copied the model's output) is acknowledged, but its impact on the central finding is not assessed. The paper reports that 28 Bing sources and 1 Gemini source were updated between the LLM checkpoints and the analysis, yet it does not say how many of the 'yes' relevant links fall into this set. Please quantify the overlap between the 'yes' links and the sources updated after the checkpoints, or otherwise bound the effect on the reported percentages. This would let the reader judge whether the 'relevant' figures are upper bounds and by how much.","section":"Section V"}],"minor_comments":[{"comment":"The sentence 'For Bing CoPilot, we obtained higher percentages of relevant links (41.03% for Ruby up to 61.79% for Java vs 24.26% for Python up to 56% for Go)' is ambiguous because it is not clear which percentages refer to Bing CoPilot and which to Gemini; please label the two assistants explicitly.","section":"Section III.A"},{"comment":"The RQ2 summary sentence ends with a comma and no period: 'Bing CoPilot is able to provide at least one “relevant” link for the majority of the considered queries,'. Please correct the punctuation.","section":"Section III.B"},{"comment":"The column headers 'One' and '> One' should be expanded to 'Exactly one' and 'More than one' to avoid ambiguity about the meaning of the columns.","section":"Table VI"},{"comment":"The paper reports the number of links classified as 'Unreachable' and 'Not a Source' together as NA in Table IV; reporting these two categories separately would help interpret the often-high NA percentages and let readers judge how many links were simply unreachable at analysis time.","section":"Section II.C"},{"comment":"The conclusion lists '(i) a general noise of the produced links, and (iii) a limited ability...' but no (ii); please renumber the items or insert the missing point.","section":"Section VII"},{"comment":"The phrase 'computed across snippets from links classified as “not relevant” and “relevant” are statistically significant' is awkward; consider rephrasing to 'computed for links classified as relevant versus not relevant'.","section":"Section III.B"}],"recommendation":"major_revision","confidential_remarks":"This is a solid empirical study with a public replication package and transparent manual annotation procedures. The main concern is the discrepancy between the abstract's headline numbers and the tables, which is fixable with a careful revision. The qualitative findings (Bing often provides at least one relevant link; Gemini rarely does) appear well supported by Table VI. I see no reason to doubt the core methodology. The paper fits the scope of an empirical software engineering journal."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read on arXiv:2501.12134. The paper is a genuine empirical contribution: first study I know of that checks whether the links that Bing CoPilot and Gemini attach to generated code actually point to code the snippet resembles. Design is careful—two annotators per link, kappa 0.77–0.91, clone detection plus textual similarity, public replication package. The qualitative findings are solid: Bing's links skew toward Stack Overflow/GitHub and often contain similar code; Gemini's skew toward official docs that rarely contain the snippet; both produce plenty of noise. Useful for anyone thinking about licensing or trust in LLM output.\n\nThe soft spot is real: the abstract's '66% of the links from Bing CoPilot and 28% from Google Gemini are relevant' cannot be derived from Tables IV or VI. Table IV is per-link relevance to the query; Table VI is per-query 'at least one relevant link' (the unweighted language average for Bing is ~67%, which may be where 66% came from, but that's not 'of the links'). For Gemini the average of Table VI's yes column is ~7%, nowhere near 28%. The paper never defines whether the headline numbers are link-level, query-level, or weighted. Since the replication package is public, this is a fixable reporting error, but as written the central quantitative claim is unverifiable.\n\nThe direction-of-reuse threat (Section V, 28 Bing sources updated after checkpoints) is acknowledged and does not break the qualitative conclusion—the links are still noisy and often unhelpful. But the paper should soften 'likely originate' or strengthen the case that the similarity is not mostly due to pages copying the model.\n\nBottom line: deserves a serious referee. The method and data are a real step forward for provenance research. The authors must clarify the aggregation metric, reconcile the abstract numbers with the tables, and frame provenance direction more carefully. With those changes I'd point others to it.","headline":"Useful empirical study of LLM link relevance, but the abstract's headline percentages are unreconciled with the tables—fix that before citing.","tokens_in":22478,"tokens_out":3783,"would_cite":true,"duration_ms":35189,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that LLM-based assistants Bing CoPilot and Gemini frequently provide links that do not lead to the code they generated, a problem the authors call 'provenance debt'.","keywords":["code provenance","large language models","LLM-based assistants","code clone detection","Bing CoPilot","Google Gemini","code licensing","empirical software engineering"],"falsifier":"Collect cases where the true origin of a generated snippet is independently known, for example, code that was published to a public repository with a timestamp clearly before the assistant's training cutoff; then prompt the assistant for that exact task and check whether its provided links include that known source and whether the clone ratio flags it. If, in such known-origin cases, the assistants' 'relevant' links are no more likely than chance to contain the true source, the provenance-debt conclusion would be weakened or overturned.","tokens_in":21638,"feed_emoji":"📎","tokens_out":3162,"duration_ms":32292,"temperature":0.7,"pith_summary":"This paper asks whether LLM-based assistants can tell a developer where an automatically generated code snippet actually came from. Using 243 snippets from Bing CoPilot and 194 from Google Gemini across six languages, it checks whether the links each assistant provides contain code similar enough to count as the likely origin. The answer, according to the study, is mixed: 66% of Bing CoPilot's links are relevant, but only 28% of Gemini's are, and many links are noisy, unreachable, or point to code in another language. If correct, this means developers cannot rely on the assistants' own citations to determine trust, attribution, or licensing terms for generated code.","feed_headline":"Copilot's code links match its output 66% of the time; Gemini's only 28%","feed_subtitle":"A 1,500-link study finds serious provenance debt: developers can't trust assistants to say where generated code came from.","key_machinery":"The argument is carried by a two-layer similarity analysis. First, the authors scrape code snippets from the landing pages of 1,520 links (1,006 from Bing CoPilot, 514 from Gemini) and run CCFinderSW, a token-based clone detector, with a threshold of at least 20 tokens, computing a cloning ratio between each linked snippet and the LLM-generated snippet. Second, they complement this with a cosine textual similarity over code tokens and, crucially, a manual annotation by five authors who judge link type, relevance to the query, and whether the linked snippet likely originated the generated code. The manual judgments are the ground truth; the automated metrics are evaluated against them.","core_discovery":"The paper's central claim is that LLM-based assistants suffer from serious 'provenance debt': the external links they provide for a coding query only sometimes lead to code that resembles the snippet the assistant generated. Bing CoPilot gave at least one relevant link for the majority of the queries studied, with the share of queries having at least one relevant link ranging from 47.8% (Java) to 84% (PHP); Gemini rarely gave any relevant link, with shares mostly below 15% and 0% for Ruby. The authors also find that the assistants' link lists are heterogeneous and noisy, that they do not replicate what a straightforward search-engine query would return, and that even when a link is relevant there may be several candidate sources, making the actual origin ambiguous.","pith_inferences":["The 66% and 28% relevance figures likely understate or overstate the true provenance rate depending on direction of reuse: the paper itself notes 28 Bing sources and 1 Gemini source were updated after the models' checkpoints, so some 'relevant' links may have copied from the assistants rather than the reverse.","A practical extension would build a provenance-ranking tool that automatically takes an LLM's generated snippet, scrapes the provided links, runs clone and textual similarity, and ranks candidate sources; the paper's methodology is essentially a blueprint for that tool.","If provenance debt is a general property of post-generation attribution, then newer or differently architected assistants that rely on the same strategy would likely show similar patterns unless they are trained specifically to emit source links.","A testable next step would apply the same protocol to assistants that expose training-cutoff dates or to retrieval-augmented systems, to see whether explicit retrieval improves the relevance of provenance links."],"forward_implications":["If the paper is right, developers cannot use assistant-provided links as reliable evidence of where generated code came from, so license compliance and trust decisions remain largely unsupported.","Users must expect to inspect several links per query, since many are irrelevant, unreachable, or written for a different programming language, and sometimes more than one plausible source exists.","Clone detection and textual similarity, as calibrated and used here, can partially distinguish relevant from irrelevant provenance links, suggesting these tools could be embedded in developer-facing assistants.","Because the assistants' link choices do not match what their underlying search engines return for the same query, provenance cannot be reproduced or audited by simply re-running a web search.","Gemini's near-absence of relevant links indicates that its attribution support, at least at the time of the study, is far weaker than Bing CoPilot's for code generation tasks."],"supporting_citations":[{"why":"Supplies the CodeSearchNet queries, the realistic coding tasks in six languages used to prompt both assistants.","marker":"[27]"},{"why":"CCFinderSW is the token-based clone detector used to compute cloning ratios between generated snippets and snippets scraped from the provided links.","marker":"[55]"},{"why":"The survey of LLM attribution strategies frames why post-generation attribution, the mechanism used by Bing CoPilot and Gemini, is the object of study.","marker":"[36]"},{"why":"Prior work using clone detection to trace code provenance, cited as the technical precedent for treating clone evidence as provenance evidence.","marker":"[19]"},{"why":"The closest prior study, on LLM copyright violations; the paper contrasts its own provenance goal with memorization findings.","marker":"[31]"},{"why":"Documents license inconsistencies in LLM training datasets, motivating the need for trustworthy provenance links.","marker":"[32]"}],"fun_headline_variants":["Bing CoPilot's code links match its output 66% of the time; Gemini 28%","LLM 'provenance debt': Copilot links more relevant than Gemini's","Study: Gemini's code links rarely match its generated snippets","Bing CoPilot beats Gemini on code provenance link relevance","Provenance debt: LLMs give code links that often don't match output"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole analysis treats high similarity between a generated snippet and a snippet reachable through a provided link as evidence that the link is the source, even though the paper cannot prove the direction of reuse and found 28 Bing sources and 1 Gemini source that were updated after the assistants' checkpoints.","fun_headline_variants_meta":{"raw":{"variants":["Bing CoPilot's code links match its output 66% of the time; Gemini 28%","LLM 'provenance debt': Copilot links more relevant than Gemini's","Study: Gemini's code links rarely match its generated snippets","Bing CoPilot beats Gemini on code provenance link relevance","Provenance debt: LLMs give code links that often don't match output"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000782,"raw_usage":{"total_tokens":3431,"prompt_tokens":898,"completion_tokens":2533,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":514,"completion_tokens_details":{"reasoning_tokens":2430}},"tokens_in":514,"tokens_out":2533,"duration_ms":19195,"temperature":1.0,"reasoning_tokens":2430,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:27:37.467077+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect cases where the true origin of a generated snippet is independently known, for example, code that was published to a public repository with a timestamp clearly before the assistant's training cutoff; then prompt the assistant for that exact task and check whether its provided links include that known source and whether the clone ratio flags it. If, in such known-origin cases, the assistants' 'relevant' links are no more likely than chance to contain the true source, the provenance-debt conclusion would be weakened or overturned.","supporting_citations":[{"cited_title":"CCFinderSW: clone detection tool with flexible multilingual tokenization,","cited_arxiv_id":null,"evidence_quote":"CCFinderSW is the token-based clone detector used to compute cloning ratios between generated snippets and snippets scraped from the provided links."}],"review_version":1}