Pith. sign in

Paper Citation Record · LEDGER

Copilot Arena: A Platform for Code LLM Evaluation in the Wild

As of 9 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 5 inbound Pith citation observations for arXiv:2502.09328.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09328 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:56:20.506734Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:14:20.943931Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T02:05:18.872540Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9694f2f-22ba-475f-8aca-298ac64193f1 · outbound

This paper cites The Impact of AI on Developer Productivity: Evidence from GitHub Copilot.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild The Impact of AI on Developer Productivity: Evidence from GitHub Copilot

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.168071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.168071Z digest=sha256:a87e888cbafd25d0b327d6ccf513ac0445be85f1b61ee8795405b8c7b5086db3

Observation 1808e117-8aa5-4a54-b304-ef711dd3c846 · outbound

This paper cites an unresolved cited work.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T21:56:21.646007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.173571Z digest=sha256:484092e48146d85a1f7d05fb89b0d7d4991557739346b4e138dac34855e4a01c

Observation b24950e9-baac-4b14-be37-e8573b892b65 · outbound

This paper cites Beyond the bar: Generative ai as a transformative component in legal document review.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Beyond the bar: Generative ai as a transformative component in legal document review

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.631032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.179095Z digest=sha256:d4c7deedba79e2a206d2e44d2eded4d84f7299ed8dbad14ca6a0b61488cc77aa

Observation 5cde8a51-e30d-43c0-8e5f-fcec4a4c94a6 · outbound

This paper cites Evaluation gaps in machine learning practice.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Evaluation gaps in machine learning practice

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.615702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.184080Z digest=sha256:75e88f8875ad3c15857bd4490116f9678758aed1d1063d61f867c34a96285b4a

Observation 2c00d2a4-52e3-48fc-8eb9-ecf6554c787c · outbound

This paper cites Benchmarks as microscopes: A call for model metrology.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Benchmarks as microscopes: A call for model metrology

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.598549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.189130Z digest=sha256:ecaf9216abf5366f85e7b93599e3dad47e0d39a0fad6d055bd36ce3ffc451a2c

Observation 240c7314-fc22-43d1-8465-4e6b8da2d046 · outbound

This paper cites AI Agents That Matter.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild AI Agents That Matter

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.194237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.194237Z digest=sha256:77406ba5591e7e081e38d5397c6ab12ec20cc053adcf407f380f57998511674c

Observation 9f70086e-37b1-4b82-a41c-c9680b1320e0 · outbound

This paper cites Gonzalez, and Ion Stoica.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Gonzalez, and Ion Stoica

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.580094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.199962Z digest=sha256:d5611c2bd8400aea46d20435b6bdf5a64bb915d2845968075239a2dc40d2b361

Observation f646da96-c0f9-4f8c-ab09-ab4e3fbd3258 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Alpacafarm: A simulation framework for methods that learn from human feedback

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.548460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.209620Z digest=sha256:563ba360e93dfad74866f137074a878a279d95b1264fe3191f4b8b18bbd843bb

Observation 6082a519-ccdc-4720-b3b5-68e5a11a0060 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.214333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.214333Z digest=sha256:646d1d18e6fd8c3d5f45928a88aa9294abb9c33e6b8d574b15b2e3743c77d173

Observation 346acf53-775c-4ac9-a2b7-5cb7936e4c2e · outbound

This paper cites an unresolved cited work.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T21:56:21.532565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.220493Z digest=sha256:47b338d50a00ab3ef72375a4823809c68a45189a82fbbadb259f4fae6b1b91a1

Observation 16463f74-676b-4bd1-821f-23b542182623 · outbound

This paper cites Github copilot - your ai pair programmer, 2022.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Github copilot - your ai pair programmer, 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.225175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.225175Z digest=sha256:f1d03da3c82b3072bff29dca98b9fb1749bc2015cb0fc0068bb99e981efcf8e3

Observation e5a815bb-1701-4abf-bd59-f9abdcca4e98 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Evaluating Large Language Models Trained on Code

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.229816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.229816Z digest=sha256:88483b4d51f95506da469cba29b4ab1b3e94c9f9e5e97b36db9a8687bd2661e4

Observation e8aa71b9-8cd6-4112-b3ac-31e0048363dc · outbound

This paper cites Program Synthesis with Large Language Models.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Program Synthesis with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.235285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.235285Z digest=sha256:e11c71ddccd922e20e0529e92b693301b0eff5d64d8f4adaaf07f9825fd14a76

Observation 55e02bf3-6dba-4b3e-9b25-6057b5377ef2 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.240188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.240188Z digest=sha256:04dbc70bea18e095cc2fed9c63e67758d4a29bfd70cf0daa7c6462f2a7d752a4

Observation f2d5671b-b15b-4c51-a418-624fb29b0a7b · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.245128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.245128Z digest=sha256:656fca274e1f691507b277eaabb034f2c801baac58a84a3c399ecbc1b97588c3

Observation ea8d7c12-bed4-4a89-a160-bb035c3a3a3e · outbound

This paper cites Expectation vs.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Expectation vs

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.506828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.250065Z digest=sha256:6c06a6f9aa40ae4ff81af767eebd764770ba775f8da33cdaa35f9559237b923f

Observation 8a4acac1-cdb5-467f-9cee-0ecc5e202288 · outbound

This paper cites The programmer’s assistant: Conversational interaction with a large language model for software development.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild The programmer’s assistant: Conversational interaction with a large language model for software development

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.491197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.254679Z digest=sha256:516d63faa15b245b57579d338917e357c21ec6ddc9c43841ba04263001c7a3fb

Observation 905af829-b390-43e0-96fa-7d373ca178a1 · outbound

This paper cites The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.259808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.259808Z digest=sha256:d225ecb32d91df4d79053c9b66c711295b69b66a8a4531ea3bfaa08e809e28a6

Observation 9c531192-deec-4d34-a691-ed4fecd7659b · outbound

This paper cites WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.264700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.264700Z digest=sha256:87991a08cbdfb6f501a713725f4be919818535e8fa6e97cfc9fb8996b228db9d

Observation 8c760a64-64f7-46de-b07d-c0724c261bf5 · outbound

This paper cites Dai, and Quoc V Le.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Dai, and Quoc V Le

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.475509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.269619Z digest=sha256:b5fb600b137b3ffe6c004e5ecc0f4d571fe7d3d038219daee11ca6d8f0c3f129

Observation 822731ef-f5bb-4846-89b1-67623ba84ad2 · outbound

This paper cites InCoder: A Generative Model for Code Infilling and Synthesis.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild InCoder: A Generative Model for Code Infilling and Synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.274059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.274059Z digest=sha256:633be7470b14c0e9355b830827fc80251286d0f6a5f3b60fbcc439fd825b8ae4

Observation 3edb20c8-079b-40e2-86b8-118afdcc51cb · outbound

This paper cites Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.278997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.278997Z digest=sha256:77250ffcad52861154e3cd5467c5d3e874cc0c9c182a53d49b58a967c6300713

Observation 312abebc-6725-424b-b48d-900cfeda23ff · outbound

This paper cites Efficient Training of Language Models to Fill in the Middle.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Efficient Training of Language Models to Fill in the Middle

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.284825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.284825Z digest=sha256:5507fe8699659cd8a21e14ecc13f6045b89e88567a18a6aefad3c16e3e9bac96

Observation 2a3788e9-ddf5-4148-b7e3-8a95edbcc856 · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.290154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.290154Z digest=sha256:1687053c14183f87390b931df969e4d62b8200afcbadcbda6372b6c31a35544c

Observation 0f491527-cc55-447c-8fe1-07393d1398f8 · outbound

This paper cites Raising the bar on swe-bench verified with claude 3.5 sonnet, 2024.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Raising the bar on swe-bench verified with claude 3.5 sonnet, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.460778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.296266Z digest=sha256:0f4e096e7ad5c544aa3f1525c6a45c93a796266de0a5dca0374fbb0223b969c8

Observation 3ff4f5ef-02ae-4937-95c4-7ec20ac61431 · outbound

This paper cites Rank analysis of incomplete block designs: I.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Rank analysis of incomplete block designs: I

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.301371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.301371Z digest=sha256:5617fa937a05efa9871c5e5284ce242e96e9825d2068440ff144f9dc4ae58df2

Observation 80d70cea-7776-4ea4-894a-4a8507497d32 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.306648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.306648Z digest=sha256:5790108d1a08b719e00560ec38970e0ebd0467924306deba17631f26f8518a22

Observation 47fc453b-f8fd-4c5f-b88a-65018b72466a · outbound

This paper cites HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.311466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.311466Z digest=sha256:b236e3bba54cf3ad12dc1b216fce7c8b17b2c07725a84390aa8f7868495f51b1

Observation 75aa717a-b98f-46bb-9619-e3f330a73e18 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.317151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.317151Z digest=sha256:43f37a6c8ae10b86b2d045cea03529e1c250868a75021af214245909a31c0cae

Observation 7e6f9726-e1cb-415b-8763-17060bc3a921 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.321916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.321916Z digest=sha256:b0261ddf6605cbb83d1af12a341d6a899626930c3c9d3dd263b895f0fd2cea69

Observation 1341167b-ca6a-44c5-9823-7b6ac76e7ffc · outbound

This paper cites Fine-grained human feedback gives better rewards for language model training.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Fine-grained human feedback gives better rewards for language model training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.326862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.326862Z digest=sha256:3bf316e8058e6cf9015cac3108c1ff3366a77f2ba8b4269970e47fa285f851fd

Observation 3ae649dc-bcb4-4e30-9704-7e0ef0cfd0c6 · outbound

This paper cites Training Language Models with Language Feedback.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Training Language Models with Language Feedback

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.331588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.331588Z digest=sha256:fe2d43065dff2a1f27ba8dd796de94f5cdf8c126d6ca47a7bd85be70c7cb33c6

Observation 09a3ee1f-f729-4d38-85e7-3627f55b2e82 · outbound

This paper cites Training language models to follow instructions with human feedback.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.337471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.337471Z digest=sha256:d75b5eb717107d380b30da24b7a82fe885c7cd84fcda939b4e6f2c87afd55018

Observation b896b753-d32f-4ac8-a8a6-34808e5425df · outbound

This paper cites VisionArena: 230K Real World User-VLM Conversations with Preference Labels.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild VisionArena: 230K Real World User-VLM Conversations with Preference Labels

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.342641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.342641Z digest=sha256:f7172178554a240f33e4a3cbf896debf9a6edba6e940427381c58c71ca10a636

Observation 21c37956-7be2-4386-bdd3-8d6eb69d2a9f · outbound

This paper cites CodeXGLUE: A machine learning benchmark dataset for code understanding and generation.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild CodeXGLUE: A machine learning benchmark dataset for code understanding and generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.416663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.348807Z digest=sha256:97b2f0eb7b10c4696d53c768e61e1216548dfff82358667963bbd89d81f66700

Observation d08b4ecc-dc20-4e2d-9c92-c1ac00d611f1 · outbound

This paper cites Codegen: An open large language model for code with multi-turn program synthesis.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Codegen: An open large language model for code with multi-turn program synthesis

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.402072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.353529Z digest=sha256:374447d4a739436a99d984129296a20d23e8c73774e738853da65f23af4f7a99

Observation 2556564c-1f34-4267-9ff8-4414bd9cb7b6 · outbound

This paper cites XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.364223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.364223Z digest=sha256:4230fcce4bde173dfc3a826d352dc5cdf05f0019f81770e17dadf37ea234914f

Observation 94c4e8c3-2e3c-4bb7-9042-54e0564e8843 · outbound

This paper cites an unresolved cited work.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T21:56:21.387056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.358644Z digest=sha256:2fd000cae868a90e996712da521bf623f278bcf2f2a8d6a8143931c476461845

Observation a3e2a942-0646-476f-9d83-bf11ac275432 · outbound

This paper cites Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.357024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.374664Z digest=sha256:505153e2a099144174ed6a94ac5fedc3737164b164aba25a08ede65bd9982bd8

Observation d8d595ac-f25a-4502-a05b-a361a7b04221 · outbound

This paper cites Recode: Robustness evaluation of code generation models.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Recode: Robustness evaluation of code generation models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.372216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.368974Z digest=sha256:c3daec6a31cebbb2bdf2f66eaeeeb3cd4a9dfdddf9a74637ce1f356ec5a20aba

Observation 7257a80f-831c-4cf0-98f3-f81fb76438fa · outbound

This paper cites xCodeEval: A Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and Retrieval.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild xCodeEval: A Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and Retrieval

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.384490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.384490Z digest=sha256:72de6d655bdef80e8ae822b832c2520e7c6e2adbaa2f684c48ba66148303b0e4

Observation 0badcd11-4377-4db8-a58c-05d692fc7401 · outbound

This paper cites Swe-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations , 2023.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Swe-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations , 2023

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.341503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.379328Z digest=sha256:4dab8c23eeb73d941959086d45a7df10ba41aa8312312921bb800447573428ef

Observation 86d980e2-f0dd-4c55-94a4-e19944a2c6d7 · outbound

This paper cites Multipl-e: a scalable and polyglot approach to benchmarking neural code generation.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Multipl-e: a scalable and polyglot approach to benchmarking neural code generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.326409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.394446Z digest=sha256:8ba4113b288d40adb38fcf43872b0b99495bc38f71a2fc7cdaadbf146d4d28ac

Observation c4398e9e-b76f-4cea-9e54-431709172ac7 · outbound

This paper cites CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.389296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.389296Z digest=sha256:9e416c91bf662ba405a378f699bb23325f4298abfbd3c1f6eb01a54c703c2ffe

Observation c79ed065-9a12-441f-95ae-8b7b01d4bc71 · outbound

This paper cites Large language models of code fail at completing code with potential bugs.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Large language models of code fail at completing code with potential bugs

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.297005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.403936Z digest=sha256:e4f4bcc04c640c9ca98d6ce27bdef291853e716ef8780c476c7fb714ea517729

Observation 5d8e2407-d303-453f-836f-8d103848fdce · outbound

This paper cites Octopack: Instruction tuning code large language models.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Octopack: Instruction tuning code large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.311951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.399150Z digest=sha256:257b93b69dbea9c5db0312d853372f8e42185735325b2b1f3258501ca964ea29

Observation 46d8f48c-8d10-42b9-919e-3086c271500e · outbound

This paper cites R2e: Turning any github repository into a programming agent environment.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild R2e: Turning any github repository into a programming agent environment

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.267410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.414718Z digest=sha256:05d5af6ce2f135c2862edea9b520a972866fe7c4575d3624aa3fc81188f0d9dd

Observation c04fb848-324c-4f81-8992-9abd600582b4 · outbound

This paper cites Intercode: Stan- dardizing and benchmarking interactive coding with execution feedback.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Intercode: Stan- dardizing and benchmarking interactive coding with execution feedback

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.282470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.409665Z digest=sha256:17d7513b5359524ff0603784b50d2f356de69da924ddc62c20482e0faae5530f

Observation 73350f3a-8230-4851-b9b8-997f8347744a · outbound

This paper cites Grounded copilot: How programmers interact with code-generating models.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Grounded copilot: How programmers interact with code-generating models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.239096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.424972Z digest=sha256:aca2daacd2c475d58ad7991fe43c0f7aa65da0404e065611c93131285a63a84d

Observation 540a8e5f-e9f5-443e-8ff2-e0bf5aaf39ad · outbound

This paper cites Bernstein, and Percy Liang.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Bernstein, and Percy Liang

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.253252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.419351Z digest=sha256:81da62935fe3b3572cb59fe45d4684a70882bda71876def8ab0955c07b107fc0

Observation a8cfb10c-ef32-4ff7-8e29-ddb62d0b92f4 · outbound

This paper cites Ai-assisted code authoring at scale: Fine-tuning, deploying, and mixed methods evaluation.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Ai-assisted code authoring at scale: Fine-tuning, deploying, and mixed methods evaluation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.209861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.435058Z digest=sha256:3a219c42145aea2e50a643fa993e26f4f61a5886da0a202a886adf6bf1a6d10b

Observation 71b5cf91-4e19-456e-a9bb-d7c98a3a2b4a · outbound

This paper cites Reading between the lines: Modeling user behavior and costs in ai-assisted programming.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Reading between the lines: Modeling user behavior and costs in ai-assisted programming

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.223990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.429404Z digest=sha256:2c350945c11743f099c0987e8200b2f8cf3c6ad5cdcdc1e3e2f3b9e11ebd3288

Observation 0e859cd5-d782-42cd-a18b-1f2bb70dc091 · outbound

This paper cites The productivity effects of generative ai: Evidence from a field experiment with github copilot.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild The productivity effects of generative ai: Evidence from a field experiment with github copilot

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.194986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.444550Z digest=sha256:48d67ff2e0154ad746043a7de4d32340af5b374ccffa90e338b9050faa5efcc1

Observation ccc36f92-77ba-42e7-ab8e-2eb8904eee4f · outbound

This paper cites Need Help? Designing Proactive AI Assistants for Programming.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Need Help? Designing Proactive AI Assistants for Programming

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.439848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.439848Z digest=sha256:5ee881f246f4d4b09d5446e7b1af6612e9a7888e9e35d6e6952185ec2d8f6eba

Observation 31933b58-dde6-4e68-8a3c-cf9bf601c591 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.454878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.454878Z digest=sha256:e9b89865301e564c5ffefe53b41f77b3cb6231f233745d2d825d32dfdee5e95f

Observation 2b9efc6b-d5d2-4904-884f-2b24fcda3f97 · outbound

This paper cites Language Models for Code Completion: A Practical Evaluation.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Language Models for Code Completion: A Practical Evaluation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.449996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.449996Z digest=sha256:082e7aa31cc45ad9afbff22e80cb116f24d14aca92decce5a0f6a54e3093e779

Observation b5e176e0-f468-4289-ba89-d9ffca8f0654 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild A Long Way to Go: Investigating Length Correlations in RLHF

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T21:56:20.465030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:56:20.465030Z digest=sha256:f4ae5e14302902c8812df344d8e59f8428a39841b31297d6595e3a2b928d08c5

Observation f7022013-a03c-4689-89cb-c5c3b4b5b5c0 · outbound

This paper cites an unresolved cited work.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T21:56:21.179604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.460514Z digest=sha256:04a8565bc4d6174f85cc67e7955fbdbd90f7ecc3cc9851167829a287d83acc37

Observation bb660929-4a3a-47fd-a5df-dbd61cad2744 · outbound

This paper cites PSM presents the code context in the order of prefix and then suffix, using XML notation to demarcate prefix, suffix, and middle segments (e.g., <PREFIX> and </PREFIX>).

Copilot Arena: A Platform for Code LLM Evaluation in the Wild PSM presents the code context in the order of prefix and then suffix, using XML notation to demarcate prefix, suffix, and middle segments (e.g., <PREFIX> and </PREFIX>)

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.164432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.470711Z digest=sha256:4d4017d8e8b18c32780f78b94d16c782881d49409a5d946ba3dbc57cb3856530

Observation 791a6138-a9d7-41d7-bdd7-cce2f8ebc414 · outbound

This paper cites SPM is identical to PSM except that the suffix appears before the prefix, which may be more natural than having the suffix appear directly before the output as is the case with PSM.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild SPM is identical to PSM except that the suffix appears before the prefix, which may be more natural than having the suffix appear directly before the output as is the case with PSM

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.148917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.475346Z digest=sha256:52c71b823fad6f44c485eaf71615e1f20bd6608c6a73bdea816024b377b407db

Observation 59563f00-3eb9-4531-a2ea-6321800fa2af · outbound

This paper cites sentinel.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild sentinel

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.130744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.480939Z digest=sha256:c44ed8b9d74f1ee7e968238272dd98aaf32c9f553f4b8e528c64eeda82424d62

Observation 9c32a364-eb9a-4742-b1de-23f3c3bfc799 · outbound

This paper cites pre-fill.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild pre-fill

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.114665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.485696Z digest=sha256:dbb3b0d5418a994c5fdac3703a19a595eef641a01a98765fb81a96e79f1f075d

Observation c358e7f6-d056-4f08-888d-1fc8afcec11a · outbound

This paper cites an unresolved cited work.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T21:56:21.099436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.490889Z digest=sha256:7f301aaa531dc89c891391a67e0075d6dc807bf77f3154167ad08128c7e751f6

Observation 1ef2a7c9-b9d0-44fb-963c-75795b01b002 · outbound

This paper cites an unresolved cited work.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-07T21:56:21.084359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.495527Z digest=sha256:defc325e9060779e6cee03b3ef4a154dc868801f056e569eb2e41bc4bea31906

Observation 2f161a4e-60b7-4ac0-a093-1ffe17e6017e · outbound

This paper cites an unresolved cited work.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-08-07T21:56:21.068356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.501488Z digest=sha256:eb175c7c212ce59663d85de5644785461cfe708dfd3c48b1128ab0211ed2c078

Observation f9adbe6a-8e84-4cf6-b88d-b811ccd129e9 · outbound

This paper cites clusters.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild clusters

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:56:21.052539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.506734Z digest=sha256:e1b30753b624fbdcdefe6a7b6bfefc2fa06f9c400059a1bb6c31a41d8846fb42

Observation 8895e605-430e-4142-b4ee-c77c5f9b66b9 · outbound

This paper cites an unresolved cited work.

Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work

Reference 2023

Resolution
unresolved
raw_fallback, observed 2026-08-07T21:56:21.564333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T21:56:20.204693Z digest=sha256:6779de86a4031828d626dc5f497cd42d9bae7866e1fce3e0837f4be78fb3f59a

Pith citing papers

Observation 8e9c250b-57ee-442f-9b79-f5c9ba4c29c3 · inbound

Structure-Aware Fill-in-the-Middle Pretraining for Code cites this paper.

Structure-Aware Fill-in-the-Middle Pretraining for Code Copilot Arena: A Platform for Code LLM Evaluation in the Wild

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:14:20.943931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:14:20.943931Z digest=sha256:54be63fd909d5a325c616a33d527158e213671c9480f7d03d79e14ee6c7cf221

Observation fc242de9-59c8-4e7b-b077-f0259d48190f · inbound

Mercury: Ultra-Fast Language Models Based on Diffusion cites this paper.

Mercury: Ultra-Fast Language Models Based on Diffusion Copilot Arena: A Platform for Code LLM Evaluation in the Wild

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:05:18.874684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:05:18.834924Z digest=sha256:7a764d575862f337add98bf2dc6f8821d7713f70d39ddc6fd0bb9ef52e2e414b

Observation 94623b0d-e1d2-4c2f-a239-9613c934811f · inbound

Nonparametric LLM Evaluation from Preference Data cites this paper.

Nonparametric LLM Evaluation from Preference Data Copilot Arena: A Platform for Code LLM Evaluation in the Wild

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-03T06:55:31.047609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:55:31.047609Z digest=sha256:4a3b20a32a4f0feb20835a2098098228378bee7b7b7c06f7e6d9c6020e8ad758

Observation 08ada4f0-b026-4806-a4ad-e031b0d3f71e · inbound

Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks cites this paper.

Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks Copilot Arena: A Platform for Code LLM Evaluation in the Wild

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:48.220455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:16:33.635516Z digest=sha256:7585e4a97a2c3450bdfa4bb4d1e52c34edfaa36bc0dc10ca303085be2df98481

Observation 81ec95f4-fff1-4c73-89c7-d11343f5d94f · inbound

RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions cites this paper.

RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions Copilot Arena: A Platform for Code LLM Evaluation in the Wild

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:06:48.354349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T18:39:05.786210Z digest=sha256:3cb5b4c4978e9bcf9beb1361ad15bfe00e24536708eac0b3115fe11fd59ef8e7