Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T21:56:20.506734Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 5 inbound Pith citation observations for arXiv:2502.09328.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T21:56:20.506734Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:14:20.943931Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T02:05:18.872540Z
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c9694f2f-22ba-475f-8aca-298ac64193f1 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild The Impact of AI on Developer Productivity: Evidence from GitHub Copilot
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1808e117-8aa5-4a54-b304-ef711dd3c846 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b24950e9-baac-4b14-be37-e8573b892b65 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Beyond the bar: Generative ai as a transformative component in legal document review
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5cde8a51-e30d-43c0-8e5f-fcec4a4c94a6 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Evaluation gaps in machine learning practice
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2c00d2a4-52e3-48fc-8eb9-ecf6554c787c · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Benchmarks as microscopes: A call for model metrology
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 240c7314-fc22-43d1-8465-4e6b8da2d046 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild AI Agents That Matter
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f70086e-37b1-4b82-a41c-c9680b1320e0 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Gonzalez, and Ion Stoica
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f646da96-c0f9-4f8c-ab09-ab4e3fbd3258 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Alpacafarm: A simulation framework for methods that learn from human feedback
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6082a519-ccdc-4720-b3b5-68e5a11a0060 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 346acf53-775c-4ac9-a2b7-5cb7936e4c2e · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 16463f74-676b-4bd1-821f-23b542182623 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Github copilot - your ai pair programmer, 2022
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5a815bb-1701-4abf-bd59-f9abdcca4e98 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Evaluating Large Language Models Trained on Code
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8aa71b9-8cd6-4112-b3ac-31e0048363dc · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Program Synthesis with Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55e02bf3-6dba-4b3e-9b25-6057b5377ef2 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2d5671b-b15b-4c51-a418-624fb29b0a7b · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea8d7c12-bed4-4a89-a160-bb035c3a3a3e · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Expectation vs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8a4acac1-cdb5-467f-9cee-0ecc5e202288 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild The programmer’s assistant: Conversational interaction with a large language model for software development
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 905af829-b390-43e0-96fa-7d373ca178a1 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild The RealHumanEval: Evaluating Large Language Models' Abilities to Support Programmers
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c531192-deec-4d34-a691-ed4fecd7659b · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c760a64-64f7-46de-b07d-c0724c261bf5 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Dai, and Quoc V Le
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 822731ef-f5bb-4846-89b1-67623ba84ad2 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild InCoder: A Generative Model for Code Infilling and Synthesis
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3edb20c8-079b-40e2-86b8-118afdcc51cb · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Evaluation of LLMs on Syntax-Aware Code Fill-in-the-Middle Tasks
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 312abebc-6725-424b-b48d-900cfeda23ff · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Efficient Training of Language Models to Fill in the Middle
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a3788e9-ddf5-4148-b7e3-8a95edbcc856 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f491527-cc55-447c-8fe1-07393d1398f8 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Raising the bar on swe-bench verified with claude 3.5 sonnet, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3ff4f5ef-02ae-4937-95c4-7ec20ac61431 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Rank analysis of incomplete block designs: I
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80d70cea-7776-4ea4-894a-4a8507497d32 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47fc453b-f8fd-4c5f-b88a-65018b72466a · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild HumanEval-XL: A Multilingual Code Generation Benchmark for Cross-lingual Natural Language Generalization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75aa717a-b98f-46bb-9619-e3f330a73e18 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e6f9726-e1cb-415b-8763-17060bc3a921 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1341167b-ca6a-44c5-9823-7b6ac76e7ffc · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Fine-grained human feedback gives better rewards for language model training
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ae649dc-bcb4-4e30-9704-7e0ef0cfd0c6 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Training Language Models with Language Feedback
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09a3ee1f-f729-4d38-85e7-3627f55b2e82 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Training language models to follow instructions with human feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b896b753-d32f-4ac8-a8a6-34808e5425df · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild VisionArena: 230K Real World User-VLM Conversations with Preference Labels
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c37956-7be2-4386-bdd3-8d6eb69d2a9f · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild CodeXGLUE: A machine learning benchmark dataset for code understanding and generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d08b4ecc-dc20-4e2d-9c92-c1ac00d611f1 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Codegen: An open large language model for code with multi-turn program synthesis
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2556564c-1f34-4267-9ff8-4414bd9cb7b6 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild XLCoST: A Benchmark Dataset for Cross-lingual Code Intelligence
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94c4e8c3-2e3c-4bb7-9042-54e0564e8843 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a3e2a942-0646-476f-9d83-bf11ac275432 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d8d595ac-f25a-4502-a05b-a361a7b04221 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Recode: Robustness evaluation of code generation models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7257a80f-831c-4cf0-98f3-f81fb76438fa · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild xCodeEval: A Large Scale Multilingual Multitask Benchmark for Code Understanding, Generation, Translation and Retrieval
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0badcd11-4377-4db8-a58c-05d692fc7401 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Swe-bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations , 2023
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 86d980e2-f0dd-4c55-94a4-e19944a2c6d7 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Multipl-e: a scalable and polyglot approach to benchmarking neural code generation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c4398e9e-b76f-4cea-9e54-431709172ac7 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild CodeScope: An Execution-based Multilingual Multitask Multidimensional Benchmark for Evaluating LLMs on Code Understanding and Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c79ed065-9a12-441f-95ae-8b7b01d4bc71 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Large language models of code fail at completing code with potential bugs
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5d8e2407-d303-453f-836f-8d103848fdce · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Octopack: Instruction tuning code large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 46d8f48c-8d10-42b9-919e-3086c271500e · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild R2e: Turning any github repository into a programming agent environment
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c04fb848-324c-4f81-8992-9abd600582b4 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Intercode: Stan- dardizing and benchmarking interactive coding with execution feedback
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 73350f3a-8230-4851-b9b8-997f8347744a · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Grounded copilot: How programmers interact with code-generating models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 540a8e5f-e9f5-443e-8ff2-e0bf5aaf39ad · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Bernstein, and Percy Liang
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a8cfb10c-ef32-4ff7-8e29-ddb62d0b92f4 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Ai-assisted code authoring at scale: Fine-tuning, deploying, and mixed methods evaluation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71b5cf91-4e19-456e-a9bb-d7c98a3a2b4a · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Reading between the lines: Modeling user behavior and costs in ai-assisted programming
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0e859cd5-d782-42cd-a18b-1f2bb70dc091 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild The productivity effects of generative ai: Evidence from a field experiment with github copilot
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ccc36f92-77ba-42e7-ab8e-2eb8904eee4f · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Need Help? Designing Proactive AI Assistants for Programming
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31933b58-dde6-4e68-8a3c-cf9bf601c591 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b9efc6b-d5d2-4904-884f-2b24fcda3f97 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Language Models for Code Completion: A Practical Evaluation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5e176e0-f468-4289-ba89-d9ffca8f0654 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild A Long Way to Go: Investigating Length Correlations in RLHF
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7022013-a03c-4689-89cb-c5c3b4b5b5c0 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bb660929-4a3a-47fd-a5df-dbd61cad2744 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild PSM presents the code context in the order of prefix and then suffix, using XML notation to demarcate prefix, suffix, and middle segments (e.g., <PREFIX> and </PREFIX>)
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 791a6138-a9d7-41d7-bdd7-cce2f8ebc414 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild SPM is identical to PSM except that the suffix appears before the prefix, which may be more natural than having the suffix appear directly before the output as is the case with PSM
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 59563f00-3eb9-4531-a2ea-6321800fa2af · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild sentinel
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9c32a364-eb9a-4742-b1de-23f3c3bfc799 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild pre-fill
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c358e7f6-d056-4f08-888d-1fc8afcec11a · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1ef2a7c9-b9d0-44fb-963c-75795b01b002 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2f161a4e-60b7-4ac0-a093-1ffe17e6017e · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f9adbe6a-8e84-4cf6-b88d-b811ccd129e9 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild clusters
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8895e605-430e-4142-b4ee-c77c5f9b66b9 · outbound
Copilot Arena: A Platform for Code LLM Evaluation in the Wild Unresolved cited work
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8e9c250b-57ee-442f-9b79-f5c9ba4c29c3 · inbound
Structure-Aware Fill-in-the-Middle Pretraining for Code Copilot Arena: A Platform for Code LLM Evaluation in the Wild
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc242de9-59c8-4e7b-b077-f0259d48190f · inbound
Mercury: Ultra-Fast Language Models Based on Diffusion Copilot Arena: A Platform for Code LLM Evaluation in the Wild
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 94623b0d-e1d2-4c2f-a239-9613c934811f · inbound
Nonparametric LLM Evaluation from Preference Data Copilot Arena: A Platform for Code LLM Evaluation in the Wild
Reference 2007
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08ada4f0-b026-4806-a4ad-e031b0d3f71e · inbound
Edit, But Verify: An Empirical Audit of Instructed Code-Editing Benchmarks Copilot Arena: A Platform for Code LLM Evaluation in the Wild
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 81ec95f4-fff1-4c73-89c7-d11343f5d94f · inbound
RECAP: An End-to-End Platform for Capturing, Replaying, and Analyzing AI-Assisted Programming Interactions Copilot Arena: A Platform for Code LLM Evaluation in the Wild
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.