Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:32:25.256651Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2505.12808.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:32:25.256651Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8fc25614-7b07-43a0-bbc7-f1662f1351b4 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23b0208d-af0f-4de3-8ff7-e9507dc15591 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models LLaMA: Open and Efficient Foundation Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fee0ad9-100f-42ca-a692-671d2671d18b · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26f4897c-4452-4277-bd1f-8225b80dbe0f · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models DeepSeek-V3 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 114d6de0-8431-47d4-befd-6f1f0380c54f · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cbe7aec-b82d-4d4f-ab35-85a34540fd7d · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models PIXIU: A Large Language Model, Instruction Data and Evaluation Benchmark for Finance
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfbc6f0e-3fbb-4084-8c6d-6dbf54fc112a · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Evaluating the Text-to-SQL Capabilities of Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0102a29a-5669-411c-a7e0-604232a25b3d · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33f72cf9-a311-4e6c-bc89-3957b12c5ffa · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Large language models for software engineering: A systematic literature review
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be6708b2-b7b7-42cd-95c8-7bd2c4b0dfb7 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Galactica: A Large Language Model for Science
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33b813c8-d48a-42d4-b5cf-555e8d17063c · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Large language models for automatic equation discovery of nonlinear dynamics
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation aea10b6d-7b96-4bf4-8077-f53d1b138dcd · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8296292-47de-4b7b-930e-bb2fc7ee466f · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Towards Understanding Sycophancy in Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c56887c-0935-4866-bb22-5c23604025f8 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models this is a problem, don’t you agree?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8198f62d-f6d2-4acd-a723-81696cd35afe · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 412a9677-1131-4737-9c8f-8c592c1a8dbb · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Alpacaeval: An automatic evaluator of instruction-following models, 2023
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f024736d-a5b7-471d-9a75-9bf5e2e8e26a · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 596fa2c5-dc32-486c-97ea-27e55c12d33e · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Llm evaluators recognize and favor their own generations
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b068fae-a0eb-45f8-9018-9d1e8a1797d5 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models The role of collective intelligence in crowdsourcing innovation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d07ecbcf-428f-43e4-ab6c-e84fccb3d834 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models The Wisdom of Crowds: Why the Many Are Smarter than the Few and How Collective Wisdom Shapes Business, Economies, Societies, and Nations
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a773fd94-506c-45eb-8645-70fda3e6b5b4 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Great Models Think Alike and this Undermines AI Oversight
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0f9eb2f-fffc-4fc5-84b3-296ba4f99c57 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e98473e-ef5f-4c12-860d-c77a831a3759 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models BERTScore: Evaluating Text Generation with BERT
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9b383d-ae84-4ae4-8ee2-4c2c25a3c1a0 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Bartscore: Evaluating generated text as text generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6416c483-82ea-4d8d-aff4-4bb57ee811c5 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Gpt-4o system card, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2314fd09-95b4-4283-8400-a71eece06574 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Llama 3 model card
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8bfcef6-1833-46bd-9068-a9860b1f7ada · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Sadler, Wei-Lun Chao, and Yu Su
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a062709f-9651-4ecc-8b22-0558fd042f98 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Openai o1 system card, 2024
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd3aa93c-eb57-49b0-b80a-9a1955847564 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89656eb2-221c-4656-95c1-d3d9a81b79bf · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef719b63-329c-478e-8cf2-b24ecacce57d · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Rethinking pragmatics in large language models: Towards open-ended evaluation and preference tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 71a8ac0a-de32-42cc-b92b-9244bb780efa · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 481cbd25-884f-475e-b9dc-9d7891821f05 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Generative Judge for Evaluating Alignment
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff496347-eb8e-41e4-b2b0-c892c5a0654b · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ae1978-d969-42d5-bda5-64247e092be0 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Verbosity bias in preference labeling by large language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a785a6d9-48f9-42ab-b81e-d2a55ad74bd1 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d778eae1-1c53-4f0d-8b73-9d6c5d8cf3aa · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 505c3919-5fee-4e00-88e1-cd116667de92 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models A bayesian approach towards crowdsourcing the truths from llms
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 6d3293d0-6dce-4158-8312-1edcc1736672 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models A multiagent approach for collective decision making in knowledge management
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c24b0dab-4135-4031-9fa6-7013030cbef2 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Swarm intelligence: A review of algorithms
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 07df0506-c6d7-4a3c-a262-c7128ca837c0 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Swarm creativity: Competitive advantage through collaborative innovation networks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c8468396-33e5-470b-86de-fe59d20a2917 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Binary search algorithm
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 721d1c75-8523-4a0e-a868-eb37fad72ceb · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models The proposed uscf rating system, its development, theory, and applications
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e780c37-4818-41de-8064-ef7a54fd1b3b · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Opencompass: A universal evaluation platform for foundation models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d86183e6-fb97-41a5-8fbd-836aa6f89cb2 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Patil, Ion Stoica, and Joseph E
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6694cabc-06d1-4ad7-a6bc-433369601a04 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Holistic Evaluation of Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1629666e-bc3e-414d-8c1c-94c61d4366cd · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 797968d4-e839-4d69-b486-843b1405cd9e · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c168072-8d54-4881-b6d9-30118242fd5f · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dad94738-cb5d-4c79-be22-4bff7999e8c2 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Open llm leaderboard v2
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ce7fc7-be5d-400a-9375-59c2bb7de809 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd1baaa6-6e8f-4848-833b-e3bd189361dd · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cdf3396-cfe5-47f9-a95a-d8892a89ec88 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Anchor points: Benchmarking models with much fewer examples
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d3bc8c80-03b7-43dd-a777-3c84aadff164 · outbound
Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Measuring Massive Multitask Language Understanding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.