Pith. sign in

Paper Citation Record · LEDGER

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models

As of 16 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2505.12808.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12808 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:32:25.256651Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved38
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8fc25614-7b07-43a0-bbc7-f1662f1351b4 · outbound

This paper cites GPT-4 Technical Report.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:24.988520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:24.988520Z digest=sha256:99342146939e7c0492b3e58fdb5d798774bb33ab2b09356eb7bf3d5941408549

Observation 23b0208d-af0f-4de3-8ff7-e9507dc15591 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:24.994060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:24.994060Z digest=sha256:e6b2a7b9274792de4eaf65c13cd82ad3f39b7194614992bc75b52770e2a3156c

Observation 0fee0ad9-100f-42ca-a692-671d2671d18b · outbound

This paper cites Qwen Technical Report.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:24.999687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:24.999687Z digest=sha256:e484d33790acf9de5d2e1d2472f5206098ee4d028b3ae570785842f36e103856

Observation 26f4897c-4452-4277-bd1f-8225b80dbe0f · outbound

This paper cites DeepSeek-V3 Technical Report.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models DeepSeek-V3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.005377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.005377Z digest=sha256:827e98aa7eab1db955380a40d7daffad9094368b3e1ff8c3f207066f4d833243

Observation 114d6de0-8431-47d4-befd-6f1f0380c54f · outbound

This paper cites Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.011656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.011656Z digest=sha256:3f6001b1a2d28a8388b13c40c516014f8650c0e5da42e0cc465c6fa141ae3ef0

Observation 4cbe7aec-b82d-4d4f-ab35-85a34540fd7d · outbound

This paper cites PIXIU: A Large Language Model, Instruction Data and Evaluation Benchmark for Finance.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models PIXIU: A Large Language Model, Instruction Data and Evaluation Benchmark for Finance

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.016672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.016672Z digest=sha256:7eb5fe745d0780e3507d784d0a2a651b1705026faefaf82f61f6d3cd67f93abe

Observation dfbc6f0e-3fbb-4084-8c6d-6dbf54fc112a · outbound

This paper cites Evaluating the Text-to-SQL Capabilities of Large Language Models.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Evaluating the Text-to-SQL Capabilities of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.022516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.022516Z digest=sha256:ca95db99b6fd79be5462d04de9c7f8b2f0406f071e26ff9f46309cebfcc4885d

Observation 0102a29a-5669-411c-a7e0-604232a25b3d · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.027531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.027531Z digest=sha256:bd12eed1b3d69e115f0725b194592d69ec3ef710f3d6e83d793ce3c599b817a5

Observation 33f72cf9-a311-4e6c-bc89-3957b12c5ffa · outbound

This paper cites Large language models for software engineering: A systematic literature review.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Large language models for software engineering: A systematic literature review

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.032353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.032353Z digest=sha256:ad2ee18b1673bd4545451c9e8cbd6f0d07572f00b2083e68f5bc03b50d396b3b

Observation be6708b2-b7b7-42cd-95c8-7bd2c4b0dfb7 · outbound

This paper cites Galactica: A Large Language Model for Science.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Galactica: A Large Language Model for Science

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.037139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.037139Z digest=sha256:db0b4220b7f3aa26ff14bc2330238091be45e2f6023366f727f2d05dd3e28e56

Observation 33b813c8-d48a-42d4-b5cf-555e8d17063c · outbound

This paper cites Large language models for automatic equation discovery of nonlinear dynamics.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Large language models for automatic equation discovery of nonlinear dynamics

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.995962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.042496Z digest=sha256:ec2d7b89fd305319ae679e9516753998d59230d212f02878f3834f0c5a5df0ae

Observation aea10b6d-7b96-4bf4-8077-f53d1b138dcd · outbound

This paper cites Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.046926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.046926Z digest=sha256:77a1750c3755f553b3cdcb06e71d4f620d656fbb86f572ebcd4124f26779d091

Observation c8296292-47de-4b7b-930e-bb2fc7ee466f · outbound

This paper cites Towards Understanding Sycophancy in Language Models.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Towards Understanding Sycophancy in Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.051364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.051364Z digest=sha256:a48bde181709f1a85736c676762a22eb567ec4df8234e6d5cf89bed283edd4ac

Observation 2c56887c-0935-4866-bb22-5c23604025f8 · outbound

This paper cites this is a problem, don’t you agree?.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models this is a problem, don’t you agree?

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.972157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.057147Z digest=sha256:d5ed047fe155eb36f4c1188a8b409f0e48ff08579c6b7d6b11076168ad722c98

Observation 8198f62d-f6d2-4acd-a723-81696cd35afe · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.062684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.062684Z digest=sha256:fafdb3bf234fb8dd1d1cc03b2f1de867588cc6f2fbf6500aeff2985009f84846

Observation 412a9677-1131-4737-9c8f-8c592c1a8dbb · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models, 2023.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Alpacaeval: An automatic evaluator of instruction-following models, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.067456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.067456Z digest=sha256:23cbe64fa64f8a3c43f3680960f2de79d5dc017a1471ba19157525ece7dd5061

Observation f024736d-a5b7-471d-9a75-9bf5e2e8e26a · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.072004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.072004Z digest=sha256:ce479dc9cc7cc7a838de73b9d310e79abc4528b35d027ebb44dfde28c4f85602

Observation 596fa2c5-dc32-486c-97ea-27e55c12d33e · outbound

This paper cites Llm evaluators recognize and favor their own generations.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Llm evaluators recognize and favor their own generations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.076932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.076932Z digest=sha256:ba780ed198da651ad4a10f701d4e507eeacf122f3a1f702e2bb9e599e8a08c86

Observation 3b068fae-a0eb-45f8-9018-9d1e8a1797d5 · outbound

This paper cites The role of collective intelligence in crowdsourcing innovation.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models The role of collective intelligence in crowdsourcing innovation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.930534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.081753Z digest=sha256:d506c1f8a2608d0e22dfc4a7d852e5ca74be4af7e0a68f783476ab790f6160b6

Observation d07ecbcf-428f-43e4-ab6c-e84fccb3d834 · outbound

This paper cites The Wisdom of Crowds: Why the Many Are Smarter than the Few and How Collective Wisdom Shapes Business, Economies, Societies, and Nations.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models The Wisdom of Crowds: Why the Many Are Smarter than the Few and How Collective Wisdom Shapes Business, Economies, Societies, and Nations

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.915592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.086345Z digest=sha256:e66bb253e83e10583177c835f35b7e67c4f123e2fc193ad3fb7564193c44bc14

Observation a773fd94-506c-45eb-8645-70fda3e6b5b4 · outbound

This paper cites Great Models Think Alike and this Undermines AI Oversight.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Great Models Think Alike and this Undermines AI Oversight

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.091216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.091216Z digest=sha256:9a75bccc87a34c2b5be8afbccadd60c3924e1e73a999171c635c6d203615430e

Observation b0f9eb2f-fffc-4fc5-84b3-296ba4f99c57 · outbound

This paper cites MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models MoverScore: Text Generation Evaluating with Contextualized Embeddings and Earth Mover Distance

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.096264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.096264Z digest=sha256:c5f5d00bc3764333ebf15c0ca53628a9cba8c9a58c2f014faca9a61ff8ce483c

Observation 5e98473e-ef5f-4c12-860d-c77a831a3759 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models BERTScore: Evaluating Text Generation with BERT

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.101473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.101473Z digest=sha256:7ea49caab7294817386e4f86eaedd923eea34c09b7957d23b1666f684bea82a2

Observation 5b9b383d-ae84-4ae4-8ee2-4c2c25a3c1a0 · outbound

This paper cites Bartscore: Evaluating generated text as text generation.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Bartscore: Evaluating generated text as text generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.900680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.106444Z digest=sha256:da03bca54148a72131455f771e3561f8966d65c7662d4da8cb20a39786da6aac

Observation 6416c483-82ea-4d8d-aff4-4bb57ee811c5 · outbound

This paper cites Gpt-4o system card, 2024.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Gpt-4o system card, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.111289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.111289Z digest=sha256:4f9cd90a0ffb3307c00438bc0348edf19c7729bcb81d30ac6e0195eac308e20a

Observation 2314fd09-95b4-4283-8400-a71eece06574 · outbound

This paper cites Llama 3 model card.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Llama 3 model card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.115818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.115818Z digest=sha256:8575150f68efe636d8f8d2ea94890c54a738e7a4d04efe4ea82443bbb2d97e9e

Observation e8bfcef6-1833-46bd-9068-a9860b1f7ada · outbound

This paper cites Sadler, Wei-Lun Chao, and Yu Su.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Sadler, Wei-Lun Chao, and Yu Su

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.867989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.121364Z digest=sha256:d21e08070570b55e8b30e95abebbe1e150b5edbfa87361987ce2b94d136cfecd

Observation a062709f-9651-4ecc-8b22-0558fd042f98 · outbound

This paper cites Openai o1 system card, 2024.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Openai o1 system card, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.126903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.126903Z digest=sha256:8f942368c67d268c2da2f03112fea5b49743c17841d09df2a371ade04ea8484c

Observation fd3aa93c-eb57-49b0-b80a-9a1955847564 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.132474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.132474Z digest=sha256:091f3d0300a63d968419d6b8d1209be84f68a630e9eb3c071006479f63d53900

Observation 89656eb2-221c-4656-95c1-d3d9a81b79bf · outbound

This paper cites an unresolved cited work.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.138585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.138585Z digest=sha256:012099454630de6ec947a5a5986f107185e51cdf6ce42d5d8ea9f63cbd1ecb40

Observation ef719b63-329c-478e-8cf2-b24ecacce57d · outbound

This paper cites Rethinking pragmatics in large language models: Towards open-ended evaluation and preference tuning.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Rethinking pragmatics in large language models: Towards open-ended evaluation and preference tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.834297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.143825Z digest=sha256:872c6d2b8f77d1f2227b4da5ac57332694f9a8b90ffed23a77bf1765ffc2c524

Observation 71a8ac0a-de32-42cc-b92b-9244bb780efa · outbound

This paper cites NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.819719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.148461Z digest=sha256:d399c60b178a48c24ea15ff37a5e77e56d317b89de8d2e3e542bfaef3bb7c48a

Observation 481cbd25-884f-475e-b9dc-9d7891821f05 · outbound

This paper cites Generative Judge for Evaluating Alignment.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Generative Judge for Evaluating Alignment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.153102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.153102Z digest=sha256:44cb82431eac43a3af9f8cf6f0e29379fde6eff99d2cfc1076b3c2c365169865

Observation ff496347-eb8e-41e4-b2b0-c892c5a0654b · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.158010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.158010Z digest=sha256:0f9acb6b8fcdd8ec1171b00ac2726b57a5e1b16b443e4640af51d02a79d88411

Observation a4ae1978-d969-42d5-bda5-64247e092be0 · outbound

This paper cites Verbosity bias in preference labeling by large language models.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Verbosity bias in preference labeling by large language models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.804847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.163164Z digest=sha256:8ab331ad35748bec7dca4adf018007df13fc9fc96b95ac746b82b84c8519a427

Observation a785a6d9-48f9-42ab-b81e-d2a55ad74bd1 · outbound

This paper cites PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.167955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.167955Z digest=sha256:01220682bcbd1a495c7dd4dee8bca3b67d507bd801c6afbbe7ecb04b41d6452b

Observation d778eae1-1c53-4f0d-8b73-9d6c5d8cf3aa · outbound

This paper cites Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.173072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.173072Z digest=sha256:79a8ec61f6a5c2f3106993ca6557cfb5713f1a61fcbc18195749b80926028137

Observation 505c3919-5fee-4e00-88e1-cd116667de92 · outbound

This paper cites A bayesian approach towards crowdsourcing the truths from llms.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models A bayesian approach towards crowdsourcing the truths from llms

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.788323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.177851Z digest=sha256:e26b2bfdd86488df4a1e7164f965ba38276e25ad38ab96412fac8bebad3947b8

Observation 6d3293d0-6dce-4158-8312-1edcc1736672 · outbound

This paper cites A multiagent approach for collective decision making in knowledge management.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models A multiagent approach for collective decision making in knowledge management

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.773908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.182847Z digest=sha256:ccd32867165d006e1f71746b9f852ca24439489eadeb3a7a9647297a8640c240

Observation c24b0dab-4135-4031-9fa6-7013030cbef2 · outbound

This paper cites Swarm intelligence: A review of algorithms.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Swarm intelligence: A review of algorithms

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.759718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.187420Z digest=sha256:da5afd90c1632e86683255719dff7464cca68bf216561bd794b39cd686eb8ed4

Observation 07df0506-c6d7-4a3c-a262-c7128ca837c0 · outbound

This paper cites Swarm creativity: Competitive advantage through collaborative innovation networks.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Swarm creativity: Competitive advantage through collaborative innovation networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.745076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.192162Z digest=sha256:637d6c13890a049cbf7e6bcb902589b6a40f085eb5da5e2661621ab1c9296d02

Observation c8468396-33e5-470b-86de-fe59d20a2917 · outbound

This paper cites Binary search algorithm.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Binary search algorithm

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.730422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.196739Z digest=sha256:c01d1287d5f57214a92752ec1b98ca737ebe1a424752f99d3402d52db6ba2d9b

Observation 721d1c75-8523-4a0e-a868-eb37fad72ceb · outbound

This paper cites The proposed uscf rating system, its development, theory, and applications.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models The proposed uscf rating system, its development, theory, and applications

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.201472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.201472Z digest=sha256:4f94823fdfd47bfb901afd4c1b83e5a84c9a20f3bc841ca5b387fa6c7f2c0b53

Observation 6e780c37-4818-41de-8064-ef7a54fd1b3b · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Opencompass: A universal evaluation platform for foundation models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.206369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.206369Z digest=sha256:cfcb359b89b0f472559899f38ff796922bcb7f617ac9ef6cc890c89bc5299e44

Observation d86183e6-fb97-41a5-8fbd-836aa6f89cb2 · outbound

This paper cites Patil, Ion Stoica, and Joseph E.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Patil, Ion Stoica, and Joseph E

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.211061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.211061Z digest=sha256:e0f6c38727cd747c8d29a92c43c1e1aba16a9476432fecdbbfd756baeb16260e

Observation 6694cabc-06d1-4ad7-a6bc-433369601a04 · outbound

This paper cites Holistic Evaluation of Language Models.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Holistic Evaluation of Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.215708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.215708Z digest=sha256:431cbfe97e00b15ff9c185d426cdec81c050de0c377833628d49d3d83e22ee2c

Observation 1629666e-bc3e-414d-8c1c-94c61d4366cd · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.220753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.220753Z digest=sha256:dea5cdd3ce0395ec5b6c4169420e217d4ccbb6ba6cf88bff824543e09488eee4

Observation 797968d4-e839-4d69-b486-843b1405cd9e · outbound

This paper cites EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.225734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.225734Z digest=sha256:954b7052204e6c9f34203f76361d8ff802125b6e1b17e01b32614957c0e38e30

Observation 5c168072-8d54-4881-b6d9-30118242fd5f · outbound

This paper cites MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models MixEval: Deriving Wisdom of the Crowd from LLM Benchmark Mixtures

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.231127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.231127Z digest=sha256:c9b89096c76478d6ffc41c12cbfc0d16cd565e583330193e727d862bb6ef5441

Observation dad94738-cb5d-4c79-be22-4bff7999e8c2 · outbound

This paper cites Open llm leaderboard v2.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Open llm leaderboard v2

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.236356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.236356Z digest=sha256:679ab78c7f8eef3d208201cadc1a04f5a0ef45db45f4ca508495b25cecb72a23

Observation 67ce7fc7-be5d-400a-9375-59c2bb7de809 · outbound

This paper cites The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.241466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.241466Z digest=sha256:adc1278206fae757b518ee5c80c1e48ab4e6901b4091e10e3d913e5ba959391b

Observation cd1baaa6-6e8f-4848-833b-e3bd189361dd · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.246442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.246442Z digest=sha256:2f7bead9a81e12d619d8558f6bd743bb6a7ba114f8e354c521d131b27eee0558

Observation 7cdf3396-cfe5-47f9-a95a-d8892a89ec88 · outbound

This paper cites Anchor points: Benchmarking models with much fewer examples.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Anchor points: Benchmarking models with much fewer examples

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:32:25.677825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T20:32:25.252100Z digest=sha256:22f4d7c75900038cbd704987d86bae239917c0e1b4c8b850ca8c8ebfd54680ea

Observation d3bc8c80-03b7-43dd-a777-3c84aadff164 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models Measuring Massive Multitask Language Understanding

Reference 54

Resolution
malformed identifier
no resolver link, observed 2026-08-15T20:32:25.256651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.256651Z digest=sha256:4d9be0493e3a99231ca89c6a26d0112d35c56e295928af04b31a32e8333e019f

Pith citing papers

No inbound Pith citation observations are available.