Pith. sign in

Paper Citation Record · LEDGER

Correlated Errors in Large Language Models

As of 14 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 15 inbound Pith citation observations for arXiv:2506.07962.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07962 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:27:56.382686Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:41:10.536734Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T14:19:53.634812Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved28
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70a57325-53df-4dd0-a6cd-dbb5bec75018 · outbound

This paper cites write newline.

Correlated Errors in Large Language Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.202558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.202558Z digest=sha256:cdb923893590848aeecd562035ac2f7b1361ffe89af20243deb7313976c2fd42

Observation f42b330d-70cf-4f50-a1de-42f09bd3c4e3 · outbound

This paper cites F., Joachims, T., and Antonio, A.

Correlated Errors in Large Language Models F., Joachims, T., and Antonio, A

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:57.079707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.207806Z digest=sha256:35b3d178e8c8a984be31454ef60705043db3c602a70dffefd6a050b0a493b4d1

Observation f8d37fdb-3653-4c1c-9d65-4aa8d1dd1156 · outbound

This paper cites Upwork job postings dataset 2024 (50k records), 2024.

Correlated Errors in Large Language Models Upwork job postings dataset 2024 (50k records), 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:57.069391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.212724Z digest=sha256:31cf431491a6556ed60c998de3ecc81511e89fbeb8c4e7b1ff8e1232e14df90a

Observation b2e59c5d-8253-4039-b930-1f5791bef919 · outbound

This paper cites Hiring under congestion and algorithmic monoculture: Value of strategic behavior.

Correlated Errors in Large Language Models Hiring under congestion and algorithmic monoculture: Value of strategic behavior

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.216131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.216131Z digest=sha256:68681fdeedd959cd9578449e705dbe244225969289a6731859fb4ab1ea529227

Observation d2250d20-a2f7-40af-baa6-fe93469b7f8b · outbound

This paper cites Resume dataset, 2022.

Correlated Errors in Large Language Models Resume dataset, 2022

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:57.058432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.219550Z digest=sha256:346cce640da4bea4d0d98dfeeebb87d89e9da0ecb1d9fa05f71518b691891afa

Observation 8e420c71-310f-4019-b003-0f8d47b39abb · outbound

This paper cites A., Kumar, A., Jurafsky, D., and Liang, P.

Correlated Errors in Large Language Models A., Kumar, A., Jurafsky, D., and Liang, P

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:57.047942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.222764Z digest=sha256:e6fd14b7935310522ddbf703a0bbd8f37d14513a659c23bfa4f7b75a7f62ebfd

Observation 022660ef-983e-4b6d-a0aa-794b58273bab · outbound

This paper cites Ecosystem Graphs: The Social Footprint of Foundation Models.

Correlated Errors in Large Language Models Ecosystem Graphs: The Social Footprint of Foundation Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.226762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.226762Z digest=sha256:c65e24ae382f3b9f0343a5250439ec061a33c994e0de57eded5201d6f67acbfc

Observation 6a7e76d8-601d-4646-944c-957b58e7d4ab · outbound

This paper cites Harnessing Multiple Large Language Models: A Survey on LLM Ensemble.

Correlated Errors in Large Language Models Harnessing Multiple Large Language Models: A Survey on LLM Ensemble

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.230771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.230771Z digest=sha256:4eeb5a0bfe7474b1ef30d040f77bc193357def0a2d14b9077b0c738c0c3014cf

Observation 45b95af4-2e89-43c8-af15-b2395f9da3a0 · outbound

This paper cites and Hellman, D.

Correlated Errors in Large Language Models and Hellman, D

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:57.038631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.234413Z digest=sha256:93237a8c13938079ede8611a5141539e2272229e575c708a9c2e049fb3a36c75

Observation e32eff6d-d795-4c47-b006-1a8b65830820 · outbound

This paper cites Auditing the Use of Language Models to Guide Hiring Decisions.

Correlated Errors in Large Language Models Auditing the Use of Language Models to Guide Hiring Decisions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.237670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.237670Z digest=sha256:6145b0cc9f38b26eefe23b1a2679ae6cc2e22c8ee3329d84bddc72f942db5658

Observation fe05791d-4e96-4706-8420-c4f3e199ef7d · outbound

This paper cites an unresolved cited work.

Correlated Errors in Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:27:57.028708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.241343Z digest=sha256:b6f53f6e30927a7a661591b05fd57bc3468883f5b17b94e5496fcddbffa1b9a8

Observation 7035dbe6-4d82-4cdc-adf5-bbb23b52220d · outbound

This paper cites Great Models Think Alike and this Undermines AI Oversight.

Correlated Errors in Large Language Models Great Models Think Alike and this Undermines AI Oversight

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.244885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.244885Z digest=sha256:3e8f82a469571d50304e8a5f6b844728b6f33347e2f5b238cb8bbfc419084d49

Observation 8434a646-576b-4526-911c-365e727e2c73 · outbound

This paper cites Auditing work: Exploring the new york city algorithmic bias audit regime.

Correlated Errors in Large Language Models Auditing work: Exploring the new york city algorithmic bias audit regime

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:57.018976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.248448Z digest=sha256:03427de69a4cad76e19df191c3425342f7fa6b10b5e620a02428d0338b628624

Observation dc3f0758-e6dd-4df5-b157-ba2f84a2d21b · outbound

This paper cites A Survey on LLM-as-a-Judge.

Correlated Errors in Large Language Models A Survey on LLM-as-a-Judge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.251435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.251435Z digest=sha256:2f8fd4007f337bf6cd3c298db35e1840de3ad9bb0d52f4c38452c7ca6998fa59

Observation 582b507b-e467-49f8-a42b-b2a6307aa3fd · outbound

This paper cites S., and Chouldechova, A.

Correlated Errors in Large Language Models S., and Chouldechova, A

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.254808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.254808Z digest=sha256:a88d2a24eb3f31e5004ec26bb6a534f983d8f52a06a3c892cc2eb4e98aa1eeae

Observation eac72792-1661-4081-bbda-9a1496e4bb85 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Correlated Errors in Large Language Models Measuring Massive Multitask Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.257636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.257636Z digest=sha256:6fcd492eeeddbe49787d608f19ad6cd040bd0a621202f820f7001a7230a2a253

Observation 7d677d86-1727-4533-a373-28bcdc90f35a · outbound

This paper cites Scarce Resource Allocations That Rely On Machine Learning Should Be Randomized.

Correlated Errors in Large Language Models Scarce Resource Allocations That Rely On Machine Learning Should Be Randomized

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:27:56.637614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.260834Z digest=sha256:575e28f8c6e697ea266add3a313e97e49dfe88cb858e3b0e669a997072ec0adb

Observation 649cfe73-94ef-493a-9b41-62aded8ddbb8 · outbound

This paper cites Algorithmic pluralism: A structural approach to equal opportunity.

Correlated Errors in Large Language Models Algorithmic pluralism: A structural approach to equal opportunity

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:57.008360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.264090Z digest=sha256:2f0c81f3058e88995e506b9b0b0b0d1f00f93e4e9f89f780d8fcc47fe92ec71c

Observation af230139-69bc-4bd1-88f3-9d7f8b67293d · outbound

This paper cites an unresolved cited work.

Correlated Errors in Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.268441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.268441Z digest=sha256:8777e1f35669580dfe129389f72d08695b4c20b1a86712787fb8ad4e86a8a938

Observation 751fd996-f515-4153-adff-64cee291e63b · outbound

This paper cites an unresolved cited work.

Correlated Errors in Large Language Models Unresolved cited work

Reference 20

Resolution
verified exact
doi, observed 2026-08-07T05:27:56.424058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.271345Z digest=sha256:28722af4b1889b8bf1eae66c0a572e349ad83b7065ad8c0fd0627f93577289bc

Observation 2fd09eab-9413-4b59-9a9c-dec5b828472e · outbound

This paper cites an unresolved cited work.

Correlated Errors in Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:27:56.994718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.274353Z digest=sha256:036b3ba6c3ef389efef0976125827dd9c608ac3deb9f711bcf9889f21d35cb4d

Observation 38d88977-e760-4732-a469-bc4660b6e058 · outbound

This paper cites When can llms actually correct their own mistakes? a critical survey of self-correction of llms.

Correlated Errors in Large Language Models When can llms actually correct their own mistakes? a critical survey of self-correction of llms

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.985138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.277312Z digest=sha256:97acbec7e14a354c0fdc4214ab70c3a34a8d590600916c39947512df527ac859

Observation 16366c34-8406-4acc-a032-c6db38377b29 · outbound

This paper cites and Raghavan, M.

Correlated Errors in Large Language Models and Raghavan, M

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.974785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.280327Z digest=sha256:82d9d8130950ccafdf4a29357f634a0c93f4703c1a5eee7dddf9c06cada5309a

Observation 277136a5-c2f7-486c-a0b3-14418df016f9 · outbound

This paper cites A., Manning, C.

Correlated Errors in Large Language Models A., Manning, C

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.963191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.284616Z digest=sha256:8176b865e399d465ef43349e1050e6da22e979ba212fea0eef3883a0580869ec

Observation a0f1ea8a-1218-407d-a4b4-559a9f278762 · outbound

This paper cites Sparse Autoencoders for Hypothesis Generation.

Correlated Errors in Large Language Models Sparse Autoencoders for Hypothesis Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.287837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.287837Z digest=sha256:64bd36cae5f76416a02133491b151d870bae5d1f65436de6252e49d0398f631d

Observation 006fe020-d1b6-4047-8a65-87d947f2ddd0 · outbound

This paper cites Llm evaluators recognize and favor their own generations.

Correlated Errors in Large Language Models Llm evaluators recognize and favor their own generations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.291205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.291205Z digest=sha256:f204a460ad2d9e2b4f370acf1ce0255a601c7ab10ea0dd8807d957bb8f9df028

Observation ced26615-ae60-4926-9014-d89f891bd0db · outbound

This paper cites and Garg, N.

Correlated Errors in Large Language Models and Garg, N

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.946811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.295361Z digest=sha256:a5df1d223dcdbb430989c472a65048ce31b2509b353eabf25438b9284a83660d

Observation 30147cd4-942f-4086-b0ba-0c00678d00b1 · outbound

This paper cites and Garg, N.

Correlated Errors in Large Language Models and Garg, N

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.936226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.300314Z digest=sha256:018d5ebbc22a8416d7b1afc734d72c841a29eb57c32cc1657767ca20bfb363de

Observation 3cbd99cd-0b9c-43d6-ab3e-4883f0c23ba4 · outbound

This paper cites Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench.

Correlated Errors in Large Language Models Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.303191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.303191Z digest=sha256:92b2a160bbeba1f7e0d3acac4eecd67d4ec336d508932cf832096cbfe64f9024

Observation 87583b3c-c713-45e3-add1-2cd721cd6157 · outbound

This paper cites Competition and Diversity in Generative AI.

Correlated Errors in Large Language Models Competition and Diversity in Generative AI

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.306759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.306759Z digest=sha256:e3dbe284d10e0f6046467384b7025d6a02a18f196255f650c081c4842d8ccf70

Observation e05a9890-8eb3-4c57-818d-d7e9bc07600c · outbound

This paper cites K., Ahmad, A., Andreetto, M., Prabhakaran, V., Prabhu, U., Dieng, A.

Correlated Errors in Large Language Models K., Ahmad, A., Andreetto, M., Prabhakaran, V., Prabhu, U., Dieng, A

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.925482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.310626Z digest=sha256:50e76c515cbd0f4cb014c47aefa068d92069366148d8c16047525dc217fc4256

Observation eef06558-03a2-436e-b2e5-15ecc7d920cf · outbound

This paper cites Large language models are inconsistent and biased evaluators.

Correlated Errors in Large Language Models Large language models are inconsistent and biased evaluators

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.913839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.313627Z digest=sha256:b25688df712792de10006a1bab3c0eb93a8e355173d1d4aac00a5dd81b159e4b

Observation 40175e95-6296-4094-95e5-4c8f9d32365e · outbound

This paper cites LLM-TOPLA: Efficient LLM Ensemble by Maximising Diversity.

Correlated Errors in Large Language Models LLM-TOPLA: Efficient LLM Ensemble by Maximising Diversity

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T05:27:56.510635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.316625Z digest=sha256:cffe95e70fcfde4ee09c005cb82f026a2a2448743f74c0244da30ca9504c6198

Observation bee61e58-541e-4645-89da-09910372b29d · outbound

This paper cites Law and the emerging political economy of algorithmic audits.

Correlated Errors in Large Language Models Law and the emerging political economy of algorithmic audits

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.902482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.319952Z digest=sha256:ed176d518fc06b134ecfcf8f5365a7a9b326ab65c689fc681cad891f7bdfffa1

Observation b7020430-28b7-47d5-9e35-4f0e0d37e87f · outbound

This paper cites an unresolved cited work.

Correlated Errors in Large Language Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:27:56.892385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.323872Z digest=sha256:218ca4b4eb1027a6cf66f03bccdfedb21cc2fb0b17f8a65202930eb8252b4aab

Observation b936e3d5-7f47-4c95-86be-8945b4b3d1ff · outbound

This paper cites Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models.

Correlated Errors in Large Language Models Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.327123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.327123Z digest=sha256:002ad13c3e1b33e9ba2535226d662f7813a887013588ec94a65393b8b911e8d3

Observation dc84554f-f5ac-46c7-9ab3-8b4ce16d6715 · outbound

This paper cites Evaluating Generative AI Systems is a Social Science Measurement Challenge.

Correlated Errors in Large Language Models Evaluating Generative AI Systems is a Social Science Measurement Challenge

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.330863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.330863Z digest=sha256:e8a748d930ccd67ff8af9291da42c2151044ab7d2f5d312a270afe8d2192a52c

Observation 5a80a956-6ae0-4b42-8e3b-5e1161634dbc · outbound

This paper cites an unresolved cited work.

Correlated Errors in Large Language Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:27:56.881442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.334387Z digest=sha256:672d32636760add2cd19eabee962708fdc441e38eb5c847d884d0cb35ead5808

Observation db25db59-1a5d-4e88-94a0-fb21fb7f69a0 · outbound

This paper cites Self-Preference Bias in LLM-as-a-Judge.

Correlated Errors in Large Language Models Self-Preference Bias in LLM-as-a-Judge

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.337602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.337602Z digest=sha256:a84eada38b0229d94ce6c0859e8e4c067bb4dec957706324db08790c00c025fb

Observation b2f468dc-1180-48b3-b383-c073f710ffd0 · outbound

This paper cites Toward an Evaluation Science for Generative AI Systems.

Correlated Errors in Large Language Models Toward an Evaluation Science for Generative AI Systems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.341173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.341173Z digest=sha256:9a6d3c2dc4f21c6de17eb853bffec4ba375a7c9bfb4ed58ca942dfe73feca9f0

Observation 251a79ab-f120-4b0a-ae24-04ac8667f492 · outbound

This paper cites We're Different, We're the Same: Creative Homogeneity Across LLMs.

Correlated Errors in Large Language Models We're Different, We're the Same: Creative Homogeneity Across LLMs

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.344423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.344423Z digest=sha256:ca8848beb4b0d600493ab93e834fe4929578c5d2ea7c02ea6a88792b384b6fde

Observation 4a2f3f31-dadd-4c62-867f-9f660d7143aa · outbound

This paper cites and Horton, J.

Correlated Errors in Large Language Models and Horton, J

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.870269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.348112Z digest=sha256:0cacedbc1eeeabcf995cc69dc2d7cca7d391105fd5a99c734c102ce1a4e1fbe0

Observation 9f72f5cb-5757-402e-9420-e5c35fa74b10 · outbound

This paper cites Algorithmic writing assistance on jobseekers’ resumes increases hires.

Correlated Errors in Large Language Models Algorithmic writing assistance on jobseekers’ resumes increases hires

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.858998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.351450Z digest=sha256:140f18901d5bef814eddd85feb0ad4b70df3c05d5a5de02db80fb3663856611a

Observation 1478fcfd-d65c-4739-b5d3-f968f268e33b · outbound

This paper cites and Caliskan, A.

Correlated Errors in Large Language Models and Caliskan, A

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.848456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.354598Z digest=sha256:1784afc2ab0fd731b02f45a59578377ce28aafa92c916a6b20c00100d2105c93

Observation 69e40af7-a9f7-4acf-be9e-5f8a339798a2 · outbound

This paper cites Layers at Similar Depths Generate Similar Activations Across LLM Architectures.

Correlated Errors in Large Language Models Layers at Similar Depths Generate Similar Activations Across LLM Architectures

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.357648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.357648Z digest=sha256:65af4f9cbee137b536a6760ef283502ed0580af35b1e356661f750b58200622b

Observation 7a0081b0-6a5b-47dc-8890-c1ba9d1d14c0 · outbound

This paper cites M., Vecchione, B., Qu, T., Cai, P., Smith, A., Investigators, C.

Correlated Errors in Large Language Models M., Vecchione, B., Qu, T., Cai, P., Smith, A., Investigators, C

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.837755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.360917Z digest=sha256:4b38828715bda620dbdc2bcea3f1f4d97fc755c6dab545a3d0896488189dd1a1

Observation c84cdb38-16d1-4f71-b58a-96ca1133c6a1 · outbound

This paper cites Generative Monoculture in Large Language Models.

Correlated Errors in Large Language Models Generative Monoculture in Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.364180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.364180Z digest=sha256:d5aeb2093e2ab112890c33d607ec755d342c94786267e2a65a89c3dacb9d0efa

Observation 03dfceba-39f6-45f9-b77c-728a5f9d3d4f · outbound

This paper cites Echoes in AI: Quantifying lack of plot diversity in LLM outputs.

Correlated Errors in Large Language Models Echoes in AI: Quantifying lack of plot diversity in LLM outputs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.367802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.367802Z digest=sha256:54da9d32636f2e1592c7224fd76d7a30a22ec680c19bafd6c44cd5267dbb4dec

Observation d55c4435-c5c4-4174-a6df-b755aeee04e4 · outbound

This paper cites One llm is not enough: Harnessing the power of ensemble learning for medical question answering.

Correlated Errors in Large Language Models One llm is not enough: Harnessing the power of ensemble learning for medical question answering

Reference 49

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T05:27:56.413896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.371877Z digest=sha256:f1f5e4d83c39e5551ccfa7e1579338bb1953a9c242befbc799e5fbeabd00c2fc

Observation e57b3245-d5b6-4f7a-9906-30b007e8cad6 · outbound

This paper cites Generative ai meets open-ended survey responses: Research participant use of ai and homogenization.

Correlated Errors in Large Language Models Generative ai meets open-ended survey responses: Research participant use of ai and homogenization

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:27:56.827566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-07T05:27:56.375374Z digest=sha256:b0e4b959f9d5a9bea3e7b755a0fd464c3822420a7db7cb93fa09d90be4ab7322

Observation 29c19df9-c8ac-43d6-972e-ec5956d452ef · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Correlated Errors in Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.379434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.379434Z digest=sha256:9c2d4d510a39027e1e76a7e94d9df74aea60dd707cf5d51369b29f76f988596d

Observation 1551c4a8-e1f1-4d07-aca3-c1ceeaca247c · outbound

This paper cites Hypothesis Generation with Large Language Models.

Correlated Errors in Large Language Models Hypothesis Generation with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:56.382686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:27:56.382686Z digest=sha256:394c01999e1c6968ceaa1ea4684ef33be784e45cf5ac07dd9f7170bd02548ba2

Pith citing papers

Observation 5e4a5119-ddef-4e4f-a568-ed6f770dd35a · inbound

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework cites this paper.

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework Correlated Errors in Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:36:24.964235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T13:34:22.790199Z digest=sha256:008b86bc68c7040572ee3a4f28981fb982aa1eba7271fb29c4e73ff779657865

Observation 12367219-7353-4362-b60b-9493e7ab957b · inbound

DIANOIA: Diagnostic Decomposition and Joint Optimization for Multi-Agent Reasoning cites this paper.

DIANOIA: Diagnostic Decomposition and Joint Optimization for Multi-Agent Reasoning Correlated Errors in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:22:16.274146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:22:16.274146Z digest=sha256:b01cb7250698e49c67c2d362f4dd2eee93920f786fb6e3d8ceb2fcb9d48849cd

Observation 4fc8be66-188a-43fb-8740-8d82a9ad6b52 · inbound

Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact cites this paper.

Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact Correlated Errors in Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:46:29.460161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T18:44:17.986764Z digest=sha256:5df53dda4e1a293a7ece50b30293696e0ef31be89589cdff91a10252eb26792e

Observation 3e1c16eb-842f-40c8-95ee-bfe2c8f1cc13 · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Correlated Errors in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:24:21.221003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T10:24:16.283375Z digest=sha256:63639d8b028ad832cd7b1a5459f94cee64d242ad8be2a1bccb13759aa6e7f299

Observation ec910fde-016d-4960-b8be-d5cd41527b5b · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Correlated Errors in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:50.086527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-11T00:47:51.440441Z digest=sha256:0fd89224183497d8214f6a239aeb2b4fa1f30dc97cca18f92aa80720048c4b97

Observation 176d43db-5200-44b4-8040-58859fcc2a1b · inbound

LACE: Lattice Attention for Cross-thread Exploration cites this paper.

LACE: Lattice Attention for Cross-thread Exploration Correlated Errors in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:28.500762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T04:05:35.063305Z digest=sha256:6a2b45f1477d31b651620a61e1d126f1f804d4495e7ff87124db6bbb19cb4296

Observation 79dcebf4-a62d-460d-9652-0f3efaf5dd92 · inbound

Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery cites this paper.

Refute-or-Promote: An Adversarial Stage-Gated Multi-Agent Review Methodology for High-Precision LLM-Assisted Defect Discovery Correlated Errors in Large Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:14:08.026843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T03:11:26.513954Z digest=sha256:3409b30a946c7c50ad622eae536ea6a7a1b359f04e327458d2b07d311466b1e0

Observation a76ebf2e-c20e-41a1-81c3-8ff71c9a67d9 · inbound

Pramana: A Protocol-Layer Treatment of Claim Verification in Autonomous Agent Networks cites this paper.

Pramana: A Protocol-Layer Treatment of Claim Verification in Autonomous Agent Networks Correlated Errors in Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:13:56.536389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T02:10:13.030288Z digest=sha256:e1ff37993cd6bb7f2451dd3c08e0df7f4acb7548598cab972bbac0091362267c

Observation dbe96fb3-cf0f-4f13-a4ac-9e24e2f22c21 · inbound

Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution cites this paper.

Design and Evaluation of Multi-Agent AI Oracle Systems for Prediction Market Resolution Correlated Errors in Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:36:11.302923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T20:59:43.961420Z digest=sha256:6a0390a5011fc90164d236a5dac3c796bdb85fca40ed66d6e9504de49ccaf351

Observation 2916fca7-771a-46ac-8b79-9ae6046aec74 · inbound

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation cites this paper.

Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling Estimation Correlated Errors in Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.599963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T22:48:20.193699Z digest=sha256:a43e002fee03077f3b5b9f25b373e8eb39d0e702c18094a404f06fb0c8caff70

Observation 93dd0bbe-ba52-4994-86c4-7581c0363fd6 · inbound

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models cites this paper.

When Does Combining Language Models Help? A Co-Failure Ceiling on Routing, Voting, and Mixture-of-Agents Across 67 Frontier Models Correlated Errors in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T14:19:53.636552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T04:20:02.722108Z digest=sha256:7344f097adfe6d8e1d7fb8ec2f6a1a25e66c48fe56641fb175633a51802c116b

Observation 714ea577-33f7-4ea4-83b1-4ef9e52358e7 · inbound

How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm cites this paper.

How Much Does Correctness Cost? Budgeted Placement of Strong Correctors in a Weak Multi-Agent Swarm Correlated Errors in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T16:11:18.394700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:11:18.394700Z digest=sha256:47c4753f37ed1a43f46fb6c265925b3fdf04b30f3d8e7ce2f57bea2f88d1ebd3

Observation 68444501-b105-4269-9ecc-3c00fdd0f4cb · inbound

PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection cites this paper.

PlanFlip: Attacking Multi-Agent LLM Systems via Planning-Phase Prompt Injection Correlated Errors in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T15:17:29.473018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T15:17:29.473018Z digest=sha256:6da5043df366d0278d61d23331dd1239fe5b46fb69b51230d91eb78cc33ac4b9

Observation 8a59f7c6-e22a-4bfa-a429-5d9b4d3af1aa · inbound

Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles cites this paper.

Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles Correlated Errors in Large Language Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T09:29:49.714215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:29:49.714215Z digest=sha256:3c70e2434f2b25349bcc8ddab9244dca6fe0c444747fd7fd406092ae7163bf6b

Observation 57ac367a-af4a-42e3-9ecb-ee8b0b44a388 · inbound

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees cites this paper.

Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees Correlated Errors in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:41:10.536734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:41:10.536734Z digest=sha256:2ab1e3d4296db6fc13801662ddde08f69b6d724bdd083225c860262df2bd8421