Pith. sign in

Paper Citation Record · LEDGER

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation

As of 18 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2411.13212.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13212 v3

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:45:16.249314Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 719a5616-f964-4136-bae7-31952d4b207c · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.172964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.172964Z digest=sha256:cb90750682b8177451a0d2dacf1183bd388b151ee49d7aa25c6a1926c147f65c

Observation 23c5bd0d-aa97-4efe-8f17-2fa2e8b26ebb · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.181554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.181554Z digest=sha256:4482cd8ceac534481727af1c274d406ee405dafff8dcecf3a847c76b7a832316

Observation 11a01bd5-e960-49ab-af5e-b27b3506cbf1 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.188659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.188659Z digest=sha256:18f851a8435c77f9f33754f0f6b12f4950d1b209f61d0aa6b8b1a458f9058f5c

Observation ba6f85e9-2379-493d-be2a-9cf6a11b8fc5 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:16.954953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:45:16.192143Z digest=sha256:152e799961a32e9dfb17cfd675c5ce806f930823791049fbfd16e93bc2e77075

Observation a319de42-8823-401c-ae81-93c3e7d9b326 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.195698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.195698Z digest=sha256:47b08c17af61ac366ca18f3ba4d3730c6de5ff41b2e153d8fd99e1f15421b9a0

Observation 2877a261-d1ea-46e7-baae-f3e2868fecbd · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.199067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.199067Z digest=sha256:ca807885c13f16a20dc615d477165a57ca2a488e5d372886ff304aa71cc7b6f1

Observation 86fd2ccb-ce24-4e1e-8005-5574ed2bc5c5 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 8

Resolution
malformed identifier
doi_truncated, observed 2026-08-12T16:45:16.286367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:45:16.202283Z digest=sha256:8b8a91c06f02b0bbb035e22338f500d9f1a177d1eaac0b264cdf0ce5ee2bef66

Observation 47810618-584b-478e-b3ab-0dc86d1b7c13 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:45:16.945776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:45:16.205442Z digest=sha256:c27ce3a3ca36757cb4ad1e8d0b7db464ac926475bb9f425f1635cfbc55025fc9

Observation 63177dff-b5d9-4d3e-957b-3c3e97414cb9 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.208443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.208443Z digest=sha256:4ef1db344bdcdbfacbdcfd3ecaf9e718dbd0c37c71c0f5012b63b966a74a8f5c

Observation 6cb2ca0e-e5b9-4f25-8638-1a7657acc8a2 · outbound

This paper cites Losada, and Álvaro Barreiro.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Losada, and Álvaro Barreiro

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.211432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.211432Z digest=sha256:29f7e9d7cba41cc7a1a466a534394605b23e0658e17b3dd072e16a538cddc056

Observation 0f406067-1086-452a-9b58-e535f94cef88 · outbound

This paper cites Rahmani, Nick Craswell, Emine Yilmaz, Bhaskar Mitra, and Daniel Campos.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Rahmani, Nick Craswell, Emine Yilmaz, Bhaskar Mitra, and Daniel Campos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.214392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.214392Z digest=sha256:ef0865fa6eae650a54545492be8ebd69f1bf7cbaa1e49c147bcf11ea6dfb1e4c

Observation 1b0a63bc-0732-433b-b94f-57f2d74f53fb · outbound

This paper cites SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation SynDL: A Large-Scale Synthetic Test Collection for Passage Retrieval

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.217587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.217587Z digest=sha256:80cd6cc9efaa5f281e07524e5700d82933755a7dea174f4690f169211c48ae2f

Observation a75077f7-f433-44f9-95b8-7acf7330784d · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.220944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.220944Z digest=sha256:dba09ac101f338b20e00dafd94bdca25a90deb64c3af701f76d6a118f6e9b8f4

Observation 2d70b950-2328-4124-8d5b-f7841ce428da · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.224025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.224025Z digest=sha256:6c14dda783abbd21d794f38921ee9c7aae26533b5df35609eb48a2915ec12806

Observation 2d0f5d2f-ac9c-4096-81c6-035bc5fa079b · outbound

This paper cites LLMs Can Patch Up Missing Relevance Judgments in Evaluation.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation LLMs Can Patch Up Missing Relevance Judgments in Evaluation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.227040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.227040Z digest=sha256:7b609e6885b7d0f0fa239a57ccfd0b213ae0367984f5146b4aedd8e67be2a216

Observation d9a39eb1-883e-4ef9-8e2f-d267468b9c6c · outbound

This paper cites A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation A Large-Scale Study of Relevance Assessments with Large Language Models: An Initial Look

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.230608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.230608Z digest=sha256:ee2d39331fd4fbc994f76a4b79a5c657efd704c6fbcd5db00fee613285209eca

Observation a28b4541-130e-4fb8-aad0-e16e78c42fa9 · outbound

This paper cites UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation UMBRELA: UMbrela is the (Open-Source Reproduction of the) Bing RELevance Assessor

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.233934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.233934Z digest=sha256:c5e17e40a39cf7557f8f67ca19e9ef67ce9fdc64f1be64681e58bbe94b93cf44

Observation 3173b6b2-a376-491c-b094-64c0585e645a · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.237208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.237208Z digest=sha256:2818bde38387bc33a30c169bd5892755251058b7a0c7408cd545bf5aca89a50e

Observation fb359f50-d6b9-4044-a36c-19f311244491 · outbound

This paper cites Voorhees.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Voorhees

Reference 20

Resolution
malformed identifier
doi_truncated, observed 2026-08-12T16:45:16.276630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T16:45:16.240158Z digest=sha256:6e258d9bff5cf849fd9a748a86b4524418ee2e35199d3ae8f0d99b100816b7ca

Observation 03e2af04-5850-4189-8f0e-3026fdc512b7 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.243268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.243268Z digest=sha256:7fbb510ba48c44568c5b9344e59d60447cdca98596e245509197ab3dd620064f

Observation 239b7110-afd6-451b-a258-12c7eaa8d807 · outbound

This paper cites an unresolved cited work.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.246309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.246309Z digest=sha256:21b8ecbe04eb156a932ee4afd9b17ea1792b7cee72d77b344472c67a0128b7e9

Observation 183aa483-9e76-4d1a-9cda-f107feacfb56 · outbound

This paper cites Aslam, and Stephen Robertson.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Aslam, and Stephen Robertson

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.249314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.249314Z digest=sha256:c3d2c6e87ffc029c4004f628a01067d25deb2adfa50e2aba1cd228a41231d38d

Observation 33d3d959-0248-4a3c-9cd5-4354eb5fa3d6 · outbound

This paper cites Can We Use Large Language Models to Fill Relevance Judgment Holes?.

Limitations of Automatic Relevance Assessments with Large Language Models for Fair and Reliable Retrieval Evaluation Can We Use Large Language Models to Fill Relevance Judgment Holes?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T16:45:16.177722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:45:16.177722Z digest=sha256:b5324d132edd100c63fb7ee192da0d38646ba9181a1213148e727cf86bef0001

Pith citing papers

No inbound Pith citation observations are available.