Pith. sign in

Paper Citation Record · LEDGER

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models

As of 20 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 2 inbound Pith citation observations for arXiv:2411.16797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16797 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:21:28.151808Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T17:12:14.146302Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T17:12:24.245674Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 95bc0cae-529f-4883-8a40-ca69200b221a · outbound

This paper cites GPT-4 Technical Report.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.011334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.011334Z digest=sha256:9c29923b703413433ebb7c71a8efde530b845bd9794ed940d17a4ca1ac74ecc5

Observation a87d9bb1-b8f4-412e-a86e-1c7728b940ef · outbound

This paper cites LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.045862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.045862Z digest=sha256:ff5174753e0ef3b744fe97bd57ae573b2df59ac9277c2dd90ca35772986e2895

Observation 19300bb6-215f-4525-8c1b-dfba43ff98a7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.057753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.057753Z digest=sha256:c10b22058cce94972fd94d80ff06104a4d683b90711aea59e4be5e23929160f9

Observation 61f6a845-ddc3-4742-a7ea-4ee0279e2fe5 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.069139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.069139Z digest=sha256:5b7f961b4d201dbee46f33cf106d49af45f334a5fa5b99b54a34b6cc488c6a36

Observation cae75261-3045-4195-a7a9-874e2055e0ce · outbound

This paper cites Improving fairness in machine learning systems: What do industry practitioners need? In Proceedings of the 2019 CHI conference on human factors in computing systems, pages 1–16,.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Improving fairness in machine learning systems: What do industry practitioners need? In Proceedings of the 2019 CHI conference on human factors in computing systems, pages 1–16,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.583008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T13:21:28.074657Z digest=sha256:51c3c9c5a67726628c0e00f9e74eb6a99f21767e004ccf58aa58fa7b2bdb2640

Observation 53caf2f2-4260-4faa-9023-f7f2b949041c · outbound

This paper cites Harnessing the wisdom of crowds in wikipedia: quality through coordination.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Harnessing the wisdom of crowds in wikipedia: quality through coordination

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.565545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T13:21:28.086707Z digest=sha256:062d53fb093c2edd25db4be50021b77d1463831b760acae8e11d4ca9c77953ee

Observation 4e4c0ce4-9ec1-4d48-8447-56aa6fd17594 · outbound

This paper cites Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Merge, Ensemble, and Cooperate! A Survey on Collaborative Strategies in the Era of Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.096521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.096521Z digest=sha256:1816fd54d9b48f3448b824e4e1d95e6b5a61e1c919f4b3a17921f48c8f1b4ca7

Observation ba4b6d1c-96ff-45b9-b14a-67e25f3d780f · outbound

This paper cites Language Models are Few-Shot Learners.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Language Models are Few-Shot Learners

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.101763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.101763Z digest=sha256:fb41885d08c78800eb87b0f58c8efb034ddf047e444aeade8beac6859eca4655

Observation 27e43cd9-f0eb-4806-acb8-f83fa111fd6c · outbound

This paper cites Brent Mittelstadt.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Brent Mittelstadt

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.547773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T13:21:28.107527Z digest=sha256:4e332dfb5e7b9ab8077e2bc0f0fc98a6369cf586a6fe7908508b84fe1e23f5d3

Observation 34c92536-b659-430a-97b9-ef7c8ab81621 · outbound

This paper cites Mitigating bias in algorithmic hiring: Evaluating claims and practices.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Mitigating bias in algorithmic hiring: Evaluating claims and practices

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.530508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T13:21:28.117355Z digest=sha256:56fc2bba5a88e91175eb22b34211dfcb2bc25f523db445844b69f40c0c5b200f

Observation 503f515d-7c8f-469f-9462-2aea3a1d1f8f · outbound

This paper cites Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Corex: Pushing the Boundaries of Complex Reasoning through Multi-Model Collaboration

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.122091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.122091Z digest=sha256:988d7727f4b25827a8d8cb56071242f3cf15476c8e7bc8c44cbb23a7d447c3a0

Observation 2601fffa-15fa-4834-b521-b3049b48088b · outbound

This paper cites Galactica: A Large Language Model for Science.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Galactica: A Large Language Model for Science

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.127021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.127021Z digest=sha256:1a0d7f504749a89dbdb9f74617e1bcd8be884047328ad14f3dda3bc68118e1e5

Observation e6b5a435-7c4b-4712-a09f-a082c53473a9 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.131826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.131826Z digest=sha256:70f302534efb195f8526d2793c83c10a14ab5f7c4be5da15fe95c70d08577d3c

Observation eb19abfe-7bb2-4ca8-933e-5ed0ada9733e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.136780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.136780Z digest=sha256:84110035ccdce6dc98ac2914f18474e065bdccace38d0d3748a752eeba92a397

Observation 54595ea8-8143-4981-a0fb-8dd4dd0d188d · outbound

This paper cites The collective intelligence of random small crowds: A partial replication of kosinski et al.(2012).

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models The collective intelligence of random small crowds: A partial replication of kosinski et al.(2012)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.513679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T13:21:28.142188Z digest=sha256:04914dfc00ad7235d080b24ecaffaba5d004010ded864323db189428a13046e0

Observation 6fe34682-5c77-4a20-96ab-403d2b2a1928 · outbound

This paper cites Deep learn- ing for computer vision: A brief review.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Deep learn- ing for computer vision: A brief review

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.495559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T13:21:28.146899Z digest=sha256:0670fb2359b8ee324137b20e019bbdbbbaab2088eccf1f734ff9d45395615c72

Observation 12b2009c-b670-415f-9e3a-f3a2d16cb20f · outbound

This paper cites Large Language Models and Causal Inference in Collaboration: A Survey.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Large Language Models and Causal Inference in Collaboration: A Survey

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.091534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.091534Z digest=sha256:f47ad4f67dba3062e55d3e198590a5e1c8fc832825063af88378b69084370d2d

Observation d569f57f-47d3-42e4-8597-0465ff4f3ff7 · outbound

This paper cites Accountability of AI Under the Law: The Role of Explanation.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Accountability of AI Under the Law: The Role of Explanation

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.051711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.051711Z digest=sha256:ece1b4b5610b218d410c15da728c2500f5df89aa3b03e9d00d3ae6c99d742d54

Observation 0854f7bf-941b-457e-9c80-7013be8bdfc7 · outbound

This paper cites Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel Collaboration.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Ensemble Learning for Heterogeneous Large Language Models with Deep Parallel Collaboration

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.081674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.081674Z digest=sha256:f8675932821f8352bdda813d22dcc15b209c306bfae763c3149eaa3bc724ce54

Observation 3c86ee35-6580-4e3e-a5d9-06b33a76a49f · outbound

This paper cites Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Exchange-of-Thought: Enhancing Large Language Model Capabilities through Cross-Model Communication

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.151808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.151808Z digest=sha256:faebd276abd33c900f43010aa8c1f07844c1849b0da5d0f0cd7765efe0dee649

Observation 9d7ca5ef-e005-41ae-b3dc-9c4b0dc066d4 · outbound

This paper cites On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623,.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, pages 610–623,

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.028565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.028565Z digest=sha256:1f1b1bbcd53d86e7034956d9c1216e383f36602942b6563d20bdd8369c05071e

Observation f8771898-a9fa-4b81-966d-fd942435a37a · outbound

This paper cites Probabilistic Consensus through Ensemble Validation: A Framework for LLM Reliability.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Probabilistic Consensus through Ensemble Validation: A Framework for LLM Reliability

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.112281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.112281Z digest=sha256:931b1b5fe67b718fd6b5ef3bf31a75268df78cf4133cf1faa2f9a60de0d52db8

Observation 48530819-7c65-4fbe-ab5d-a983f3184455 · outbound

This paper cites Open Problems in Cooperative AI.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Open Problems in Cooperative AI

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.040010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.040010Z digest=sha256:e9956d0442a60726b410a1bdb3803516b2812701f0da53c6c69fe2639c0a5381

Observation 05d23a2e-e037-42cb-94c6-a462d9c0d66b · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models On the Opportunities and Risks of Foundation Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.034743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.034743Z digest=sha256:ff9bead4a87f0a9310db9cbcb2f745fa257f9f202db5a467d25d9f92f3406604

Observation 09915226-6215-4d05-90b2-57f47b6f486b · outbound

This paper cites Fairness without demograph- ics in repeated loss minimization.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Fairness without demograph- ics in repeated loss minimization

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.599437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T13:21:28.063661Z digest=sha256:47070051b654d891c784f3ab49989e95468da8329f2a42b9099e3ee5dbbcfa8b

Observation 6d595faf-8b1c-4b46-bc1c-6a69ebc3f796 · outbound

This paper cites Large Language Models for Mathematical Reasoning: Progresses and Challenges.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Large Language Models for Mathematical Reasoning: Progresses and Challenges

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T13:21:28.017481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:21:28.017481Z digest=sha256:987839122e642711cdc5bd3720695003a58b31bb0ff68c322e7e9f74ee7e25f8

Observation 42229a8a-2355-4e89-9290-239f8a47b5af · outbound

This paper cites Anthropic.

Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models Anthropic

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:21:28.626772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T13:21:28.022885Z digest=sha256:b36a33570819ee049702010bf63ad303a57803cc6b1c3fff27b798b1cd8a69fc

Pith citing papers

Observation b6c747e7-9df5-4b07-b6d4-2880cad55f2f · inbound

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning cites this paper.

SIV-Bench: A Video Benchmark for Social Interaction Understanding and Reasoning Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.735769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T11:36:36.687324Z digest=sha256:91c1a1ec57ec1e1302720a2d94c1b01bcfc8f4fed4568c61db69bfeb1712d07d

Observation aed2077d-01a0-4a6c-ae0b-56bd1a8115af · inbound

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment cites this paper.

Truthful AI Advisors: A Pre-Specified Benchmark for Large Language Model Honesty Under Preference Misalignment Enhancing Answer Reliability Through Inter-Model Consensus of Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:24.247068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T17:12:14.146302Z digest=sha256:e5acf96994d2dbbb2c7be2b56489511a07f2d9ff097d0ebe6ebcad0c262b9091