Pith. sign in

Paper Citation Record · LEDGER

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities

As of 17 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2506.12376.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12376 v2

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:10.932988Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d050e6a-69c9-40f2-929a-0363616f7a0b · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.821107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.821107Z digest=sha256:aedac8c4d3cd612979dc5f4aca4d1476e78da66248f752da5719c94627860246

Observation ad1334f8-3b07-44fe-ba21-7ea852ba5ddf · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:11.385892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:10.825536Z digest=sha256:72437134aeb4722707469e2d0d96b83aa0e1fe92a76ce0aebc96c03ceb15d065

Observation 2f840291-5fd4-41d5-aefd-257aa267c3cd · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.830526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.830526Z digest=sha256:a506356711f55c452991eb3049f72827f2ab998e8e605bb502606801d7e4b9c4

Observation 3bd6c51b-ed0b-4e60-8830-8115d7f6ea1a · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:11.375353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:10.834151Z digest=sha256:cfca7ed510b04f20e2f3448e4eca2754460d40e1fb77b4296a6885febca1ffba

Observation 8778659d-18f0-4202-a4ae-92c2799d5c1f · outbound

This paper cites Language Models are Few-Shot Learners.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.837613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.837613Z digest=sha256:61ff48ec379a80c8c9b15437fadfd429455d66b99b171ae2bd0796b69d5a08db

Observation 7cfb0def-27cc-495f-a134-7dd54ae0430a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.840876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.840876Z digest=sha256:652c909bf748cdc320eabd9aecd497d17390a7346d1e7dd5a69ec9e92165ae09

Observation cb38e085-ba45-4357-baa4-dc0d0c94f7a4 · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:11.366029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:10.844622Z digest=sha256:1fd8e97b469486cad2aa7e1134aef5a3d37023409a11bb9af73c8e84f5586af6

Observation 08eb1bc9-ba21-4081-938e-685779d06d5d · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:11.356190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:10.847412Z digest=sha256:bdf4a112a739de451051b5dc5da9b292badd6e903954febc4f23db836fd05faf

Observation 0433ecc0-9904-4509-9613-f476493d1bda · outbound

This paper cites Malin, Sricharan Kumar, and Jiaxin Zhang.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Malin, Sricharan Kumar, and Jiaxin Zhang

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.850485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.850485Z digest=sha256:0e18b2cd2c71d83a3470da00c06c6b92e25dcfc6f4a5757f90928dd76c7771d5

Observation f19a296a-b8c6-4db2-a841-9e573f80affb · outbound

This paper cites The Llama 3 Herd of Models.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.853445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.853445Z digest=sha256:f50c6cc5d6f190116deb69492f239f256f2af687bda697d05290a88c97e5a7f5

Observation bf334976-5606-46ea-920e-c8115ebb92d4 · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:11.346614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:10.856654Z digest=sha256:4cfecf38b2524ed761ebd9058e2cb6a83ddf9b468038627cd7d859f6c59adc62

Observation 7fb5c1bc-f0a8-4ee5-adf9-03af834ef47a · outbound

This paper cites Mistral 7B.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Mistral 7B

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.859209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.859209Z digest=sha256:73d21909ce4db668550598891cb5e2eb51bf205b5da9cf2755ed41b6037760d8

Observation 3c540a06-3b16-476f-90cb-032bd58883f7 · outbound

This paper cites Mixtral of Experts.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Mixtral of Experts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.862036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.862036Z digest=sha256:c042f282873ae35433580dad0707add28efccb0bb23fc8082753e1564d788335

Observation 766c3ef3-49ac-4d5a-b56d-6f7560111c2d · outbound

This paper cites A Survey on Large Language Models for Code Generation.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities A Survey on Large Language Models for Code Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.864820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.864820Z digest=sha256:cca0a442c3671d222e8560d6a6bf94f812f3832aaa470c2092c5e9b7ce14ccd5

Observation 42595550-5a18-4443-a1df-1435d5ba8007 · outbound

This paper cites MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities MetricX-24: The Google Submission to the WMT 2024 Metrics Shared Task

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:57:11.208030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:10.867801Z digest=sha256:5b93cbc6f0cf7c510b89c1557da0952106a90da47273ae3766092a1cd9426ea8

Observation bd23c936-8285-4087-8950-e6fb2c9d78ed · outbound

This paper cites Preliminary WMT24 Ranking of General MT Systems and LLMs.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Preliminary WMT24 Ranking of General MT Systems and LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.870910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.870910Z digest=sha256:33fd1f05ffe7c72eae7517e4c03886fb212b3fde57f8c580f3bb9f548a8b7259

Observation 6371c8b8-2223-47bb-ae69-5fa74cfed51e · outbound

This paper cites MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.874023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.874023Z digest=sha256:d8b52b96cefc63e1f652d92e142103c77f8d23031f7ede6317c68ae223025fef

Observation 6edf4539-aaf9-4890-85f6-e98b888a3944 · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.877303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.877303Z digest=sha256:a9af219d1a8191ac7996c64f04bff1068c1e2b22ac09240349ace3ba8bf45c94

Observation f35e663f-0bc7-4a89-8eb9-efd88d9bc290 · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:11.330094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:10.880538Z digest=sha256:8e46e4aaaabad692520f8e56983eea8a29299678a6cfc5c16b2071af14076b37

Observation bdd275b8-91d1-4a23-960b-b60932644982 · outbound

This paper cites Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChain.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Beyond Accuracy: Evaluating Self-Consistency of Code Large Language Models with IdentityChain

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.885298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.885298Z digest=sha256:bf379745af35b964685e35f5754bab18ad1e84903d07a6b71630038010055087

Observation 1acbad04-91ff-47d4-9a66-16e92ff28842 · outbound

This paper cites Large Language Models: A Survey.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Large Language Models: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.888528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.888528Z digest=sha256:9840803eb5edaf929594575c15daba07da7cd2a96796a08985d76dec53c9da7a

Observation d6f0a735-8925-4de6-9463-b6d053771fc0 · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:11.320159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:10.891954Z digest=sha256:cbca82b6078f2ee9660fd5aa773057b922e09211fe768a7865631c516b19b4d3

Observation 15b82ac5-ab48-4da9-9588-337a88546413 · outbound

This paper cites GPT-4o System Card.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities GPT-4o System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.895294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.895294Z digest=sha256:08e3eff978b0e84f2332c43a15eb7f1bafbad110b8e06cb109d8eb92165ba0cb

Observation 0617a9c6-65c8-4ca7-ab18-677fab056174 · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.898912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.898912Z digest=sha256:a1d76e19b14d7602d235e10bd5665687ad3b1cbdfb89f8cecd21e17d324d3e7c

Observation ca15a532-652c-452d-a539-bbda996e90c4 · outbound

This paper cites Scaling up COMETKIWI: Unbabel-IST 2023 Submission for the Quality Estimation Shared Task.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Scaling up COMETKIWI: Unbabel-IST 2023 Submission for the Quality Estimation Shared Task

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:57:11.055345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:10.901731Z digest=sha256:646aef04ae125b1e151192d835ad96a5a66956ed8b4bb04815cb08827ab75e97

Observation b78b934b-c0c5-48e4-a050-5cde0f8bda71 · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.905184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.905184Z digest=sha256:e03301b42d0a00076b6a2c9fcb4c9fb1e1daf830284443db443402e8436c7e58

Observation 75ecb9cb-f076-4b2a-86fe-a714d4b318db · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.908249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.908249Z digest=sha256:cdc69e9536938d52676dd2caf89cdd5b2d98e2e5b919394ce64b9c44276c6707

Observation 9426f444-e2dd-4eeb-9a11-66b08b1675ca · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 28

Resolution
verified exact
doi, observed 2026-08-07T00:57:10.963083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:10.912099Z digest=sha256:5790163b28fbee1cb9c19860fe58e49d36c22baa3860f766b21eacb712e58f69

Observation 0ebd2b94-b99e-498f-8f9c-f19c1620819a · outbound

This paper cites Gomez, ukasz Kaiser, and Illia Polosukhin.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Gomez, ukasz Kaiser, and Illia Polosukhin

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.915097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.915097Z digest=sha256:c45622b668082fe4a275baa1bc95163415adce72972a676125a8b9a8a90a839b

Observation b2db46cb-b472-4c53-a9af-9c12381f108c · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.917840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.917840Z digest=sha256:7511a22a6693e19dad2d02dbb3f7d624af26fd572584835d26f15f9540689695

Observation e9a8494b-d2f9-47f1-aa87-7a6f86346bc7 · outbound

This paper cites Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Latent Space Chain-of-Embedding Enables Output-free LLM Self-Evaluation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.920765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.920765Z digest=sha256:3fcfd756a89c62c5da3ac5d9f85e882e9dd60f598a92b41433a115ef38e11e0a

Observation 0a388d33-2556-4b29-bc0e-f8fd39ff955e · outbound

This paper cites an unresolved cited work.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:11.295020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:10.923723Z digest=sha256:1ee617d4c217508ba92b310df96907bb0638ecf44aba8a04d928c6135fdae11a

Observation 7262afd5-a0b6-4607-a037-09c8a69a14d0 · outbound

This paper cites A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.926390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.926390Z digest=sha256:f46a7e0a32a50fb64d197ad9438ae2f23e39ea862a1247a513c6bdecc775f548

Observation 5349fefa-a554-421c-82ef-0da9cf739b54 · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities Benchmarking Benchmark Leakage in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.929888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.929888Z digest=sha256:1f9e5012130630793e03dd077e37dd102ce1e196cc093dd505915a59cf6d6eea

Observation 8c4f554b-11c9-4c03-88a8-301010ae9143 · outbound

This paper cites LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks.

ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:10.932988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:10.932988Z digest=sha256:5d1273fc553807b26bd74c20139608fc064e9486095cbbe5a823e5ed1efb7348

Pith citing papers

No inbound Pith citation observations are available.