Pith. sign in

Paper Citation Record · LEDGER

Self-Taught Evaluators

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2408.02666.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.02666 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:52.667559Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T11:39:47.080762Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e284f656-bc61-432e-bee0-b54d8ff14625 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge Self-Taught Evaluators

Reference 162

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:35:44.077704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:d80b08bd0bb1b840daece492c8f40b6f09339eabfc40e797362cb52d535ecf95

Observation 3a599e2b-4a53-405d-a5ba-9b88dda5f78f · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods Self-Taught Evaluators

Reference 240

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:35.619226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:9997aa8f29545744b5b10f1165b901528c05d3e3c5238a69e4f4c10626c44c7b

Observation d720deb4-efae-4f39-b636-0d2064001d37 · inbound

Reward Reasoning Model cites this paper.

Reward Reasoning Model Self-Taught Evaluators

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.667559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.667559Z digest=sha256:bc12a65a2e6788eb3996704089f182d821b7c1063f79b674663eba532d5074b9

Observation f42bd0e0-562c-4f4d-b9c7-f08dab4cba0b · inbound

Learning to Reason via Mixture-of-Thought for Logical Reasoning cites this paper.

Learning to Reason via Mixture-of-Thought for Logical Reasoning Self-Taught Evaluators

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.997675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.997675Z digest=sha256:3778f0ea4b5d0888624cc9589be2718a9a0902aa534e80830dee582af6ab55fb

Observation 2d11fb40-7bb4-40ab-854d-10e171d8a16b · inbound

SkillVerse : Assessing and Enhancing LLMs with Tree Evaluation cites this paper.

SkillVerse : Assessing and Enhancing LLMs with Tree Evaluation Self-Taught Evaluators

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:36.267450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:12:36.267450Z digest=sha256:0959f696d02f31fc87d881e5e39ea353a182acea6aa7d826ed35701bee36f874

Observation f6fc7ac3-3d55-4304-b3b2-e556f1ec9ee3 · inbound

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning cites this paper.

Beyond Gold Standards: Epistemic Ensemble of LLM Judges for Formal Mathematical Reasoning Self-Taught Evaluators

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:18:38.084976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:18:38.084976Z digest=sha256:f65c36d6ff526f2b5c27a96c6f4346f524e497d8b8072ca1f818073f4eb8bd19

Observation 164df63e-321b-459d-a263-6131ebd555aa · inbound

From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation cites this paper.

From General to Targeted Rewards: Surpassing GPT-4 in Open-Ended Long-Context Generation Self-Taught Evaluators

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:47.837626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:49:47.837626Z digest=sha256:8ad69d7fe5d58d66823f5f2e8d5571dad482b7f33df3ded6128b96457f180fbb

Observation ac0be8b9-04c1-4ea2-a14f-64a21bb03f77 · inbound

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks cites this paper.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks Self-Taught Evaluators

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:33.368009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:33.368009Z digest=sha256:60240337203177071b0255ce1665a38bc01eb8ee70fb7b994f5a42449400530a

Observation 2d9e0ca0-a2e4-428f-a9f1-f5eb5f577acd · inbound

Improving Data and Parameter Efficiency of Neural Language Models Using Representation Analysis cites this paper.

Improving Data and Parameter Efficiency of Neural Language Models Using Representation Analysis Self-Taught Evaluators

Reference 145

Resolution
unresolved
no resolver link, observed 2026-08-06T17:03:46.209554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:03:46.209554Z digest=sha256:11479a1b8bbb503d9ad651d7abe5c65e083e12169362bda5e770c6d13f2078dc

Observation 6339ba3e-8db3-49cc-ae12-7296d9190e59 · inbound

Multilingual Self-Taught Faithfulness Evaluators cites this paper.

Multilingual Self-Taught Faithfulness Evaluators Self-Taught Evaluators

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:24.899310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:22:24.899310Z digest=sha256:691263da1eea55100bba4184fcee3ad7ac58433e7e8ce2103cd197ab0bf86e2c

Observation 13502761-0064-481f-a9ac-e087347d6afa · inbound

Improving Large Vision and Language Models by Learning from a Panel of Peers cites this paper.

Improving Large Vision and Language Models by Learning from a Panel of Peers Self-Taught Evaluators

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T12:27:27.624059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:27:27.624059Z digest=sha256:c1c531547c78e863eeef8910512f2d7a93a142fc7c0e22cefbbd579f2a0f566c

Observation b5f00401-d095-4c5f-86a6-8d031751195e · inbound

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization cites this paper.

On the Shelf Life of Fine-Tuned LLM-Judges: Future-Proofing, Backward-Compatibility, and Question Generalization Self-Taught Evaluators

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:56:24.630408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T12:53:45.767341Z digest=sha256:e746486d487e2b916c16f23f1611b8a562f72e99d45b388c7faf15d501413f0a

Observation ef3f7e79-e7b6-4bb2-b2b8-0f13d14ed7a5 · inbound

Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation cites this paper.

Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation Self-Taught Evaluators

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:07:43.335795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T10:04:44.379128Z digest=sha256:0f75b2496b0db9ab9e16c8514d6f6b3668cc6fce8aa9c09d85f80771e5f6d3ad

Observation c956993e-8033-4bb3-b88b-75403199d96d · inbound

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation cites this paper.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Self-Taught Evaluators

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:30.964243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:30.964243Z digest=sha256:aa2db1e41cf98daa28b2c56c27ce8800a9806366920e8e1d7a53415020fe90a6

Observation b809b336-e4db-44ff-922c-99944a7e300d · inbound

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety cites this paper.

MultiBreak: A Scalable and Diverse Multi-turn Jailbreak Benchmark for Evaluating LLM Safety Self-Taught Evaluators

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:31:01.129656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T16:00:32.413225Z digest=sha256:b37149f66c0a7d4014cb2f3a795b5d6c8bf785495c40dd85252ec2190c04f416

Observation a996b724-8eb1-4702-9738-752da9dde08f · inbound

Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation cites this paper.

Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation Self-Taught Evaluators

Reference 97

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:56:08.424802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T16:53:00.860162Z digest=sha256:087357c781b147dd802880f1effdca9bc690ab23af9a61d97b3ddf32f375c865

Observation d10d6478-a27e-4e88-998b-15103a3cb8f2 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs Self-Taught Evaluators

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:39:47.082804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T07:49:36.816100Z digest=sha256:c755e6ac5b092149b290aa55d770ee3e77f1f426543cffd0262651335ed2fdbf

Observation db66d51a-4494-44a0-9a20-39e034d00ae3 · inbound

Towards Spec Learning: Inference-Time Alignment from Preference Pairs cites this paper.

Towards Spec Learning: Inference-Time Alignment from Preference Pairs Self-Taught Evaluators

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:44:39.540134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T10:17:33.176525Z digest=sha256:e40c3bb99a4f73ea3c111d8959ee64be3c0b0672c56376e047b9f6884a82258e

Observation 83be3beb-3088-4a67-a07d-45c960555ef8 · inbound

PASTA: A Paraphrasing And Self-Training Approach for Knowledge Updating in LLMs cites this paper.

PASTA: A Paraphrasing And Self-Training Approach for Knowledge Updating in LLMs Self-Taught Evaluators

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:44:37.486590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T09:42:40.357949Z digest=sha256:0cc0233b77817336b727bd3c11117ef52f0a014537d81aee93c57c5b867fa1f9

Observation 551babe2-fe27-43d0-a6b4-11cc905ba28b · inbound

Codifying the Judge: Scalable Evaluation via Program Distillation cites this paper.

Codifying the Judge: Scalable Evaluation via Program Distillation Self-Taught Evaluators

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T12:49:09.307551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:49:09.307551Z digest=sha256:6137128a7d7f9639b7dd83c963f21314b237bc6df96114768756bb620ff07922