Pith. sign in

Paper Citation Record · LEDGER

The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2406.05761.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.05761 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:34:42.280569Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:36:27.452622Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3f67dbd6-a5e2-4f4f-9e15-6e82d5d3ef93 · inbound

OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs cites this paper.

OpenScholar: Synthesizing Scientific Literature with Retrieval-augmented LMs The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T15:29:32.632423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:29:32.632423Z digest=sha256:8f74e1e6f3d8fabdcda21a5810f2e4fe5948507fbd68305d6652f04eb4a87a99

Observation 14d1d0cc-b665-4a4b-b844-73eec63fc032 · inbound

Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning cites this paper.

Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:54:52.920460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:54:52.920460Z digest=sha256:882e66af19e6f70672f576a4f6d3ca43fc3680db366d2fba9802f4531925baba

Observation d0960f4e-fff1-46d6-96c4-885849d85175 · inbound

GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking cites this paper.

GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T12:32:23.877682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:32:23.877682Z digest=sha256:3cd6399db601b6853a7d647cd2cd299fd523127d8fe0b937275e74a5b977d0dd

Observation edf4a044-44f1-432b-a217-65ca0d19f1f8 · inbound

MetaSC: Test-Time Safety Specification Optimization for Language Models cites this paper.

MetaSC: Test-Time Safety Specification Optimization for Language Models The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:44.981245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:44.981245Z digest=sha256:c7f76b5ecc1f75dc43cbb4215ca5758cbce61e6c6dbc7b47dffe1c6ddc197ba2

Observation 1f62dacc-111b-4c33-981d-f83a21f7aca7 · inbound

Salamandra Technical Report cites this paper.

Salamandra Technical Report The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-08T04:58:32.668823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:58:32.668823Z digest=sha256:5880a1991890afa2e426c93361e8c87312432bda4791124007b43c641f31b053

Observation 52971542-9699-47ac-9a56-320483853afa · inbound

EvalAgent: Discovering Implicit Evaluation Criteria from the Web cites this paper.

EvalAgent: Discovering Implicit Evaluation Criteria from the Web The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:34:42.280569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:34:42.280569Z digest=sha256:d5ac2221c93f17fc52dab63442fae66c22c5ac2052f042ceaf3ffc5e4cd55889

Observation 022857c0-015a-464c-bfbd-d7c1624d3580 · inbound

Trillion 7B Technical Report cites this paper.

Trillion 7B Technical Report The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:32:05.941324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:32:05.941324Z digest=sha256:1f30f667cc87862ee91bcc78db553217263c4443d50d65ea1324560a2d12cc26

Observation 67ce7fc7-be5d-400a-9375-59c2bb7de809 · inbound

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models cites this paper.

Decentralized Arena: Towards Democratic and Scalable Automatic Evaluation of Language Models The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T20:32:25.241466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:32:25.241466Z digest=sha256:c7a72fa1f6f9027694077ad9db235045dba37856a28d0bf495ab09104d26ece1

Observation 9bcc8e31-95e2-4215-9e50-b123a60c7ee0 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.036636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.036636Z digest=sha256:8e4d0ec7cfc80749826eb9b61cef5c3fb1b147ae0cdcdee1c6d9e920bd62849b

Observation 2f7fa853-9edd-4b9b-b9f2-82d2ba362b9c · inbound

Aligning Large Language Models with Implicit Preferences from User-Generated Content cites this paper.

Aligning Large Language Models with Implicit Preferences from User-Generated Content The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.001931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.001931Z digest=sha256:119e7f2d6fa3c06c5ac8ef007d74013c4fd5ae9eff962c371673c60c8ef81432

Observation 840b7fc9-1e0e-4dcc-99f5-49bfd2af4791 · inbound

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability cites this paper.

An Empirical Study of LLM-as-a-Judge: How Design Choices Impact Evaluation Reliability The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:31:35.564117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:31:35.564117Z digest=sha256:ed241358ceb62c809f7849839b835bac0cbbd83cfe18dbb149e2388de2ffa67d

Observation 7d636483-3ae0-4150-890f-df653d0b1174 · inbound

Evalet: Evaluating Large Language Models through Functional Fragmentation cites this paper.

Evalet: Evaluating Large Language Models through Functional Fragmentation The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:01:40.061056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-18T16:57:25.259866Z digest=sha256:23fb8971ba74f01e2f9920447aa767aa54e200c085ac933f34686456b3c2ca9d

Observation d9e87539-1a81-450a-b433-0784c542ef95 · inbound

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy cites this paper.

Detecting Unfaithful Chain-of-Thought via Circuit-Guided Internal-External Discrepancy The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T21:53:59.562279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T21:47:17.894881Z digest=sha256:7621661cead89964250523a192d7c42c0fc0ad4229331ea1a02d10ccae1bf9a7

Observation d42c853c-44aa-4875-99eb-f390e9b72887 · inbound

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks cites this paper.

CoEval: Ranking Language Models for Custom Tasks Without Labeled Data or Trustworthy Benchmarks The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:27.454057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T10:46:24.554332Z digest=sha256:40796c8941f859e9f249aae7df260afa8a0969a8defbafd89e56ecf7773362db

Observation acad77c1-3bac-4ea6-ac20-e78029278a30 · inbound

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation cites this paper.

FinReportBench: Measuring and Improving Institution-Grade Financial Report Generation The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T18:53:02.007620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:53:02.007620Z digest=sha256:1c66558df80b4f43577d757506b59dfcd144ec3d53d547c679d834f162d9ce03