Pith. sign in

Paper Citation Record · LEDGER

FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2404.06003.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.06003 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:40:13.907632Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:10:41.139960Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a3c7e143-51de-4d68-9ec0-e218f84b6280 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

Reference 175

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.142309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:fca9b71286eb0af6cbd0cfb85018f3d8dd488eb81818de484924ac0d580a6c3a

Observation 7dca7866-e9b6-499c-960e-481489b1c060 · inbound

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation cites this paper.

Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T11:40:13.907632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:40:13.907632Z digest=sha256:ad232286c59ef58d01cedf86399bbf4c5fb0fbacf302ec648d3cb8f2eb0ff2d3

Observation aac12416-9c89-4f07-a085-38a7e4404567 · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.109652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.109652Z digest=sha256:a9d0ed761675e28467a1a7c99044317f34001080c0cb530408ed8e92023a3cac

Observation 0c2c9326-3176-47d3-bd94-fa9397f04101 · inbound

Behavioral Fingerprinting of Large Language Models cites this paper.

Behavioral Fingerprinting of Large Language Models FreeEval: A Modular Framework for Trustworthy and Efficient Evaluation of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:04:08.940093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:04:08.940093Z digest=sha256:81ad0b8610f6d228d0908dbd08f1ab4d3fd908f88e86527aa8e54f5f279a12c6