Pith. sign in

Paper Citation Record · LEDGER

Can we trust the evaluation on ChatGPT?

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2303.12767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.12767 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:28:14.265774Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:29:51.149973Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e9f7d148-839e-42ce-8b10-84a39255a7fe · inbound

AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction cites this paper.

AI-Augmented Surveys: Leveraging Large Language Models and Surveys for Opinion Prediction Can we trust the evaluation on ChatGPT?

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-24T08:49:14.002548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-24T08:47:23.231930Z digest=sha256:9ba722b66ddeeb99ae358cf0a68def7f1e462d6858db9f2621a3c43d316226ee

Observation 1c92d103-afe7-40c6-9443-8016d59add23 · inbound

Chain-of-Verification Reduces Hallucination in Large Language Models cites this paper.

Chain-of-Verification Reduces Hallucination in Large Language Models Can we trust the evaluation on ChatGPT?

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:06:50.474295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-18T01:06:49.811982Z digest=sha256:352142fda8e87148d57ece913c0d30de645c97749883341de625588eec9ad960

Observation e3a868ec-c0ed-4d52-9f25-df2169db957e · inbound

Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models cites this paper.

Identifying the Achilles' Heel: An Iterative Method for Dynamically Uncovering Factual Errors in Large Language Models Can we trust the evaluation on ChatGPT?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:43:53.980663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-24T04:41:48.938278Z digest=sha256:9fca79a6194f4760259d1cbf9664f3519237c73c171e9ec269f28ee19c626ec9

Observation d8f18681-9e81-4f7e-8a9d-e64bbd1f26d4 · inbound

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive cites this paper.

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Can we trust the evaluation on ChatGPT?

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:04:44.505064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-17T23:04:44.287660Z digest=sha256:66424bbe0775a8f6338f4a53a33baf22ec3855a11bff91ee295d24f690757a5a

Observation cc7a962d-bb78-447a-9fc9-4a979c470ac7 · inbound

Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks cites this paper.

Strategic Prompting for Conversational Tasks: A Comparative Analysis of Large Language Models Across Diverse Conversational Tasks Can we trust the evaluation on ChatGPT?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:28:14.265774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:28:14.265774Z digest=sha256:27bf0b4651307890eb2f2c26e58369738bbee3106f17d44675353e4b0c129fa5

Observation cc855fc7-a873-4d5b-b9c1-b4c173e2cd46 · inbound

Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora cites this paper.

Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora Can we trust the evaluation on ChatGPT?

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:16.300342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T03:13:30.648190Z digest=sha256:2a2aec26978d9f978fafd58a3e366acdedadef6253333c4405a21b5af4884664

Observation 906ae647-b257-48e8-9f6a-320c6f5ad4ee · inbound

Structure Before Collapse: Transient semantic geometry in next-token prediction cites this paper.

Structure Before Collapse: Transient semantic geometry in next-token prediction Can we trust the evaluation on ChatGPT?

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:29:51.151419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-26T05:14:07.208255Z digest=sha256:81fe2d09c7c0b018b852582be17aca8c2be6f634e50b86eef5b2e659340b1bc6

Observation e535b227-9148-4d67-8367-004f59e2c19c · inbound

Prompt engineering using order-of-addition experiments: An application to generating two-level fractional factorial designs cites this paper.

Prompt engineering using order-of-addition experiments: An application to generating two-level fractional factorial designs Can we trust the evaluation on ChatGPT?

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-11T06:09:35.110633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T06:09:35.110633Z digest=sha256:293fd6aadca1ac621c2e2f545b8954087d2e58f6e34bf60b76fd9723ecf8c291