Pith. sign in

Paper Citation Record · LEDGER

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation

As of 12 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 0 inbound Pith citation observations for arXiv:2412.11988.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11988 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:26:02.765845Z

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

11 of 11 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a060630d-61b3-487b-97a4-ce393b1bba62 · outbound

This paper cites an unresolved cited work.

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:26:02.689181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:26:02.689181Z digest=sha256:601e4a64d005b33eb6c70201f86b7403e7d698990cb29a5f24f7ddbbd8adb1e4

Observation 7975dc92-c889-4679-b8b5-7caf0e115050 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T14:26:02.698707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:26:02.698707Z digest=sha256:c784e442f5e06e7009ed2e8976c4d5b82e720ddaab5a576fa129f4a2a31bc466

Observation c790df0f-0c01-491a-8b9f-dae81159b7bf · outbound

This paper cites AI now beats humans at basic tasks — new benchmarks are needed, says major report.

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation AI now beats humans at basic tasks — new benchmarks are needed, says major report

Reference 3

Resolution
verified exact
doi, observed 2026-08-11T14:26:02.844233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:26:02.704292Z digest=sha256:fff4cfa78ff2d07550bc7289ef6f86233da72833d14eac3634af57e673bc4b84

Observation 5a36e48c-08d1-43dc-88e7-8e1e1e71d5e8 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T14:26:02.711431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:26:02.711431Z digest=sha256:d4b9e3eb4c0838ab2132a41de1c19dad2827554f091c820e79f2605e9f6711bf

Observation a97c27ac-a0a6-4193-bf31-d979cd9bc06f · outbound

This paper cites StylePTB: A Compositional Benchmark for Fine-grained Controllable Text Style Transfer.

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation StylePTB: A Compositional Benchmark for Fine-grained Controllable Text Style Transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:26:02.722224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:26:02.722224Z digest=sha256:74766ded2a6ced6cace8bb6743e2f15b349203b76eeb2af57a9203e92d17e083

Observation 0211efbe-614f-420d-aa31-5eda3a7a4614 · outbound

This paper cites Fine-grained Text Style Transfer with Diffusion-Based Language Models.

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation Fine-grained Text Style Transfer with Diffusion-Based Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:26:02.729133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:26:02.729133Z digest=sha256:75e7d89600de6b64ca5f22393927498c25570e01b865464e15090aa2e6ed7ec8

Observation e3d5ea51-3a4e-49b9-951d-11f840bf0365 · outbound

This paper cites C., Shoham, Y., Wald, R., and Clark, J.

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation C., Shoham, Y., Wald, R., and Clark, J

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:26:03.135865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:26:02.738232Z digest=sha256:3adbc0ea1073ccb54424be70049488f5216dfd223fac5e2d60c03afc78599674

Observation 7b3f1a3b-871c-4b00-951f-6764fc93bd23 · outbound

This paper cites Learning to reason with LLMs.

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation Learning to reason with LLMs

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:26:03.114324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T14:26:02.743712Z digest=sha256:d4842e886cebbcc6538e14e3dabfb8eb67bd0cd273cb088d88a10309f206c881

Observation 54d87bdb-d81b-4287-bd93-3599c098793b · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:26:02.751752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:26:02.751752Z digest=sha256:98ee0717ca4f5d17bd55b3adb33d076c26ff5b7a9c53fc8df0f354f95bf9e0ab

Observation d9f1481b-f470-4768-be29-cc5edd664e29 · outbound

This paper cites Crowdsourcing Multiple Choice Science Questions.

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation Crowdsourcing Multiple Choice Science Questions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:26:02.758615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:26:02.758615Z digest=sha256:4ecf4ec322a169b4a00ce2afcfe1860eea89b7bfdbd5d5f00472e2eabda96c05

Observation 09f2343c-70f9-4924-a5ac-bca3add8e83a · outbound

This paper cites write newline.

SciFaultyQA: Benchmarking LLMs on Faulty Science Question Detection with a GAN-Inspired Approach to Synthetic Dataset Generation write newline

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:26:02.765845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:26:02.765845Z digest=sha256:11364a786c81dddbfc7b66f3a96a0f43fd9d1a50a2005a6c43ef8b83cc635b78

Pith citing papers

No inbound Pith citation observations are available.