Pith. sign in

Paper Citation Record · LEDGER

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories

As of 10 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2502.07747.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07747 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:44:59.786646Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:57:23.282052Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T05:46:09.807404Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39b2c24b-2545-4049-b454-603a1a4f1cb4 · outbound

This paper cites GPT-4 Technical Report.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.700476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.700476Z digest=sha256:730a629e9e78db015186b38003d701afffb5111f7227e3d88b958fe36ef828c7

Observation d5e90bf2-a2d1-41c1-943e-4a42362d8876 · outbound

This paper cites an unresolved cited work.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T11:45:00.069383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T11:44:59.706172Z digest=sha256:d16c53a208edb8d517958f1a734f5ddbb2b5814b38f00522b2d57418363288f6

Observation 50b29953-3905-433e-9a69-512a0e5c8c4e · outbound

This paper cites Language Models are Few-Shot Learners.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Language Models are Few-Shot Learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.710993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.710993Z digest=sha256:1e3092bf60e3686fd5c11dfba0bc3b08f6239b5473bdb20354f6d938a0904760

Observation 010dfd47-a8f6-406d-bebb-f6b29bc53777 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Measuring Massive Multitask Language Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.715986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.715986Z digest=sha256:e8e075358762103281ec25b800b26d1d17ff5680786531b53e5f4b1ac139f565

Observation 091c956d-2a2b-46be-bd39-1828c1575ef4 · outbound

This paper cites an unresolved cited work.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T11:45:00.050955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T11:44:59.722751Z digest=sha256:ba0523afa7388f68270b1c7dd0dfe4f4991e4a1c67d0d632d3bb0eb74e0d62e0

Observation 600c0ef7-64ef-43bd-9648-089d3d84fe33 · outbound

This paper cites an unresolved cited work.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.727737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.727737Z digest=sha256:74884235be8658f48295e727a0b639564376f9527d189c72b0aa1c8e08e8ea9e

Observation 6ccb9e5d-77dc-4605-97b2-3ff4ae39264b · outbound

This paper cites Holistic Evaluation of Language Models.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Holistic Evaluation of Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.733056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.733056Z digest=sha256:8923ea322a925a979d5068a6b17ef35ea7a5d6e145a8d3c33347438c055413e0

Observation 8f06e9fa-0bdd-4c18-bd1b-7e6a95367a23 · outbound

This paper cites an unresolved cited work.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.738022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.738022Z digest=sha256:2c8b6423d2367f8aba6ad39fac1858428a9dfdfddd0fde046c650aab8d3961a7

Observation f9df960a-91e3-40da-8691-2bce538e669e · outbound

This paper cites an unresolved cited work.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-08T11:44:59.999161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T11:44:59.742707Z digest=sha256:04ecbc4385ee5a51b1574c6225eeee05c25049d433163bcb085be3d4be60bc59

Observation ea275fd5-fdff-4702-9748-9a4287871438 · outbound

This paper cites an unresolved cited work.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.747533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.747533Z digest=sha256:93074f6a45f7fc56eef007700cbd0bad9eef8b04798524292b0252aaac5f4593

Observation a0f75bec-e159-4903-a39a-d7db722d934f · outbound

This paper cites PlotMachines: Outline-Conditioned Generation with Dynamic Plot State Tracking.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories PlotMachines: Outline-Conditioned Generation with Dynamic Plot State Tracking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.752378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.752378Z digest=sha256:0dd903db1f421f3e0fc7d04c589121224abb29b7cd4ea3448b6f0b9edb9ae49c

Observation 0f087741-f23f-44b2-91d8-e6093f350b0f · outbound

This paper cites an unresolved cited work.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.757427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.757427Z digest=sha256:b7184035a2dae6b60b4933d18ffc43a477518a444b24d6b2c72a3b785f007ba1

Observation f26113ed-0664-4453-8dee-31c96fbbe350 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.762175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.762175Z digest=sha256:aced46d29000956aaea9a0bc384dd8b9c002d837f36daae699399d6d7f8e22bc

Observation 494b9779-6d1e-48d0-8252-c4273c606e7a · outbound

This paper cites an unresolved cited work.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.767095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.767095Z digest=sha256:02628fb8d3b51c5e8fc69d772d93b371f8cb45cc4dcb189d1245ff66b3473d15

Observation b1137e2f-cc6c-44c3-9429-dd16f6a4aebf · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.771725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.771725Z digest=sha256:613c06ee1124d52756d98abd6f5531b3eed550d8fd3da265cd1cab056b8a1d3d

Observation db0776ab-3549-4225-b694-1d124fb1f7d8 · outbound

This paper cites an unresolved cited work.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.776792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.776792Z digest=sha256:b8e939512c5ded8a1f41fad2076f79ff9f88cce3136c126b6565ccdb9914276b

Observation 92941092-748c-4246-9086-a0422cb9996c · outbound

This paper cites online" 'onlinestring :=.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories online" 'onlinestring :=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.781345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.781345Z digest=sha256:bae481250c7746e4ce1297162b235005d3a2e83e4f099bb05d98144f3ed065a4

Observation 3c46f2b9-10f3-40b8-bfb5-19a543093cba · outbound

This paper cites write newline.

WHODUNIT: Evaluation benchmark for culprit detection in mystery stories write newline

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:59.786646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:59.786646Z digest=sha256:e4aa08e9948956ef90ed01d17b6e1cfb7ff0938aaadc7dfdbd2e2653793cea1d

Pith citing papers

Observation 5195f490-acfd-420c-9c88-98dbe2a01f82 · inbound

Wiring the 'Why': A Unified Taxonomy and Survey of Abductive Reasoning in LLMs cites this paper.

Wiring the 'Why': A Unified Taxonomy and Survey of Abductive Reasoning in LLMs WHODUNIT: Evaluation benchmark for culprit detection in mystery stories

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:46:09.877297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T17:57:23.282052Z digest=sha256:9b69ad39d1f63466161e7b8e343dc9dc3f842eeceb6454cc0e0a14bd32884f60