Pith. sign in

Paper Citation Record · LEDGER

WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2311.15930.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.15930 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:38:06.553485Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T09:54:34.363302Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f8f0eff5-13ee-4c79-8580-69b6724a0d57 · inbound

What makes a good metric? Evaluating automatic metrics for text-to-image consistency cites this paper.

What makes a good metric? Evaluating automatic metrics for text-to-image consistency WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-11T12:38:06.553485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:38:06.553485Z digest=sha256:85c72d2d865dca6d38bc7497c471291c39dc8c229eecc8de4202d6a876ecda67

Observation d939dd53-b3f1-4aae-adff-335a7b2bd631 · inbound

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions cites this paper.

AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:08.270084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:08.270084Z digest=sha256:2256916908bce37c3ebe9463129f35030dece7c7cb423dbc9a0c063132665c29

Observation 9ee9361e-f6db-4092-97b8-f337a7bbc683 · inbound

IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments cites this paper.

IntPhys 2: Benchmarking Intuitive Physics Understanding In Complex Synthetic Environments WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:00.504029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:00.504029Z digest=sha256:8382695aae9eb2cb7b0c17503d3b9b9bedbae67d4b63ae860de977cd083015b8

Observation 87a0a974-f900-43b5-b4e8-7a7afb72d8e2 · inbound

Enhancing LLM Metacognition via Cognitive Pairwise Training cites this paper.

Enhancing LLM Metacognition via Cognitive Pairwise Training WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:33.886009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T19:01:18.153145Z digest=sha256:4610e8a9a0b31db81df35e6f8ae852d835ce8ec55de593781b9cd7cadb0344de

Observation 67e23275-1227-4c0f-9ea2-54dc0da607df · inbound

Agentic Abstention: Do Agents Know When to Stop Instead of Act? cites this paper.

Agentic Abstention: Do Agents Know When to Stop Instead of Act? WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:54:34.365047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T09:54:07.138157Z digest=sha256:73ffcf13309aea795237756ec64f34e311652a71422057864bcf9d1e719fcbe5

Observation f46d6a71-f30f-4a3e-a6e9-16b47e02561d · inbound

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models cites this paper.

MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models WorldSense: A Synthetic Benchmark for Grounded Reasoning in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T00:53:20.749426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:53:20.749426Z digest=sha256:84c2c92aa45c6e160c3e33be1af8cc0f1a57feae03fda5f5156f4031ee01c8f6