Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2406.11020.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:36:46.216027Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T19:07:17.629184Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 6966cf7e-bc05-4e2f-82bd-9e14df86f8a6 · inbound
Generative Agents for Multi-Agent Autoformalization of Interaction Scenarios RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aeaf685-397b-4fd3-a57d-7c85cfd8e104 · inbound
Investigating the Robustness of Deductive Reasoning with Large Language Models RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa4dcf69-077d-49ba-8cdc-3ef7994a1ff0 · inbound
Statistical Runtime Verification for LLMs via Robustness Estimation RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ce7da9e-3f97-470c-a0c5-91b59f47040e · inbound
A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 181
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb2b3ddc-6b9f-499f-924c-71350f93ab6e · inbound
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed6e48bb-5ca5-437e-bea2-6117ceec8fb8 · inbound
Towards a Science of AI Agent Reliability RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 786d5343-04c4-40a3-87fe-ee41f2fe97df · inbound
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 632b6b7c-1189-41fa-a37c-072bbc04f0e9 · inbound
GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d13b6aa6-beb6-4b4b-80fb-57dd5289f287 · inbound
Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2b5dec2f-bf4f-42e9-a639-87bacd7473c5 · inbound
When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 32d1364f-0811-4c0b-b22c-bc37564cea51 · inbound
Implicit Reasoning Steering via Concept Chaining RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 525f9cf0-1081-4e5c-9915-8614244d99de · inbound
Implicit Reasoning Steering via Concept Chaining RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 173
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65bb48e2-b5b4-44b8-8708-0e44caf57646 · inbound
Understanding the Impact of Linguistic Realization Choices on LLM Stance with Causal Tracing RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbcbfa2d-e625-4dfc-81af-1e632d639675 · inbound
Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66ad5cfa-ba6c-4d63-835b-bfe1016bfa12 · inbound
Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b58994f9-d4e6-4c6d-94d9-884a55ac8053 · inbound
Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 311eb501-d071-4230-9228-4c7e3f35a4eb · inbound
Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks RUPBench: Benchmarking Reasoning Under Perturbations for Robustness Evaluation in Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.