Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Reasoning Robustness in Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2503.04550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.04550 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:52:41.500608Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T16:53:40.544660Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ece21f9f-633a-43c6-9300-958f80c86e96 · inbound

Wireless Agentic AI with Retrieval-Augmented Multimodal Semantic Perception cites this paper.

Wireless Agentic AI with Retrieval-Augmented Multimodal Semantic Perception Benchmarking Reasoning Robustness in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:41.500608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:41.500608Z digest=sha256:e6914b2e4e7e9faa897f6aa850e070612926a570c3e13de68caee72b619bcfc4

Observation addce766-70cb-442f-9db5-af65c27dccac · inbound

Bayesian Social Deduction with Graph-Informed Language Models cites this paper.

Bayesian Social Deduction with Graph-Informed Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:32:09.373274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T07:27:20.027179Z digest=sha256:f13f723172d5e8d2d3cfd8bef7489d7d387ac3cabb50e3ab4cc7f01ca2480c21

Observation a99acdcf-a909-44f6-afdf-ec0a1265bd8d · inbound

A Large Language Model-Empowered Agent for Reliable and Robust Structural Analysis cites this paper.

A Large Language Model-Empowered Agent for Reliable and Robust Structural Analysis Benchmarking Reasoning Robustness in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.584043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:21:11.584043Z digest=sha256:0ac2e7bfbde53e405d62a25a2cf443276b1f735166ff515eea120c0c0c4a10b1

Observation 3b6d29b8-3d4f-40aa-a26c-e2931ad03d7c · inbound

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies cites this paper.

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies Benchmarking Reasoning Robustness in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:37.741861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:58:37.741861Z digest=sha256:dc8bcedc0c53a23eacedba3d36e40d913d3152767feb33a1bb5b8be50eae474d

Observation 343f7b8a-0f47-48cc-a7ea-b84176a34a89 · inbound

Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need cites this paper.

Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need Benchmarking Reasoning Robustness in Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:41.606874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:41.606874Z digest=sha256:9c417a5888ccb4246c4d747ece39e83c9bc8194064eea7381c33243b243e878a

Observation ab2b6813-c8b3-4b9b-b62b-dc614589eade · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:10.164969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:10.164969Z digest=sha256:61314afa611110ec779b8ed4feb4c9827479b687478d74951072d3813c9e8c11

Observation 49485f00-064e-4416-8d89-69c43dac6161 · inbound

Throttling Web Agents Using Reasoning Gates cites this paper.

Throttling Web Agents Using Reasoning Gates Benchmarking Reasoning Robustness in Large Language Models

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:05.016955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:05.016955Z digest=sha256:c623cef776432491e622ef5786aea80f537f70a12460212bb27269d8cb17feaa

Observation e460ba93-291d-45bf-97f4-e0a883657636 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:07.143764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:07.143764Z digest=sha256:3bf873020713450f6df785c728cde29af54377927bdd8eb6d8de42373f73dedc

Observation 21109c5e-c722-4689-b483-84b58368f1a7 · inbound

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation cites this paper.

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation Benchmarking Reasoning Robustness in Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:20:06.816253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T12:18:18.407788Z digest=sha256:aa171a7bdf5e1f9c37d88b022ed08ba7d25ffebb7d0ad695ec32b413aa614ccb

Observation 985e2736-6776-4983-a8be-cc8b35157f77 · inbound

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning cites this paper.

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning Benchmarking Reasoning Robustness in Large Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:22:01.794846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T01:19:49.761472Z digest=sha256:6b101b67e3ccd079ba372ccc942869c8253c8271981d63ad64c09cf6c2d7354a

Observation 6c9f0343-f6fb-446a-a119-2f5a828c2255 · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments Benchmarking Reasoning Robustness in Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.546212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:b4d3edae9b326016ca57a8df58faa3178a4de4d6bfa155bfc96de5c429a8a2f1

Observation 827760a3-26cf-493a-a2ca-eebe1f86bf14 · inbound

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models cites this paper.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:55.417781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:55.417781Z digest=sha256:8d60ffb66a9fa0e29ae6885b2f2f60d7d23196eb99a542681f8f9808a61e9f7b

Observation 07a6f5f0-324f-4787-9690-08dd164475b6 · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining Benchmarking Reasoning Robustness in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:22.985400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:22.985400Z digest=sha256:0b2f3d2daf3b8cdb765c12017a09558dab6841426ae5bf7d8b71c2794ed23350

Observation 81fb91db-cea3-4b1d-af33-11f627ce29fd · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining Benchmarking Reasoning Robustness in Large Language Models

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:37.046778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:37.046778Z digest=sha256:e5519061295ab90a0f0183d5dd762d03fe523bea10629f0b2a7b523aa484ffa3