Pith. sign in

Paper Citation Record · LEDGER

Reasoning Capabilities of Large Language Models on Dynamic Tasks

As of 17 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2505.10543.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10543 v2

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:12:01.571261Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:36:44.033596Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:25:46.619265Z

Reference resolution

23 of 23 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 775d0281-04ef-4006-b35f-d0d7237facb2 · outbound

This paper cites Language mod- els are few-shot learners,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Language mod- els are few-shot learners,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.466773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.466773Z digest=sha256:daef8e80725ee13c6823738fe6bb7375dcea4f3d567bc2f11ca5b30206c09a05

Observation ca51aa6e-73d8-4458-87c9-cb654602f008 · outbound

This paper cites Reflex- ion: Language agents with verbal reinforcement learning,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Reflex- ion: Language agents with verbal reinforcement learning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.471784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.471784Z digest=sha256:23390a5be527eb2259d3750d0d542aca760f430cdb2f562e224a1a74d3042319

Observation 5d6b2e2a-80bf-4586-a364-39ea9594b2ec · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Reasoning Capabilities of Large Language Models on Dynamic Tasks ReAct: Synergizing Reasoning and Acting in Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.477653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.477653Z digest=sha256:306a346935196a2fc72f8590402fd0873cf045bca8a9e47897a1368a0ee5993a

Observation 0ff6ec71-be91-42a3-8f17-8315825a5e20 · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

Reasoning Capabilities of Large Language Models on Dynamic Tasks The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.483633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.483633Z digest=sha256:80437c224ad7e36b506302d65d48d8d7a325a441639d0a4617eaa5bc92c2245f

Observation 397d00e7-3ba5-429e-a182-d106554c23df · outbound

This paper cites Text-based games as a challenging benchmark for large language models,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Text-based games as a challenging benchmark for large language models,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:01.979481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:12:01.488755Z digest=sha256:0cc665e2f9c8cc93526eb43c4b8bfd93be0870be72a80dc9095980d9683254c7

Observation 1f7aa33f-e1db-4eb1-b571-debb7e5e8640 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Training language models to follow instructions with human feedback,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.493247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.493247Z digest=sha256:022e56836453d5926c925bc99e3ac049b72c8728e072c0866e5d0d3a63bfc232

Observation ba4a0d14-a25e-4426-9429-0b67fc784714 · outbound

This paper cites Prompt programming for large language models: Beyond the few-shot paradigm,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Prompt programming for large language models: Beyond the few-shot paradigm,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.498450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.498450Z digest=sha256:54b4f9bab9e34fb27d3fb77cd01b668f0face6b0dfb90d8a0ec10bde5eaae38e

Observation 564526f6-4b44-4d2e-9eb5-c1230ed93105 · outbound

This paper cites SmartPlay: A Benchmark for LLMs as Intelligent Agents.

Reasoning Capabilities of Large Language Models on Dynamic Tasks SmartPlay: A Benchmark for LLMs as Intelligent Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.502630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.502630Z digest=sha256:8bd2af66b5bc7319b1c8a7bd27a3cdee825dc4d955c5d4f550d26da02259789c

Observation ac619244-a0bb-40b6-a592-7810617ec795 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Chain-of-thought prompting elicits reasoning in large language models,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.507211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.507211Z digest=sha256:74448cb89544eea99e2a96264e8b6fb676b4ce9210cb411d54a40b52ca7a971c

Observation 50feaa8f-be13-44e6-a508-f20848d88d13 · outbound

This paper cites Self-refine: Iter- ative refinement with self-feedback,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Self-refine: Iter- ative refinement with self-feedback,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.511339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.511339Z digest=sha256:7709ee012b553fff8239a16a04080c2e065e7caa6f95f4a0bed838d91d9034a5

Observation 0259a033-897e-4c0e-8a1c-2a357cedce90 · outbound

This paper cites Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Agent-Pro: Learning to Evolve via Policy-Level Reflection and Optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.515887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.515887Z digest=sha256:bf33628b1e1c17c858e4b66003637937dd2ad41b7b9242d61e3c7fc531213448

Observation b6cafd80-2a29-4413-8728-be1593027a34 · outbound

This paper cites AutoPlan: Automatic Planning of Interactive Decision-Making Tasks With Large Language Models.

Reasoning Capabilities of Large Language Models on Dynamic Tasks AutoPlan: Automatic Planning of Interactive Decision-Making Tasks With Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:12:01.733559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:12:01.520534Z digest=sha256:d885d03b4a3c5b6b5c3fd6dd559b70b3a1c77108f312476479e51c978de9c8e9

Observation 377af352-bb2d-4d06-9771-d332ad7d3e72 · outbound

This paper cites Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.524931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.524931Z digest=sha256:8b7cb6818812b0c83ec279288de497c083c6774cf93c512eb7ecd5cde77abdc0

Observation 4b535f04-ce16-415b-bc38-e7c19f0013d8 · outbound

This paper cites Mental Modeling of Reinforcement Learning Agents by Language Models.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Mental Modeling of Reinforcement Learning Agents by Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.529561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.529561Z digest=sha256:05779d3f9ab44e5b2b7af102682afb80195ea05709a463e3cbf5891becff6a16

Observation 62adbb95-fe10-4f37-be0c-6b79d85bdc0a · outbound

This paper cites EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers.

Reasoning Capabilities of Large Language Models on Dynamic Tasks EvoPrompt: Connecting LLMs with Evolutionary Algorithms Yields Powerful Prompt Optimizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.534625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.534625Z digest=sha256:ba50fcb8cea4e0aad495ca9d731abd10c6f4b75ad4d6598d063b06acd41ae664

Observation db7d2981-7220-4eb8-b4f8-d329ee8e128b · outbound

This paper cites Llamea: A large language model evolutionary algorithm for automatically generating metaheuristics,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Llamea: A large language model evolutionary algorithm for automatically generating metaheuristics,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:01.925131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:12:01.539176Z digest=sha256:0238705c2c2607d6b4d303f739a8accefe4b093883bc9049a9d81890dd4a65f4

Observation 9d576973-9a70-4cc9-9ba4-d9da422eac3b · outbound

This paper cites Focused transformer: Contrastive training for context scaling,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Focused transformer: Contrastive training for context scaling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:01.909695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:12:01.543682Z digest=sha256:214d96c56ce806ed4e98a110acd8d8495645ee0b9786f20a08bb9c1b531c1b68

Observation e1a413d7-8a89-420c-928b-dfd617713924 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Lost in the Middle: How Language Models Use Long Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.547721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.547721Z digest=sha256:ab7268b2252f09a6e19e7bb237d7141b0c9472cf4748461698b007539572ad53

Observation 528de17c-6008-4c33-b393-b4b9f20b0797 · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.552813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.552813Z digest=sha256:70bb11da405b6dbbd03d40d452c95b6e1240c54c4e4ecca39ed78f45d8a78f57

Observation 937adfb6-0f8f-40dc-aed3-cc8084c7ca7d · outbound

This paper cites Chain of thought- lessness? an analysis of cot in planning,.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Chain of thought- lessness? an analysis of cot in planning,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:12:01.819066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T21:12:01.557485Z digest=sha256:ffe5bccf274664bc37d1c8e3ec3f721896a80de3068750e8561fd589f484ed4d

Observation 6b30c080-4519-49ea-ae4d-2a4feb7e57d4 · outbound

This paper cites Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?.

Reasoning Capabilities of Large Language Models on Dynamic Tasks Missing Premise exacerbates Overthinking: Are Reasoning Models losing Critical Thinking Skill?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.562040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.562040Z digest=sha256:74d849c2cf02322abbcdc6281bd4cb76266caa9a001ea1bc3fcef9e974fe2765

Observation a1c75472-7ada-44f0-b4fb-ed9f59678d52 · outbound

This paper cites HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering.

Reasoning Capabilities of Large Language Models on Dynamic Tasks HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.566767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.566767Z digest=sha256:ed871797a21aa1bdbb2e5b61c48316589151c40c45095d0af3bd33ff68375813

Observation f0664dfa-eb55-42ae-9c61-6d5a51f64d30 · outbound

This paper cites BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games.

Reasoning Capabilities of Large Language Models on Dynamic Tasks BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:12:01.571261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:12:01.571261Z digest=sha256:673ee4bac4e49bf12eec690bc46a30a08e8eeee621164e45e8fe1b739efaa780

Pith citing papers

Observation 82031c4c-49ec-4ff8-99a9-b8e562076058 · inbound

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform cites this paper.

Why We Need World Models for AGI: Where LLMs Fail and How World Models May Outperform Reasoning Capabilities of Large Language Models on Dynamic Tasks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:25:46.621005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T21:36:44.033596Z digest=sha256:797bde81a06a31828e7d96b2c28491113128333da7ae228aef7d1443b07a09fb