Pith. sign in

Paper Citation Record · LEDGER

Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2407.08440.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.08440 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:01.512719Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T07:46:25.994614Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ff534545-f620-4513-9392-6df483194e48 · inbound

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios cites this paper.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.650877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.650877Z digest=sha256:a6ed76901ecc0945789df7954b9076017e3e3e69c4ea304c93792cc2130a44a6

Observation cd2a23dd-2983-4ffa-bfa5-071455945bde · inbound

Shuttle Between the Instructions and the Parameters of Large Language Models cites this paper.

Shuttle Between the Instructions and the Parameters of Large Language Models Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T12:41:02.371174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:41:02.371174Z digest=sha256:bfa02b99f777744d00b62faec5215a9f428d03ec0f9d47985d1a37196a115ee9

Observation 27595f1c-dc36-4ba9-820c-408fdf0f5987 · inbound

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks cites this paper.

Toward Generalizable Evaluation in the LLM Era: A Survey Beyond Benchmarks Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 278

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:01.512719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:12:01.512719Z digest=sha256:60476264d99312585a5fc30f332d1a92eb2d57c1cc7cc82d640c7d883f91c85a

Observation dee4cf25-94f2-496e-812d-a9a03724b887 · inbound

Improve Rule Retrieval and Reasoning with Self-Induction and Relevance ReEstimate cites this paper.

Improve Rule Retrieval and Reasoning with Self-Induction and Relevance ReEstimate Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:06:47.732299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:06:47.732299Z digest=sha256:a44a2d0dd27980936db3a3d69b0a739880c271b823c27e9436f6975719bb313c

Observation 954b8358-7a3d-4754-8802-5aaf54b48113 · inbound

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models cites this paper.

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:56:15.498222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:56:15.498222Z digest=sha256:7ad1883eada59b71e8ee00ed861f93198ca7e6f8e2e92351149b12349091edfc

Observation d9977b53-dbfa-4a20-80e4-e8be69543975 · inbound

LLMs for Customized Marketing Content Generation and Evaluation at Scale cites this paper.

LLMs for Customized Marketing Content Generation and Evaluation at Scale Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:04:00.296919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:04:00.296919Z digest=sha256:e2bb9c39feab016c0c4a6da19e96ec8a41a5ebf2fe876c928f1b2bb37ffc5bc6

Observation 83565a4c-80b7-444a-9126-35304835ed13 · inbound

GENIE-ASI: Generative Instruction and Executable Code for Analog Subcircuit Identification cites this paper.

GENIE-ASI: Generative Instruction and Executable Code for Analog Subcircuit Identification Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T16:57:19.179609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:57:19.179609Z digest=sha256:b808b8a825508194d332ed62bda861664765e2376b7bbbc0b7063a1d432b0693

Observation 35815867-dbe3-4cb1-b11d-5136d5061d30 · inbound

Evaluating Language Model Reasoning about Confidential Information cites this paper.

Evaluating Language Model Reasoning about Confidential Information Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T16:52:47.524750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:52:47.524750Z digest=sha256:40febbc5f879166d812226454e1be4bd23893720a398aedcca673ab37fe97f1b

Observation 6c17cd8d-4e75-40a9-ae2f-976ef6ddab5e · inbound

Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities cites this paper.

Absurd World: A Simple Yet Powerful Method to Absurdify the Real-world for Probing LLM Reasoning Capabilities Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:46:25.999930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T02:02:06.788056Z digest=sha256:eb2ea89f0f8bdfec5894c3953cc6d9d8773916edd52c46c5217c81c0cf227c48