Pith. sign in

Paper Citation Record · LEDGER

FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2311.09829.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.09829 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:21:59.082243Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T01:47:04.050716Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b707f8a9-0b69-4068-9772-b0a142164054 · inbound

LCTG Bench: LLM Controlled Text Generation Benchmark cites this paper.

LCTG Bench: LLM Controlled Text Generation Benchmark FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T13:53:45.047113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:53:45.047113Z digest=sha256:5b005cac666eefb2e9a787ca83f3c149c115eab56a48b5a7423faa599ed6f6fa

Observation 80cd3bb1-90be-4ef7-bb44-8c12aa87e8be · inbound

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models cites this paper.

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:02.097569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:02.097569Z digest=sha256:59ce036d50d5dae30d40906b3ca0e3af35b6726e97947af63adb705b3f7303c0

Observation 01b32f49-366c-49ab-8178-7a476c009360 · inbound

How Many Instructions Can LLMs Follow at Once? cites this paper.

How Many Instructions Can LLMs Follow at Once? FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:13:47.227859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:13:47.227859Z digest=sha256:d60203e4cad41ecb6f0dbfebb6a4a8ff549601b471ff1eff4625139075c0af1a

Observation 3ccc14a2-14db-4e67-8759-426f51fc79a4 · inbound

Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? cites this paper.

Inverse IFEval: Can LLMs Unlearn Stubborn Training Conventions to Follow Real Instructions? FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:16:51.156699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:16:51.156699Z digest=sha256:1a89819e5f262b76dc44c2f2e0930c3f3a010c02e0ec281093188885222bdc8e

Observation 6f08d9cb-51ea-4058-b4a9-d3b13a98c906 · inbound

ReAD: Reinforcement-Guided Capability Distillation for Large Language Models cites this paper.

ReAD: Reinforcement-Guided Capability Distillation for Large Language Models FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.053427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-13T01:45:26.316655Z digest=sha256:5c0e318a36ed5411e6454d13bb52eed6ca777d19c398746c65f4e4ebdef8b75b

Observation e1fda741-9ab2-465b-8824-813dcb2ae246 · inbound

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information cites this paper.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information FollowEval: A Multi-Dimensional Benchmark for Assessing the Instruction-Following Capability of Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.082243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.082243Z digest=sha256:67ed1e8c5e743ca0e02076f48e02020e26fa7a4f0e7dfcc7af91fc6ffb575599