Pith. sign in

Paper Citation Record · LEDGER

FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2310.20410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.20410 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:25:31.573876Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:59:44.890478Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 84047906-9d58-4a8b-b19f-970db3f7ca96 · inbound

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios cites this paper.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.573876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.573876Z digest=sha256:96340bb31fb0298596a42f3ef13c5ef14a2fc9e3bd7a68f777eae5c5ee990298

Observation 716ad6ad-80f5-412e-a7af-39dba8696df4 · inbound

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models cites this paper.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.605394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.605394Z digest=sha256:906c2acdfea59ecdd3a9f8ed91fee2d389a4f0b82784d9804e7d2f2c43cfc8b0

Observation dafef778-a906-4c63-a13c-4fba0027a983 · inbound

Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning cites this paper.

Gensors: Authoring Personalized Visual Sensors with Multimodal Foundation Models and Reasoning FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T14:03:07.024354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:03:07.024354Z digest=sha256:ed0ac6794be9eb0638adff6c72ec16b4cb284e75dfddd117b0d9ba2af74cb9a1

Observation cea1df16-de13-4803-9e53-5ce7018ccd4a · inbound

LCTG Bench: LLM Controlled Text Generation Benchmark cites this paper.

LCTG Bench: LLM Controlled Text Generation Benchmark FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T13:53:45.041143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T13:53:45.041143Z digest=sha256:cf67b0c7d9eede063842ad8b0e87f88f87a749edbfcf69a2439e8ba6033a111d

Observation 03b59bb3-e6c4-428e-bdeb-90f0d24ae605 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:23:31.015188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:11d33c6cb22ca053fb2b5075b5c5076bfb736f5fc6d67fdf7c83eea45166f66a

Observation 4e7e4ece-aeed-4bc9-b68a-bcab20c887d0 · inbound

LLMs can be easily Confused by Instructional Distractions cites this paper.

LLMs can be easily Confused by Instructional Distractions FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T10:50:12.328219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:50:12.328219Z digest=sha256:ebddf103b73170c754a7d5caf4b30c1ec776fad26e671ec489ccc887faf44cfa

Observation c84ff9eb-af47-4bf8-879a-c83a77da64db · inbound

Verifiable Format Control for Large Language Model Generations cites this paper.

Verifiable Format Control for Large Language Model Generations FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T22:36:35.237097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:36:35.237097Z digest=sha256:434648710035fe72837ad3235d5f22bfc132f1590a434d5c590a6668b13189fd

Observation b88fc9a4-424f-48eb-833c-b2a1a5d412fe · inbound

IHEval: Evaluating Language Models on Following the Instruction Hierarchy cites this paper.

IHEval: Evaluating Language Models on Following the Instruction Hierarchy FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:56:00.043351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:56:00.043351Z digest=sha256:f51724761714c199b2ba37cb38429009572f409e67becca63164d7a3c05faeeb

Observation 6e794753-fe77-449e-8d01-208d32f78d11 · inbound

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models cites this paper.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.940404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.940404Z digest=sha256:83c7b5775dc1deca0381ccb0a861461a2e85a6b0b5f3f97fafcd8b7be789a345

Observation 78054f9a-c2f4-46be-9c4d-6224a5874124 · inbound

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios cites this paper.

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:06.978691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:06.978691Z digest=sha256:31da3301f7f16dc57e3fc39ee11aa79104aa59cc073e43dffb0338df3848b45c

Observation d99caefb-5ec6-40dc-bab6-a222343171bf · inbound

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components cites this paper.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:20.709513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:20.709513Z digest=sha256:5a7daa24b25c7f70fc568ccfe5a2b08ce3a33a84a508755cb29e57ae1ec85de1

Observation cea143ab-1230-4175-9893-9bb48f9b74f8 · inbound

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback cites this paper.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.694748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.694748Z digest=sha256:119741927527777106f0ce73f5d184dcf9cc28186a0a4025df44ae3129b71f59

Observation 3a2b176a-8bcf-4207-b9d0-f2b39abd00f9 · inbound

How Many Instructions Can LLMs Follow at Once? cites this paper.

How Many Instructions Can LLMs Follow at Once? FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:13:47.092552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:13:47.092552Z digest=sha256:8bf07153adc32480ecf47fd99e20ac0e4927cb99fe0eca55b0d8ce64ff9f6a47

Observation a50026f2-3946-4563-a8c4-4e0bca440f0f · inbound

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting cites this paper.

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T17:39:22.252419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:39:22.252419Z digest=sha256:f3d808fbb3ef14c13242b0545d828c3168c48a9a53f0931830a32813a6baf258

Observation e5b2b43c-60c3-4bcc-b384-4614c96dfddc · inbound

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization cites this paper.

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:52:06.021759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T20:50:53.673037Z digest=sha256:c0ca27025efcf2a6eaa3e5485c41f5d09274ed612efb2c1cb6e4c2cee09f4e4d

Observation 7b785a7e-2495-4583-9e12-2a48d583ac09 · inbound

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems cites this paper.

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:02.499731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T14:56:00.449776Z digest=sha256:d430c90649f7a264a144041871ebff3230294a1c7f30639a2985b71afa6824e3

Observation d863c62e-c4c5-4539-932c-7380e5556d66 · inbound

ComplexConstraints and Beyond: Expert Rubrics for RLVR cites this paper.

ComplexConstraints and Beyond: Expert Rubrics for RLVR FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:31.354635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T16:37:11.141846Z digest=sha256:2cec5a40b0a20cf2edf078870c8b756e438d00c3a8d0e8aa3d51f09ffce4a4b2

Observation 776a9f8e-bbe3-420f-8cb0-85c8cf6d1e8c · inbound

LLM-as-Code: Agentic Programming for Agent Harness cites this paper.

LLM-as-Code: Agentic Programming for Agent Harness FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:18:44.164149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T04:16:58.964046Z digest=sha256:1cf9690228623596f0a3cd90188cdf91d182dc13cdcb6309552ef606f2a1135e

Observation a1f6e6a3-bd3a-4068-8ad4-f8288a904160 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 232

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.891870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:c1e71e74025b76f775fdc3e7032b81d1e5244dfb18e78781da2b45c54b06695c

Observation 008130e9-02a3-4996-b4dc-c1ed6818f878 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 231

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:59.626380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:f6df36d97cd1b822c334261da13dac3b3b61940cea5259d5f8b04886e2d7f27b

Observation e08387bd-acb7-4937-b3e6-d36ba320e29e · inbound

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following cites this paper.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:36:07.970657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:36:07.970657Z digest=sha256:a5f07c53179ebb358c02f2134535e72de26d1e0280e93ebe59f41ab343dd350e