Pith. sign in

Paper Citation Record · LEDGER

FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2310.20410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.20410 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:50:12.328219Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:59:44.890478Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 03b59bb3-e6c4-428e-bdeb-90f0d24ae605 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:23:31.015188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:528067fb227353618fdb982c4b6f9c1269d62c6c4f21fd445611d79ab7468b88

Observation 4e7e4ece-aeed-4bc9-b68a-bcab20c887d0 · inbound

LLMs can be easily Confused by Instructional Distractions cites this paper.

LLMs can be easily Confused by Instructional Distractions FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T10:50:12.328219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:50:12.328219Z digest=sha256:1a9f7405023af159c495d3566aec6a937e0e25331df8151b395decdd66907033

Observation c84ff9eb-af47-4bf8-879a-c83a77da64db · inbound

Verifiable Format Control for Large Language Model Generations cites this paper.

Verifiable Format Control for Large Language Model Generations FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T22:36:35.237097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:36:35.237097Z digest=sha256:b0240c8e781dccb0ff6f4ae03d6076ea44e77c7015cf8bdefd07de7dc0bbb516

Observation b88fc9a4-424f-48eb-833c-b2a1a5d412fe · inbound

IHEval: Evaluating Language Models on Following the Instruction Hierarchy cites this paper.

IHEval: Evaluating Language Models on Following the Instruction Hierarchy FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:56:00.043351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:56:00.043351Z digest=sha256:bf2413ff704b7ccd83183c3c9410f7f8ed7cb588dea604ef218188f4bd912b3d

Observation 6e794753-fe77-449e-8d01-208d32f78d11 · inbound

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models cites this paper.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:10.940404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:10.940404Z digest=sha256:0ef32f338719d75cfdb0aa9172b5e3745ecc3826b6a4d97aa9db794e0cb4f5f7

Observation 78054f9a-c2f4-46be-9c4d-6224a5874124 · inbound

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios cites this paper.

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:06.978691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:06.978691Z digest=sha256:926307cbf89519977c675aed9d04f59aab507c39d1644e65045711b48950140e

Observation d99caefb-5ec6-40dc-bab6-a222343171bf · inbound

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components cites this paper.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:20.709513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:20.709513Z digest=sha256:2b57b9132cc5f765a3f3a437cf79bb7a3454c88558e8c618525ce7eae97b8cfa

Observation cea143ab-1230-4175-9893-9bb48f9b74f8 · inbound

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback cites this paper.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.694748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.694748Z digest=sha256:51afa8d3eb18f056fd29f488c7e67e398199e29dff0e33e93a4ac944a58a41c9

Observation 3a2b176a-8bcf-4207-b9d0-f2b39abd00f9 · inbound

How Many Instructions Can LLMs Follow at Once? cites this paper.

How Many Instructions Can LLMs Follow at Once? FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:13:47.092552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:13:47.092552Z digest=sha256:ff1179fa98b322d1a02d70f51311d93ab4d5a362aad0bd7c31f8662e65921396

Observation a50026f2-3946-4563-a8c4-4e0bca440f0f · inbound

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting cites this paper.

QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T17:39:22.252419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:39:22.252419Z digest=sha256:845ced47576bc135c0704d8fb42b2b39343e81525a0f871f18b891f7d5ada20f

Observation e5b2b43c-60c3-4bcc-b384-4614c96dfddc · inbound

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization cites this paper.

Chinese Short-Form Creative Content Generation via Explanation-Oriented Multi-Objective Optimization FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:52:06.021759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:50:53.673037Z digest=sha256:bc284a7ad8460f1aff154c3e322a527193dee106a0e922c4d1e0677036e65b2b

Observation 7b785a7e-2495-4583-9e12-2a48d583ac09 · inbound

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems cites this paper.

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:02.499731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:56:00.449776Z digest=sha256:94228fa5d98ecd8e1cffdb479136c062e201cf9627bba1746059a9e4ac581d03

Observation d863c62e-c4c5-4539-932c-7380e5556d66 · inbound

ComplexConstraints and Beyond: Expert Rubrics for RLVR cites this paper.

ComplexConstraints and Beyond: Expert Rubrics for RLVR FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:31.354635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T16:37:11.141846Z digest=sha256:0c48d8dba063b37b48f61cc29914db6464d213358bd9f21f83b940f65ca01495

Observation 776a9f8e-bbe3-420f-8cb0-85c8cf6d1e8c · inbound

LLM-as-Code: Agentic Programming for Agent Harness cites this paper.

LLM-as-Code: Agentic Programming for Agent Harness FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:18:44.164149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T04:16:58.964046Z digest=sha256:250ec41c84b536a0b868206d3ed7087a1ea9115f7db6c7627c1c042cdea80ad9

Observation a1f6e6a3-bd3a-4068-8ad4-f8288a904160 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 232

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:59:44.891870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:26829617cb378260ba1594fedc601937cd18f9f23bee365107a5710d89d5b590

Observation 008130e9-02a3-4996-b4dc-c1ed6818f878 · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 231

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T18:55:59.626380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:84117ec4979e56429221fdfe4dd4e163b7b0fe53fc44010043f90a7285f7b250

Observation e08387bd-acb7-4937-b3e6-d36ba320e29e · inbound

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following cites this paper.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:36:07.970657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:36:07.970657Z digest=sha256:116c55bb2f1158acd3d8e1e9ac0345e9466ceca2c216268b3ac988e223e057ca