Pith. sign in

Paper Citation Record · LEDGER

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

As of 10 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 5 inbound Pith citation observations for arXiv:2507.00699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00699 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:16:42.794738Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T08:17:10.481202Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact3
  • verified fuzzy7
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 98959035-344d-4dd5-b524-8cc2a7ddf776 · outbound

This paper cites Soen-101: Code generation by emulating software process models using large language model agents,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Soen-101: Code generation by emulating software process models using large language model agents,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.649612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.649612Z digest=sha256:02def030193613798a3504204429a21218e91d70b6feea5553f5c4b2f4bb9d3c

Observation d930fc1b-9079-4376-9294-e8d94752079b · outbound

This paper cites ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.840152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.840152Z digest=sha256:8c256c5746c1c9902c3a4e5b0e607fb28d565b7f55038add2fe224544016dad4

Observation 815bdd69-f192-49d1-a45f-c0730df2b1fd · outbound

This paper cites Skcoder: A sketch-based approach for automatic code generation,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Skcoder: A sketch-based approach for automatic code generation,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.912600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.912600Z digest=sha256:9de5a5e789ed450fc6725f43829da256f502b711158b983d1fd7fb3325ade0b3

Observation 05f51ad5-7b73-48d2-af70-60563fdb60ed · outbound

This paper cites Enhancing Code Generation via Bidirectional Comment-Level Mutual Grounding.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Enhancing Code Generation via Bidirectional Comment-Level Mutual Grounding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.926501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.926501Z digest=sha256:2200a9c4a8af1fbc3b2b5f8d3628a44e304218407750ba4c32a0fbf4b12ab8b7

Observation e0d3c8c6-b55a-49ef-ad5e-effa4a22987b · outbound

This paper cites Fixing Large Language Models' Specification Misunderstanding for Better Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Fixing Large Language Models' Specification Misunderstanding for Better Code Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.068895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.068895Z digest=sha256:3f74c39c3f3cebb561f0179e740b100f83a4a12d48714adb9f0a8053872add13

Observation b61f1b99-39fe-4665-b59b-aa73c79502a7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.283330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.283330Z digest=sha256:456a109d01250b25cdf40f032f80c4f081fe43d5c1d4e13f438496125eadc4ae

Observation b47c095f-2342-44b3-9824-95727e3a39bc · outbound

This paper cites Program Synthesis with Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Program Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.360449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.360449Z digest=sha256:b94cafdb27f4d9da77c5d962a9306856a1bef15c0bb987b1762bfd0ac5a4ea49

Observation 3bc9bdaa-f0df-4836-91cb-e64f7b267922 · outbound

This paper cites A Survey on Evaluating Large Language Models in Code Generation Tasks.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.365515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.365515Z digest=sha256:3ebee26768be8567243534a94f3290acadd4a6b61003502bea15587980f456f7

Observation 43c39c9e-b5aa-4659-bec4-0e0753fedd37 · outbound

This paper cites Instruction-following evaluation for large language models,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Instruction-following evaluation for large language models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.453046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.453046Z digest=sha256:a54e8ae42b0d7304f3f1ab541c793ced5dcf8ae68c26fb87750033267935699a

Observation da751ba6-c1d1-47df-985f-5a1f713dda1c · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.489193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.489193Z digest=sha256:88cdd6af437b5979787ca31bcc681e499881710af64da3c5f87192c0bc2f47a1

Observation f1774027-b0c4-4f17-81a5-df9a12cc59c7 · outbound

This paper cites CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.524744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.524744Z digest=sha256:e9c88c8ac4387ce33686e89b51a95f6cfb64a66799cebb8ffd653170c6bb1a90

Observation 1193c980-a560-4bd8-85e2-898adf58dbbe · outbound

This paper cites Codeif-bench: Evaluating instruction-following capabilities of large language models in interactive code generation,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Codeif-bench: Evaluating instruction-following capabilities of large language models in interactive code generation,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.559939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.559939Z digest=sha256:1670ae856b3f93024912361717eb15802ae2f60d8c73299b741f20b96c61c48d

Observation 79701b5d-baf1-4604-9754-541bf883a75a · outbound

This paper cites A hierarchical and evolvable benchmark for fine-grained code instruction following with multi-turn feedback,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback A hierarchical and evolvable benchmark for fine-grained code instruction following with multi-turn feedback,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:48.065482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:16:41.648086Z digest=sha256:713118fdb13125a49c6ad8813415c9c85fcc57ce72cb53ca266b15fc33af08a3

Observation ea3e0791-063d-489c-a370-71143a52ca3b · outbound

This paper cites Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.671896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.671896Z digest=sha256:fb4ce93bece7549dbbfb26c0995aa015359dd8247a50e54768cdbf4fb81ffe84

Observation cea143ab-1230-4175-9893-9bb48f9b74f8 · outbound

This paper cites FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.694748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.694748Z digest=sha256:51afa8d3eb18f056fd29f488c7e67e398199e29dff0e33e93a4ac944a58a41c9

Observation bb171d85-fe5a-43b1-837c-f0361ca7f398 · outbound

This paper cites Sharegpt,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Sharegpt,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:47.895647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:16:41.734747Z digest=sha256:afd9dd2d6d879e905604cc746bfc52fffaab35d7217922158ca910d337b5361f

Observation f6eabef0-6a5c-4206-bbcc-a25731a628f8 · outbound

This paper cites Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:16:45.463009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:16:41.784746Z digest=sha256:4128c1528f0a0f4a8a9696cad8508fd599e00c93ee669562a5a615dbc922660e

Observation 1e647956-e148-494d-8d08-dc3e467a846e · outbound

This paper cites tree-sitter/tree-sitter: v0.25.5,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback tree-sitter/tree-sitter: v0.25.5,

Reference 22

Resolution
verified exact
doi, observed 2026-08-06T21:16:43.134792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:16:41.819239Z digest=sha256:ac3b9ae66ed5eb3a2bbb6b271674a0456bb132a341e45a17e7bcda141f730a89

Observation a19eca15-2d4a-4f3a-9595-400190df6f7e · outbound

This paper cites GPT-4 Technical Report.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.874749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.874749Z digest=sha256:a1cc2211ff3efedecff56097a3a8cd4c2a613eaff8d2078d2d0aba840746e187

Observation e6c825af-d137-4486-9936-ea7107ce2302 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback ROUGE: A package for automatic evaluation of summaries,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:47.604770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:16:41.914739Z digest=sha256:1db5e954693bf881ad949142743cc15cafba9781d946d7142c72fca8cfde97bb

Observation 5a0c9c0b-bf7f-4272-9ae4-e7c2665ef132 · outbound

This paper cites GPT-4o System Card.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.937129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.937129Z digest=sha256:4a8fe42033ae81a681189854d405d331a092c2d714984559228d29c18928378a

Observation 26e3e459-d30c-460a-a7ed-295320b0c912 · outbound

This paper cites Claude 3.7 sonnet and claude code,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Claude 3.7 sonnet and claude code,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:47.186490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:16:41.963659Z digest=sha256:fbdd486a868abace58e3bfadebaf0ef6a448ccb540da61cb4980cb2a9a8f3ca3

Observation 84baed67-510b-43ca-b183-a206c5ce0c6f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.994749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.994749Z digest=sha256:d3c2ce8468167723c86c16a8273b3371e1b8de0342f46711e7112b9bdedde531

Observation 8abd2f57-2fd9-40f2-8480-f908485b89dd · outbound

This paper cites DeepSeek-V3 Technical Report.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DeepSeek-V3 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.034748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.034748Z digest=sha256:982a43c9f4b2a87c7a62804c522d6cd1cb9a64b4bb876434d67350405bb29cea

Observation 7a297632-837c-4105-8ad7-12f0e53f9b6d · outbound

This paper cites Qwen3 Technical Report.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Qwen3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.094905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.094905Z digest=sha256:b121a31f2b013ac1f250932ce9b9f89e27edc44792c2048829d6e09e8de071d4

Observation 1927e71d-37ba-4f26-8e8d-3d54415b000e · outbound

This paper cites The Llama 3 Herd of Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback The Llama 3 Herd of Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.134745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.134745Z digest=sha256:7ae786107f60a8966f1466dd9d5c2046bea8096d9e842a37294940a6d7759d63

Observation 37067ce1-3533-4b5b-ac4f-c3abc6105d35 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Efficient memory management for large language model serving with pagedattention,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.174750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.174750Z digest=sha256:201c6d4f1940218220e2fc73f961b36d4aa1c77c6bf3a1c3c4f65d1335ee6a21

Observation 984bcaa4-5042-4ed1-8926-f2ea96675d1d · outbound

This paper cites Guesslang: A neural network to guess the programming language from code snippet,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Guesslang: A neural network to guess the programming language from code snippet,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:46.819549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:16:42.225023Z digest=sha256:4aff534f9436ef85c4d65948db24a6bf12712b839e9ffcdd6d074236316eded4

Observation 154ca5db-694d-4241-b4e1-f0df76995d2e · outbound

This paper cites DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.274834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.274834Z digest=sha256:27725f020df4bb8dd6c38e57d1fb5ab25fe138a920e1d1321ad563b1e87892cd

Observation a938fd60-4e67-407d-864f-5b65fe7c3f08 · outbound

This paper cites ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.308763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.308763Z digest=sha256:d0fd4969fc9e934d0f18edb94640963b434101dff3549693bf0dbf37e7aac5ea

Observation 72cdc08a-3578-47b8-8024-0aced3a8923c · outbound

This paper cites Learning-based widget matching for migrating gui test cases,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Learning-based widget matching for migrating gui test cases,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.314114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.314114Z digest=sha256:3df090a0510461721411f0f22a3fb9d433126ca210270bbc742901a2f3dfd4f7

Observation e6008267-b874-4237-a2fc-97efca23c204 · outbound

This paper cites DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.331104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.331104Z digest=sha256:8dbdebb060cd4eb148c0ea4713b8465e5d2167f6b4cdfe3e970cb497e972e780

Observation 459eb448-1417-48ee-bb16-61b7931bcb71 · outbound

This paper cites RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.334629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.334629Z digest=sha256:8283a3849d68c126574406f3f054ed2e9593200ba4d3b3bd9eed56be1808f9b8

Observation b0b49b14-e3ac-4ed8-9abc-f6418ac8b6db · outbound

This paper cites Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.374745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.374745Z digest=sha256:f21e685804e16518c0b834c7c5497f27a6fed9a327fa9a453964e9882c0151e8

Observation a95a8c12-ebd9-42f2-bcf9-72e7ba49ac5e · outbound

This paper cites Benchmarking complex instruction-following with multiple constraints composition,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Benchmarking complex instruction-following with multiple constraints composition,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:46.644479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:16:42.424811Z digest=sha256:320b6a4647c939d7771783264b1d364d4aa2103189b10f955574ba713bb0ff1c

Observation 472da2bc-29fc-4846-9f68-71cf6927d5ff · outbound

This paper cites Generating Equivalent Representations of Code By A Self-Reflection Approach.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Generating Equivalent Representations of Code By A Self-Reflection Approach

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:16:44.004826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:16:42.503377Z digest=sha256:8faf0617a18037109eb2af31fdf671790b6edd5693d56043e6184434ce0b8033

Observation 674a015e-81df-4cc0-8811-24c3c0e410d4 · outbound

This paper cites Codescore: Evaluating code generation by learning code execution,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Codescore: Evaluating code generation by learning code execution,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.546851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.546851Z digest=sha256:0edc56c06be703eac4ec6e2134605869c7ea8165433d6fa2aff73b60432d9078

Observation d0639506-78c1-4147-a23d-9a7d2faeb7e1 · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.464751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.464751Z digest=sha256:d803874b99b13418cd7d01f04a8d44ecae4b75f3b28fead9031e2381b6e90c0d

Observation 40aa013e-3916-4b50-84d9-62400e7e56f0 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.626496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.626496Z digest=sha256:5025f77b8beb1e79fe441895f8d9f6fa9c24369cfcbc4bd2d3972ca6bbd80476

Observation ca44739c-0b3f-4e91-aadd-61e8e4bc0971 · outbound

This paper cites Magicoder: Empowering Code Generation with OSS-Instruct.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Magicoder: Empowering Code Generation with OSS-Instruct

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.654739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.654739Z digest=sha256:a62eaa1b03c33e95df857617d769a540cf791f5ab90584d9debca6faaaa3ab43

Observation 7d317827-9bbe-47f6-a99c-de1dfb3fb9b5 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.591153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.591153Z digest=sha256:10dc7ca952722c5543779c8d2d8414c3a9209ec1978425993e3471a3dd57591e

Observation 57834f6e-a9ec-4a0c-8657-fbc15a330194 · outbound

This paper cites Genetic instruct: Scaling up synthetic generation of coding instructions for large language models,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Genetic instruct: Scaling up synthetic generation of coding instructions for large language models,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:46.608394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:16:42.750003Z digest=sha256:3b595e3623279df47c0c864cfb6dec6ee302dd0eeb6ab2c879f0400bbffa29d6

Observation 9b693b4d-0478-4898-8651-a21d4f48dba9 · outbound

This paper cites WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.694742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.694742Z digest=sha256:6a92f35dafe7ff0f780670ee2e782487a7208e2481ffd83948d9cadbb13577ce

Observation 54853ea3-73e6-4c99-80e1-4982df7be0bd · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Instruction-Following Evaluation for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.484523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.484523Z digest=sha256:be0ef66c50c0592ab631bae2dcff31b7cbf706cd2afdd7c38322c0d36ff7c6fc

Observation 918c7b8c-6c40-4878-8711-7b74779eb127 · outbound

This paper cites SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.724767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.724767Z digest=sha256:a0662a1b6f9a6b5e423fa63d9db137ded95622705e44ebb4d1618c894a2f7a13

Observation eabb4b14-24c4-49af-a418-bedcec7e29c1 · outbound

This paper cites Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.794738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.794738Z digest=sha256:28b211917c88e2afe4dbbe2e52e832d7cf7829dad801aefd694ac8141263364c

Pith citing papers

Observation 39c86638-d467-4dad-9eba-bfbb7c127715 · inbound

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code cites this paper.

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:11.028874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:37:51.790000Z digest=sha256:4dafd4f3a00ed5536a9b646781efb6f3176edcdd9875315b2f9ebda7217f3cbf

Observation 290c59ff-805d-4373-ba60-59b11e8a9b3e · inbound

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models cites this paper.

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.976976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T07:17:07.720589Z digest=sha256:aa45be47ae9d37a28798918eca3f0d5539c891327d9eba51167b802df0e56d35

Observation 57e95625-d60e-4a7b-a963-890c4f78eb3d · inbound

Prompt Governance? On Governing Technologies Governed by Natural Language cites this paper.

Prompt Governance? On Governing Technologies Governed by Natural Language A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:25:32.385646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T08:17:10.481202Z digest=sha256:56e81fefd8a6373c6bd12804bf3af9a0dde5635d724dfbb97106fc0b6dd49fbb

Observation c805b868-143f-4aaf-89c6-f46868efae58 · inbound

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues cites this paper.

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:06.984330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T20:22:29.287534Z digest=sha256:e4d1b07523fe9fe63bf74b18300501093acf60a37443d089dac8626282b90845

Observation 9f50b92b-d06f-40a4-a0b3-2d24dcfef8ae · inbound

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues cites this paper.

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.347673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:05:50.811872Z digest=sha256:29c2b99625af696f0cbd1de352432407e8097cbe430b041e13c56bbcb7d331ee