Pith. sign in

Paper Citation Record · LEDGER

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

As of 8 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 5 inbound Pith citation observations for arXiv:2507.00699.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00699 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:16:42.794738Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T08:17:10.481202Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact3
  • verified fuzzy7
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 98959035-344d-4dd5-b524-8cc2a7ddf776 · outbound

This paper cites Soen-101: Code generation by emulating software process models using large language model agents,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Soen-101: Code generation by emulating software process models using large language model agents,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.649612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.649612Z digest=sha256:2214e9c77cbb8c1c6b412ecb7df5ac33d775f7b9b6adafd6a4c5282aadc2e9c5

Observation d930fc1b-9079-4376-9294-e8d94752079b · outbound

This paper cites ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback ROCODE: Integrating Backtracking Mechanism and Program Analysis in Large Language Models for Code Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.840152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.840152Z digest=sha256:e1ef791cfb8548831fccf036c69d11a8590a31d39acd4e447091451bf9fbdea9

Observation 815bdd69-f192-49d1-a45f-c0730df2b1fd · outbound

This paper cites Skcoder: A sketch-based approach for automatic code generation,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Skcoder: A sketch-based approach for automatic code generation,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.912600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.912600Z digest=sha256:26564e7c92762776975789a4a4fd6cfedd294b1951e912db60071358b2ce6e04

Observation 05f51ad5-7b73-48d2-af70-60563fdb60ed · outbound

This paper cites Enhancing Code Generation via Bidirectional Comment-Level Mutual Grounding.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Enhancing Code Generation via Bidirectional Comment-Level Mutual Grounding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.926501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.926501Z digest=sha256:367dbec6f999602d63b7d6917b1fd413255bcf88b8115ac8d97b977aa90f03f9

Observation e0d3c8c6-b55a-49ef-ad5e-effa4a22987b · outbound

This paper cites Fixing Large Language Models' Specification Misunderstanding for Better Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Fixing Large Language Models' Specification Misunderstanding for Better Code Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.068895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.068895Z digest=sha256:2b77ff64611abd2dc1c1a473fd9e438b9c2a1cb0929865e0ea026ad89e40e90d

Observation b61f1b99-39fe-4665-b59b-aa73c79502a7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.283330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.283330Z digest=sha256:fab45eba33df9c606228cdb8ca7a6acb8df7fef18713914c13d54ab31d7c3432

Observation b47c095f-2342-44b3-9824-95727e3a39bc · outbound

This paper cites Program Synthesis with Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Program Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.360449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.360449Z digest=sha256:4173449fa0ef7e4926f57e3528d3072cdbf15f387539d92521cbc3f0efb7003f

Observation 3bc9bdaa-f0df-4836-91cb-e64f7b267922 · outbound

This paper cites A Survey on Evaluating Large Language Models in Code Generation Tasks.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback A Survey on Evaluating Large Language Models in Code Generation Tasks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.365515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.365515Z digest=sha256:6380c7a6094f66990812ff98b9ce4f866cae76265bcf24fb81001099d367ca72

Observation 43c39c9e-b5aa-4659-bec4-0e0753fedd37 · outbound

This paper cites Instruction-following evaluation for large language models,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Instruction-following evaluation for large language models,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.453046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.453046Z digest=sha256:46b5740ece19638cf4c3062a0f6ea257226529bb044c3d051c37bb5677c822c8

Observation da751ba6-c1d1-47df-985f-5a1f713dda1c · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.489193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.489193Z digest=sha256:6aaf82b5d92291d22a06fd8aa9b584988ba564cde724444fadc7c7f7365ac0e6

Observation f1774027-b0c4-4f17-81a5-df9a12cc59c7 · outbound

This paper cites CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.524744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.524744Z digest=sha256:5c658db5a5a4e2ca1d84100c621f606fd1e1a911992384724f927bfedcb7f77d

Observation 1193c980-a560-4bd8-85e2-898adf58dbbe · outbound

This paper cites Codeif-bench: Evaluating instruction-following capabilities of large language models in interactive code generation,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Codeif-bench: Evaluating instruction-following capabilities of large language models in interactive code generation,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.559939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.559939Z digest=sha256:55850a7b20247ed6e326dd57827c924bbfc0b80aa503215b59f1fef7bd329587

Observation 79701b5d-baf1-4604-9754-541bf883a75a · outbound

This paper cites A hierarchical and evolvable benchmark for fine-grained code instruction following with multi-turn feedback,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback A hierarchical and evolvable benchmark for fine-grained code instruction following with multi-turn feedback,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:48.065482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:16:41.648086Z digest=sha256:a13d69ba1dd8273af64246e46f40e2610642084820f1499bb0db1ecce368ce8d

Observation ea3e0791-063d-489c-a370-71143a52ca3b · outbound

This paper cites Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Beyond Functional Correctness: Investigating Coding Style Inconsistencies in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.671896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.671896Z digest=sha256:ca44245f7107bbcc630ee09a798e08912af3f1bbf8316381f00cd8cee08b4b0f

Observation cea143ab-1230-4175-9893-9bb48f9b74f8 · outbound

This paper cites FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.694748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.694748Z digest=sha256:b713d09ec1235ea8b21412c85979845ebeff72447f7d34a3337c0a409c053b5a

Observation bb171d85-fe5a-43b1-837c-f0361ca7f398 · outbound

This paper cites Sharegpt,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Sharegpt,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:47.895647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:16:41.734747Z digest=sha256:9d105e47ee09829ba44226c55c009f1fc1052a7a262166a0452bdd0d9420c87e

Observation f6eabef0-6a5c-4206-bbcc-a25731a628f8 · outbound

This paper cites Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Prompt-Based Cost-Effective Evaluation and Operation of ChatGPT as a Computer Programming Teaching Assistant

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:16:45.463009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:16:41.784746Z digest=sha256:aff5d4afdb2e07141dcb26068e84254c2e287cc921dbf96434f427b58d99757f

Observation 1e647956-e148-494d-8d08-dc3e467a846e · outbound

This paper cites tree-sitter/tree-sitter: v0.25.5,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback tree-sitter/tree-sitter: v0.25.5,

Reference 22

Resolution
verified exact
doi, observed 2026-08-06T21:16:43.134792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:16:41.819239Z digest=sha256:3630770f7c0fe143d2568209bafd4fc6301c7a6981e986bd14660511d6f06b48

Observation a19eca15-2d4a-4f3a-9595-400190df6f7e · outbound

This paper cites GPT-4 Technical Report.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.874749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.874749Z digest=sha256:472fce07c3981d2c4cef6fe28467810a8c2a0cbf59d1d62222c4a4240ff59221

Observation e6c825af-d137-4486-9936-ea7107ce2302 · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback ROUGE: A package for automatic evaluation of summaries,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:47.604770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:16:41.914739Z digest=sha256:e0c454d38cb0e87f7ab2f317aaadddc9d78775e5ee7be1916e5fafd017ec939d

Observation 5a0c9c0b-bf7f-4272-9ae4-e7c2665ef132 · outbound

This paper cites GPT-4o System Card.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.937129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.937129Z digest=sha256:eabcc052c32222acb92f0da40dd8f07168e35683573197c396cc22c5de35958d

Observation 26e3e459-d30c-460a-a7ed-295320b0c912 · outbound

This paper cites Claude 3.7 sonnet and claude code,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Claude 3.7 sonnet and claude code,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:47.186490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:16:41.963659Z digest=sha256:cd8adae5c8b18095a27a1cc5901f0679b922bf2fe7ca789b18dbb58228ffa1a2

Observation 84baed67-510b-43ca-b183-a206c5ce0c6f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.994749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.994749Z digest=sha256:c547305857cfa89a3a2dbb8c590791d79b705528cc6a41dce4556466bf88f38c

Observation 8abd2f57-2fd9-40f2-8480-f908485b89dd · outbound

This paper cites DeepSeek-V3 Technical Report.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DeepSeek-V3 Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.034748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.034748Z digest=sha256:6def2a6b1eb8d299750b168675374d203c51d8dceaa4e728c9afe078c2597398

Observation 7a297632-837c-4105-8ad7-12f0e53f9b6d · outbound

This paper cites Qwen3 Technical Report.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Qwen3 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.094905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.094905Z digest=sha256:fe09359464ff57650cb59dd093d1a4c3f88ab302469146c7a7b9eb6a2d1c1618

Observation 1927e71d-37ba-4f26-8e8d-3d54415b000e · outbound

This paper cites The Llama 3 Herd of Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback The Llama 3 Herd of Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.134745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.134745Z digest=sha256:bd59de78cd6e5c11d302fd68e2adbb82023f3f425bca14bddf87f35bb6f0e0bc

Observation 37067ce1-3533-4b5b-ac4f-c3abc6105d35 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Efficient memory management for large language model serving with pagedattention,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.174750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.174750Z digest=sha256:2a0804e4a69a255527f8a5fd701474ab1b55f495912d30be8fe6f4ab7f189c52

Observation 984bcaa4-5042-4ed1-8926-f2ea96675d1d · outbound

This paper cites Guesslang: A neural network to guess the programming language from code snippet,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Guesslang: A neural network to guess the programming language from code snippet,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:46.819549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:16:42.225023Z digest=sha256:3bbd3f8903301bcc4ec62b5168454f8786043ace8ee87e1e76f156b65dc8678a

Observation 154ca5db-694d-4241-b4e1-f0df76995d2e · outbound

This paper cites DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.274834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.274834Z digest=sha256:cb92330866bd6aee9e96063e4aeb742c08afbf5790138da4b8372f575ef240c6

Observation a938fd60-4e67-407d-864f-5b65fe7c3f08 · outbound

This paper cites ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.308763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.308763Z digest=sha256:b1acbc1eb90526b7da0c8435188fadc204447a3b8fb8d4bc1552b9e46fad8864

Observation 72cdc08a-3578-47b8-8024-0aced3a8923c · outbound

This paper cites Learning-based widget matching for migrating gui test cases,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Learning-based widget matching for migrating gui test cases,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.314114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.314114Z digest=sha256:ab29a9ec21e22dc4084878b4033d7f9bf39fd0501f20e1b01e28da8f9dff8cec

Observation e6008267-b874-4237-a2fc-97efca23c204 · outbound

This paper cites DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.331104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.331104Z digest=sha256:d56966462c2979ef156f00ea31a9c0fa80ba4fb78df55e0be9399dae2093b4d6

Observation 459eb448-1417-48ee-bb16-61b7931bcb71 · outbound

This paper cites RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback RustEvo^2: An Evolving Benchmark for API Evolution in LLM-based Rust Code Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.334629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.334629Z digest=sha256:5bdd3131146703eaa525fdbcfa3558a3f9303a16b8d61b4b211d7882665f8d2f

Observation b0b49b14-e3ac-4ed8-9abc-f6418ac8b6db · outbound

This paper cites Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Feedbackeval: A benchmark for evaluating large language models in feedback-driven code repair tasks,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.374745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.374745Z digest=sha256:cd5d826adfaa7faae3231c5f32f39665611c3f52201b5b5c7112c3d0a76b93d5

Observation a95a8c12-ebd9-42f2-bcf9-72e7ba49ac5e · outbound

This paper cites Benchmarking complex instruction-following with multiple constraints composition,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Benchmarking complex instruction-following with multiple constraints composition,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:46.644479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:16:42.424811Z digest=sha256:8ac4ac98cc445a7417978fff1d963465ad31de6c895e3ed46f1712e085d3253f

Observation 472da2bc-29fc-4846-9f68-71cf6927d5ff · outbound

This paper cites Generating Equivalent Representations of Code By A Self-Reflection Approach.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Generating Equivalent Representations of Code By A Self-Reflection Approach

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:16:44.004826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:16:42.503377Z digest=sha256:c2a0b825eccff55038e683d481b993c1846e065e1abe066bb77eecdc4257d0ce

Observation 674a015e-81df-4cc0-8811-24c3c0e410d4 · outbound

This paper cites Codescore: Evaluating code generation by learning code execution,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Codescore: Evaluating code generation by learning code execution,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.546851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.546851Z digest=sha256:3056e84ee28f0e1e7f88633bf9022e76d592d8c57cdf099d9c03e5959d13e04f

Observation d0639506-78c1-4147-a23d-9a7d2faeb7e1 · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.464751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.464751Z digest=sha256:e7561c64a33701bd5698291a19845ce3cd4b7928cac0296a4109bf51fd7efdce

Observation 40aa013e-3916-4b50-84d9-62400e7e56f0 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.626496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.626496Z digest=sha256:14e658238cb83b84558656cd27db243afabfe11d44647f33e1bf695e37a0ef1f

Observation ca44739c-0b3f-4e91-aadd-61e8e4bc0971 · outbound

This paper cites Magicoder: Empowering Code Generation with OSS-Instruct.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Magicoder: Empowering Code Generation with OSS-Instruct

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.654739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.654739Z digest=sha256:f5072279f681757080609323d9b69596e8839f7b24f796c17112dd3a8aed282a

Observation 7d317827-9bbe-47f6-a99c-de1dfb3fb9b5 · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.591153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.591153Z digest=sha256:ad0ca3371b1418e6945b092924d9858c2a4a08f3100c821e6dcced5dfd1a073c

Observation 57834f6e-a9ec-4a0c-8657-fbc15a330194 · outbound

This paper cites Genetic instruct: Scaling up synthetic generation of coding instructions for large language models,.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Genetic instruct: Scaling up synthetic generation of coding instructions for large language models,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:46.608394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:16:42.750003Z digest=sha256:4c5de3dc8c51903d80f06ad05bb9aede2fe939269d1eb86e890a694b0fec220d

Observation 9b693b4d-0478-4898-8651-a21d4f48dba9 · outbound

This paper cites WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback WaveCoder: Widespread And Versatile Enhancement For Code Large Language Models By Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.694742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.694742Z digest=sha256:f1032034525d33e64aeac5737f454a4ce5a97e19011f00ad55360de0aef9d539

Observation 54853ea3-73e6-4c99-80e1-4982df7be0bd · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Instruction-Following Evaluation for Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:41.484523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:41.484523Z digest=sha256:c0ee42df10a486d6d02d58ad9ab57a51fdd1f07f82bd383f7c3a579de3453a5e

Observation 918c7b8c-6c40-4878-8711-7b74779eb127 · outbound

This paper cites SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback SOEN-101: Code Generation by Emulating Software Process Models Using Large Language Model Agents

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:40.724767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:40.724767Z digest=sha256:61d94f5f99f8df897e5d26a2a5cc3092fe4cd956dc8bc000dfddc49fd048f740

Observation eabb4b14-24c4-49af-a418-bedcec7e29c1 · outbound

This paper cites Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models.

A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:42.794738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:42.794738Z digest=sha256:7bde5599221616db2b74d4bf160ccd1008aecb83a83b3653088cdb6cf25d7c8a

Pith citing papers

Observation 39c86638-d467-4dad-9eba-bfbb7c127715 · inbound

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code cites this paper.

Bridging Generation and Training: A Systematic Review of Quality Issues in LLMs for Code A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:21:11.028874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T17:37:51.790000Z digest=sha256:23d8e3418b806298d3c596888d212cdac4c3ef647bfb6c0db6e0b04447f5abd6

Observation 290c59ff-805d-4373-ba60-59b11e8a9b3e · inbound

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models cites this paper.

When Should Models Change Their Minds? Contextual Belief Management in Large Language Models A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.976976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-29T07:17:07.720589Z digest=sha256:31c641fe284264d90fdcd03e1975f4874791e1c3323fe12eac9dbf8799a0e36f

Observation 57e95625-d60e-4a7b-a963-890c4f78eb3d · inbound

Prompt Governance? On Governing Technologies Governed by Natural Language cites this paper.

Prompt Governance? On Governing Technologies Governed by Natural Language A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:25:32.385646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T08:17:10.481202Z digest=sha256:ae1b9c1d7a538495995be0870559a20d038131f1c1b658ba6caac075a36627d5

Observation c805b868-143f-4aaf-89c6-f46868efae58 · inbound

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues cites this paper.

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:20:06.984330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T20:22:29.287534Z digest=sha256:f9d3392b8046e40e1959beedb2e15a83c6070c071afebba0025b02256dc8a2f2

Observation 9f50b92b-d06f-40a4-a0b3-2d24dcfef8ae · inbound

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues cites this paper.

CodeChat-Eval: Evaluating Large Language Models in Multi-Turn Code Refinement Dialogues A Hierarchical and Evolvable Benchmark for Fine-Grained Code Instruction Following with Multi-Turn Feedback

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.347673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T07:05:50.811872Z digest=sha256:f95f87446c28f5d3498d9f6e5737078870486937d115cc5863f721d176adbba5