Pith. sign in

Paper Citation Record · LEDGER

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents

As of 23 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2608.11727.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11727 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:35:53.062182Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved20
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dd50a0ba-0227-47cf-b045-7c1c0072b35d · outbound

This paper cites Claude code, 2024.https://claude.com/claude-code.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Claude code, 2024.https://claude.com/claude-code

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.920647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.851448Z digest=sha256:0beee0d0e15447153002266e4a255b1b0c50ad80c50dc89471265b536a558e40

Observation 7557ce4c-1c0b-48be-a577-805828f92a99 · outbound

This paper cites Models overview.https://platform.claude.com/docs/claude/docs/models-overview, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Models overview.https://platform.claude.com/docs/claude/docs/models-overview, 2026

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.906832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.856428Z digest=sha256:180ebab9750e7130b3ed7482f0c2071df9af5627abf1040d686cf371d7539397

Observation 560ffd42-11fb-40ea-b414-76b7949bbc90 · outbound

This paper cites Building effective agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Building effective agents

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.892417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.860808Z digest=sha256:535e6565ce6ba7b5aa06c6755b040a5aec2e0e8a2812c93c5beac686fd452172

Observation f2924955-bd89-4804-b29b-550ed3a30566 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.865427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.865427Z digest=sha256:6aad0c5ef12239c3b0a90dfcd6e1e4d677a3a6380ecb452b3d711ef03b60eb4b

Observation 73fe18b2-1629-4890-a28d-c9f2fc1df10e · outbound

This paper cites MLE-bench: Evaluating machine learning agents on machine learning engineering.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents MLE-bench: Evaluating machine learning agents on machine learning engineering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.878606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.870160Z digest=sha256:0c14d94cf43e339224d595951fd4a0ec0ef95d5ec174c230f6b908015168270f

Observation e4edd941-19be-441f-87f2-4f4d0ff8f748 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.874574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.874574Z digest=sha256:cd6e8c03ce49c4efcb4fb2d7a6f526c7e45df20d2c51834b0c96f6e48d93f11e

Observation e6047447-9ca6-4feb-8f1d-075c78ce9727 · outbound

This paper cites Datasheets for datasets.Communications of the ACM, 64(12):86–92, 2021.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Datasheets for datasets.Communications of the ACM, 64(12):86–92, 2021

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.879593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.879593Z digest=sha256:2c9fde50da5b973c7158171f058ddf90d789b3766de45737861ff1736f2d1a06

Observation 3b2aadf6-456c-45d5-9404-a5c67300fee5 · outbound

This paper cites Gemini 3.1 Pro: Model card.https://deepmind.google/models/model-cards/ gemini-3-1-pro/, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Gemini 3.1 Pro: Model card.https://deepmind.google/models/model-cards/ gemini-3-1-pro/, 2026

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.856111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.883711Z digest=sha256:d0caffea340c20a011cc5439e4b5ce92d53e43a333831ebe8c9542fb2ccbe330

Observation 94310d55-504c-4172-8c4c-91f3f2f8d770 · outbound

This paper cites Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Model Context Protocol (MCP) Tool Descriptions Are Smelly! Towards Improving AI Agent Efficiency with Augmented MCP Tool Descriptions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.887791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.887791Z digest=sha256:1626e270c98b3053d0bd06e7e6863107f4dbfba5c45d7ab56d992f3de6f11169

Observation d262a897-5a56-4027-8122-d9875218a7ec · outbound

This paper cites ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents ProcCtrlBench: Evaluating Process-Level Defects and Control Preservation in LLM Coding Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.892302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.892302Z digest=sha256:80951fd56813ee966f96f0f55f875a68787a75f0bafe3d812c8654de21e07f10

Observation 1ec3ce90-7e45-4177-ba2f-09dbbbe802f9 · outbound

This paper cites Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.896736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.896736Z digest=sha256:abdb0791bb77c4188c85967e00efabf78c71bf81792449cfef882a0ff0a5b3aa

Observation 32704800-6895-4df4-93cc-8bb392eb205f · outbound

This paper cites MLAgentBench: Evaluating language agents on machine learning experimentation.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents MLAgentBench: Evaluating language agents on machine learning experimentation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.843216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.901665Z digest=sha256:75cf0f52ca214c277fe6a96033719bc33d020a023b937c0ffc96e49bf72ee22f

Observation 8f2da929-ae61-4916-8c45-5c4111efbbbe · outbound

This paper cites FollowBench: A multi-level fine-grained constraints following benchmark for large language models.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents FollowBench: A multi-level fine-grained constraints following benchmark for large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.829881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.905780Z digest=sha256:0476fb726add9525265751013e740095926109a3c1d3cf9c498afb3a5dc1e6d1

Observation 6b06da15-8e42-4bf1-8756-504e09057943 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.816783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.909830Z digest=sha256:61a598d5e3b076ea14626b1cb174e3b80ca26935ba2845252bf1fd519ccc562d

Observation 28198d65-8b6a-4848-b2d1-a9e0f1108062 · outbound

This paper cites AgentBench: Evaluating LLMs as agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents AgentBench: Evaluating LLMs as agents

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.803265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.913841Z digest=sha256:407c64d995cd1580c39fa77e30095a14a73ffb8c34c9dbe804a5f3c6123a054b

Observation df2458e0-0b6a-4ac4-9ef9-0fa7cf30241d · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.917646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.917646Z digest=sha256:0ce800292c8c8fb54eb4ab852c3e7d0a9383f70c1f94f5dc9080bba8ea0ce59f

Observation efb42d1a-3422-4ce0-9941-37f27846325d · outbound

This paper cites GAIA: A benchmark for general AI assistants.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents GAIA: A benchmark for general AI assistants

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.789693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.922037Z digest=sha256:7309df1d33498f1a72e21c942e92fc588927dcc612ffe4a1949602339f70dd5c

Observation d3118d91-10d5-4cf4-b83f-706f44eddd52 · outbound

This paper cites MiniMax M2.7: Model self-improvement.https://www.minimax.io/models/text/m27, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents MiniMax M2.7: Model self-improvement.https://www.minimax.io/models/text/m27, 2026

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.775883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.925970Z digest=sha256:de9df2bac3560a59f1ad4617533eac04eacf486ca96675967c25703249dc7516

Observation e054397b-6d58-42a5-8ac3-d116d6fe3abc · outbound

This paper cites Kimi K2.6 model card.https://huggingface.co/moonshotai/Kimi-K2.6, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Kimi K2.6 model card.https://huggingface.co/moonshotai/Kimi-K2.6, 2026

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.762751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.930082Z digest=sha256:8311f0836edbc30030efb204b94f6346ae44db5e69fb64f75e1df2ae13ded8d7

Observation 5ae1cba8-2279-44c7-b66a-0cd602440146 · outbound

This paper cites GPT-5.5 model.https://developers.openai.com/api/docs/models/gpt-5.5/, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents GPT-5.5 model.https://developers.openai.com/api/docs/models/gpt-5.5/, 2026

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.749856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.934313Z digest=sha256:af480a4bf8b9072349128c336051ef731991530fa1c968e095963da7939622ae

Observation 7a44892d-486a-4291-9b8c-752d88af48bd · outbound

This paper cites Introducing SWE-bench verified.https://openai.com/index/ introducing-swe-bench-verified/, 2024.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Introducing SWE-bench verified.https://openai.com/index/ introducing-swe-bench-verified/, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.736649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.938862Z digest=sha256:8b93a2db58fa740507c3f30443d8666654a2de4dc4277d24ed0acd673af074a1

Observation ff6098a0-7b02-4743-a8fd-44d32f79c60b · outbound

This paper cites Patil, Tianjun Zhang, Xin Wang, and Joseph E.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Patil, Tianjun Zhang, Xin Wang, and Joseph E

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.722887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.942970Z digest=sha256:d78719e959c98c78b0122bacf957418fc71100d77ccf4ea6729a5ce5d11d016b

Observation 767cbce2-2a76-4483-a3df-ecf87742d934 · outbound

This paper cites Patil, Huanzhi Mao, Fanjia Yan, Charlie Ji, Vivek Suresh, Ion Stoica, and Joseph E.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Patil, Huanzhi Mao, Fanjia Yan, Charlie Ji, Vivek Suresh, Ion Stoica, and Joseph E

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.709443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.946957Z digest=sha256:b5ad621a6ee2d93437595afc3311abdc2f27434f4ef4d1dbbacdfd5ac4c11532

Observation e7ed8de2-7101-4010-ad4c-8a6e5f775006 · outbound

This paper cites Generalizing verifiable instruction following.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Generalizing verifiable instruction following

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.696155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.951036Z digest=sha256:169a19c28d14eb59b2dfae90c2c204bfc92321264e81b34aabaa67b6e0a642c9

Observation 45da52c9-23f5-44ad-961a-597793a83213 · outbound

This paper cites AgentIF: Bench- marking instruction following of large language models in agentic scenarios.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents AgentIF: Bench- marking instruction following of large language models in agentic scenarios

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.682836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.955062Z digest=sha256:1251c6ed7782f083d0ad2eafc1fb77e9ff94b44aa3a23bad53b2dda0775ff69c

Observation 87e0b0e3-8571-4ad4-9c60-ec7efb7e25f3 · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.959112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.959112Z digest=sha256:9e99ba6e5602d56edfe2aab49b157f83d3bf4d0fa9d2bc054faf898fdcb10990

Observation d094cc47-8480-4ff5-8985-7aeb92da1199 · outbound

This paper cites Qwen3.6-Max-Preview released.https://qwen.ai/blog?id=qwen3.6-max-preview, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Qwen3.6-Max-Preview released.https://qwen.ai/blog?id=qwen3.6-max-preview, 2026

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.669629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.963336Z digest=sha256:aaac31ef0eae26d02f3062aee92b782ca30edd8c57b614d47f5078013690e081

Observation dbbf2186-5bb2-4760-95e3-1dfd29bb93e2 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Toolformer: Language models can teach themselves to use tools

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.655948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.967780Z digest=sha256:add48d96cccb4836ed8e5e34f0111a19d2746da4efd6b33842c66280d7d0ba0b

Observation 35d5b2e3-80c5-44ac-9dc9-57b4ad1685d0 · outbound

This paper cites Seed2.0 model card.https://yfz.ai/Seed2.0_Model_Card.pdf, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Seed2.0 model card.https://yfz.ai/Seed2.0_Model_Card.pdf, 2026

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.642324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.971632Z digest=sha256:0d148020ef3b10a9e561d73f14feb0074e426def73af6d16685f95db63b34c2e

Observation 90050bbd-78d8-4c3a-828c-624629e18d88 · outbound

This paper cites OpenAI GPT-5 System Card.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents OpenAI GPT-5 System Card

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.975569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.975569Z digest=sha256:2c9e95513099f9e87be3676592227b59ec8235b6964fa7c65f3deb1c50230863

Observation 82bd7606-6097-4408-aa42-c0d2a839f38e · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.979863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.979863Z digest=sha256:e3bb5641435c5d13019ad857abb028d52838305915094510b3b4c445c08889f1

Observation 4a86962c-805c-4b22-bedf-f6a827ea6503 · outbound

This paper cites Step 3.5 Flash: Open frontier-level intelligence with 11b active parameters.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Step 3.5 Flash: Open frontier-level intelligence with 11b active parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.984189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.984189Z digest=sha256:1e8a1d8d8777c3ebb64b255369f2a13d32c3d48c8b6d942684254a4e59bf0142

Observation 64808c17-4970-4704-b91e-663890d60418 · outbound

This paper cites Tencent unveils Hy3 preview.https://www.tencent.com/en-us/articles/2202320.html, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Tencent unveils Hy3 preview.https://www.tencent.com/en-us/articles/2202320.html, 2026

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.988276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.988276Z digest=sha256:9ad51d1d89f915059cf95543d4a6408038502c005f10fae7efe5f204cc238fae

Observation 77423fd4-60a9-4daf-8854-e3c86d5d6246 · outbound

This paper cites AppWorld: A controllable world of apps and people for benchmarking interactive coding agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents AppWorld: A controllable world of apps and people for benchmarking interactive coding agents

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.629240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:52.992293Z digest=sha256:aee7eef7839ab65bdf8de7b315e8bce691749ea730b15032887e76d98ec87454

Observation 2e12bf06-e39b-43d8-8e1e-5e0da8119158 · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.996449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.996449Z digest=sha256:611c760aded172bc57678171879f979478701e2e1f44bab3f885980075e57c56

Observation ee682b72-a39b-4e0c-9b19-2a99d4a285e7 · outbound

This paper cites CodeIF-Bench: Evaluat- ing instruction-following capabilities of large language models in interactive code generation.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents CodeIF-Bench: Evaluat- ing instruction-following capabilities of large language models in interactive code generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.001057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.001057Z digest=sha256:4ff5b9edabbd80e741c7da558ae440ad7590ada88ba5ec1866cf398ffc0a60f6

Observation 45288d20-abf6-4f03-a50f-de7b520a4308 · outbound

This paper cites Benchmarking complex instruction- following with multiple constraints composition.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Benchmarking complex instruction- following with multiple constraints composition

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.615614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:53.005172Z digest=sha256:462966f6e83ac2258d6a99cf86fe9183df54b8fae787ce05ce6f1aa877719458

Observation aae303cf-a1b7-4cc0-bbe1-371988a61791 · outbound

This paper cites LIFBench: Evaluating the instruction following performance and stability of large language models in long- context scenarios.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents LIFBench: Evaluating the instruction following performance and stability of large language models in long- context scenarios

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.602044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:53.009595Z digest=sha256:81bb3a59bbeda824dd1a2be23a2f2af87a2821ad6eff9892759e5558b3e4deb2

Observation 037af9e7-2a24-4442-a52b-9270db7beb63 · outbound

This paper cites OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.013801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.013801Z digest=sha256:a9e46c23a93fccb4733bf158b70fd6f23661f620ef3d2343b4a7312bbed275be

Observation 87759daf-12d3-4382-96a8-f99e0c5820a9 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.018025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.018025Z digest=sha256:7ffc7aad468b872738d0dc1d0bfb3c6f0594fd8d22f7438998607209d580b591

Observation 6592ac29-096c-42b7-8f8a-2e77a9c3a174 · outbound

This paper cites Jimenez, Alex L.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Jimenez, Alex L

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.587390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:53.022196Z digest=sha256:6aed3dd19dc4e0cedddfc75c297f0451eda84aaed54861caa74df92223a9581f

Observation 7b207577-7c42-4392-ac3c-2ae67eeabfd9 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.026282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.026282Z digest=sha256:60107b31edf2e9e44cd1e8950d8d1ca20cbc4c9c4366680598a5a21596d52a46

Observation fca0b68c-e694-4d6d-a48d-84bda6132c8f · outbound

This paper cites OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents OpenSkillEval: Automatically Auditing the Open Skill Ecosystem for LLM Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.030478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.030478Z digest=sha256:d083328bfdab5a671aa08d73d6d2bea5b40136cef4e096da57a8610b41090385

Observation ccd56daf-e901-4d80-a080-3f9f47707b8f · outbound

This paper cites GLM-5.1 release notes.https://docs.z.ai/release-notes/new-released, 2026.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents GLM-5.1 release notes.https://docs.z.ai/release-notes/new-released, 2026

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.572368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:53.035260Z digest=sha256:450c54364920ff97cc39d9338f08900436ce443a2307f7241a034513f20aa56b

Observation fa3a73aa-147a-4fc1-b001-5baadf511a37 · outbound

This paper cites CFBench: A comprehensive constraints- following benchmark for LLMs.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents CFBench: A comprehensive constraints- following benchmark for LLMs

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.558499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:53.039245Z digest=sha256:a5258241f3c0a701603a4eeb9ef0071f340ac797a6217ac31d9b13103817a6ba

Observation 44cd9b76-4a7b-4d10-8b90-bf3e3b219e0c · outbound

This paper cites IHEval: Evaluating language models on following the instruction hierarchy.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents IHEval: Evaluating language models on following the instruction hierarchy

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.545016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:53.044353Z digest=sha256:322a917a5f49d191a232bc06bdeb0b2ba1314cd4058c9abab61735bfff0e457d

Observation 422a8ed1-6684-4f66-b8f2-120899eaada8 · outbound

This paper cites SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.049011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.049011Z digest=sha256:3b852c8bcb34132981976c00bd1f61a95f9ca78bc8ffa4476ea778b0abc4e53b

Observation 53a884e0-5fd7-4236-8f10-b025d740cf6a · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Instruction-Following Evaluation for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:53.053616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:53.053616Z digest=sha256:cb352618eec7edd5699f66af0e42f4e5e71c26a743cbac73063043cf5555de08

Observation 34ebd60a-8d73-4711-ba6a-41722ec77425 · outbound

This paper cites Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried, Uri Alon, and Graham Neubig

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:35:53.531867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:53.057988Z digest=sha256:dd08f63a5254a6313b2e0f23290119a5b78ecfd3ead09b4e0ac3eb580f2b2c47

Observation 0c81da32-8ed6-4150-904c-04ab39407e5a · outbound

This paper cites Keep generated summaries compact,.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Keep generated summaries compact,

Reference 50

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T00:35:53.518154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T00:35:53.062182Z digest=sha256:f11337f3b6978ae67a81aba9cda02f26c90108149b4563b4fba36dca2153b907

Pith citing papers

No inbound Pith citation observations are available.