Pith. sign in

Paper Citation Record · LEDGER

Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2410.15553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.15553 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:35:52.896736Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5ce290d2-75f5-434f-88c4-d596368f464b · inbound

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness cites this paper.

Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:54:51.947800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:54:51.947800Z digest=sha256:b655eb1cbdc7330d3cddbace7339b44cfcc93c3cb3143b150429fa0dcd1cb533

Observation bc105988-11dc-494e-93b3-553d4eda689c · inbound

A Survey on Multi-Turn Interaction Capabilities of Large Language Models cites this paper.

A Survey on Multi-Turn Interaction Capabilities of Large Language Models Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T19:32:43.952706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T19:32:43.952706Z digest=sha256:18d85bfdbae74cbef37c3b9d6b261d3a9aa0123a346b439cf8c891c114cc7fc3

Observation 47c23202-94f2-4fd8-a96f-18cc631a6762 · inbound

Process Reinforcement through Implicit Rewards cites this paper.

Process Reinforcement through Implicit Rewards Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 125

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:23:31.002560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T20:23:30.763794Z digest=sha256:4966ce312a5825ae68dbd98e8257b9e8340bf83aec8b5f0f3601606d31023703

Observation 7fd68216-9254-4892-baa3-93af3a2ea5a3 · inbound

Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages cites this paper.

Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T21:10:57.092941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:10:57.092941Z digest=sha256:4996b6e55f99655939b704056992c883affdc617983d57503885bf711b0076b4

Observation 23eff24a-7695-4818-97ec-809f5d827dfb · inbound

Qwen3 Technical Report cites this paper.

Qwen3 Technical Report Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:28.517996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T06:35:27.813995Z digest=sha256:3dc0432ac90dec0edebe001ad36152f03b9ca27156520a3687c6e20dc4c741dd

Observation 18e46598-e187-45a9-a98c-2130544ba8ec · inbound

BELLE: A Bi-Level Multi-Agent Reasoning Framework for Multi-Hop Question Answering cites this paper.

BELLE: A Bi-Level Multi-Agent Reasoning Framework for Multi-Hop Question Answering Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:21.440670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:21.440670Z digest=sha256:cc31c9ddb21bd0c0541f25844ca215e97811679bbeef9c3d505c8241ecff2309

Observation 8bb70e02-d4e0-48a9-bcff-a8cc018a5b53 · inbound

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models cites this paper.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.154089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.154089Z digest=sha256:7c28f1c6f1ca3b274e6bc6fdef9a832ccab7bd8159771d97cb12eb46bf6c8325

Observation f3ab0426-6758-4aac-bd61-241b9d68ca63 · inbound

LIFEBench: Evaluating Length Instruction Following in Large Language Models cites this paper.

LIFEBench: Evaluating Length Instruction Following in Large Language Models Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:08:04.207512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:08:04.207512Z digest=sha256:c18df77ded62bd78702cf49036093b451291e9fdb31d0ac60ec0c822aca0e057

Observation 05b5fbcd-31bc-4eb1-aaba-4da4d714413a · inbound

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios cites this paper.

AGENTIF: Benchmarking Instruction Following of Large Language Models in Agentic Scenarios Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:06.783841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:06.783841Z digest=sha256:adeb453a9980efd006ad8f252d1b082975eb3759fb237650fbc8de29b2327bdf

Observation c9dcc368-e1f4-4179-abc3-6f483870ce25 · inbound

ImgEdit: A Unified Image Editing Dataset and Benchmark cites this paper.

ImgEdit: A Unified Image Editing Dataset and Benchmark Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:17:45.234132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T18:17:45.123690Z digest=sha256:0fad4839386c81cb0941a4c75237b58c2ab0a4a8deac244d981e4dc5dc87b7f5

Observation 6ae36c02-44fc-4021-a281-2a6033109155 · inbound

Evaluating the Sensitivity of LLMs to Prior Context cites this paper.

Evaluating the Sensitivity of LLMs to Prior Context Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:46.479494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:46.479494Z digest=sha256:5502881b9001ff61d21d13c11258232cf67a784452448b6b0f84afdd3992c171

Observation 543ffd96-3bd1-4f99-bef0-58df714337f9 · inbound

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following cites this paper.

VerIF: Verification Engineering for Reinforcement Learning in Instruction Following Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:18.196734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:18.196734Z digest=sha256:2d6cbb3d46167e31831425b2c17fc182968e9c115ee4a130485a22e8966a684c

Observation 5f0dbac0-5e3c-4117-afbd-ad55364688dd · inbound

How Many Instructions Can LLMs Follow at Once? cites this paper.

How Many Instructions Can LLMs Follow at Once? Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:13:46.778264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:13:46.778264Z digest=sha256:b21f63d4e652390b6d6cfcb7531be47747ecb3b54b97633c357598913627fdb7

Observation d3451686-b692-40a5-910a-f735c3fde375 · inbound

Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models cites this paper.

Marco-Bench-MIF: On Multilingual Instruction-Following Capability of Large Language Models Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:04:22.608924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:04:22.608924Z digest=sha256:23f4f8890a3ac0e908f4f3f0157713e4cbdcbf413f51b752ffae6f5d17552af2

Observation 6931e197-79b0-4ee8-955f-8f83aa2a1dbf · inbound

TextQuests: How Good are LLMs at Text-Based Video Games? cites this paper.

TextQuests: How Good are LLMs at Text-Based Video Games? Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.654873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.654873Z digest=sha256:b3b0645bb9b64a952787bf822bf514171cafbfc8b6d9cc9e226e6b032970019a

Observation fe15080a-19c9-44e1-8f79-c440907cd167 · inbound

SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages cites this paper.

SEADialogues: A Multilingual Culturally Grounded Multi-turn Dialogue Dataset on Southeast Asian Languages Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T22:22:20.257178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:22:20.257178Z digest=sha256:e2c2c8ec9a8ee60f5488ffc96165cb4d0fe601bed62d0368b9fba2dc83db80e1

Observation dd7fe4ae-66b5-49da-9188-119d91e55c32 · inbound

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications cites this paper.

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:28.833271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:26:28.833271Z digest=sha256:dc1829c5a50fdc4a35b1ed59ba3b77c4e2ce6bc8873c8fae0e1f16502136f4b1

Observation 34cafd27-05fe-4369-9257-7813e9389d4b · inbound

Qwen3-Omni Technical Report cites this paper.

Qwen3-Omni Technical Report Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:20:37.803106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T00:20:37.406351Z digest=sha256:6a6cbd7e97f9a96a65a959d6bdaf39ca953f5a969096b41234c0c9f42281340b

Observation e9fa95bc-041b-4cf2-9c9c-3dbc0706a646 · inbound

Qwen3-Omni Technical Report cites this paper.

Qwen3-Omni Technical Report Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:20:37.526718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T00:20:37.406351Z digest=sha256:77c2c2b7f87f2c3ef1e1b43c318d93ff30eed8ece7e0a82510c88b6e07b22166

Observation 92da0776-98c2-4f65-8985-74d5e47917c4 · inbound

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models cites this paper.

AAPA: Adversarially Anchored Preference Alignment for Post-Training of Large Language Models Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:58.190607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:49:58.190607Z digest=sha256:c17172759007faac91b340a14c00088f5355f87212e8a2cb104d5f20467ece7c

Observation 5fa6c9ae-3293-4511-94e0-478007e2e4c7 · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 130

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:13:16.097093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:233afd21b746fc9ef43d48a97525d668eba7756e38cc0c6d9631ec691daf82ef

Observation a1bb53e8-6162-40b9-82f2-661310aa9f41 · inbound

Efficient Evaluation of LLM Performance with Statistical Guarantees cites this paper.

Efficient Evaluation of LLM Performance with Statistical Guarantees Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:00:51.716840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T10:58:40.958435Z digest=sha256:aa1c8b187c3c406c1246ab22ba5ac46f278f3ddb7b1aa5fd8024a90ee3cd5efa

Observation c65110b4-8ebf-4fd2-977e-ec4c63dd4b63 · inbound

SAGE: A Service Agent Graph-guided Evaluation Benchmark cites this paper.

SAGE: A Service Agent Graph-guided Evaluation Benchmark Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:21:00.857352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:41:23.956104Z digest=sha256:d82b706c01136648df493c8fb31967a239ee088d24799159391180cf757f8cd6

Observation 85dfff82-621f-44da-864a-afc68d82b734 · inbound

Alignment has a Fantasia Problem cites this paper.

Alignment has a Fantasia Problem Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:31:07.561994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-09T21:37:54.545865Z digest=sha256:077f3b5b6263bc1ec76c041cf6590558d331324d4c39d3ff715e352ab231785d

Observation 75faee17-ffac-4d3d-aa77-d67d838e6eb6 · inbound

SEIF: Self-Evolving Reinforcement Learning for Instruction Following cites this paper.

SEIF: Self-Evolving Reinforcement Learning for Instruction Following Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:55.925346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T01:51:09.514927Z digest=sha256:08406e8d98ea4fd12093f55e0bc905eb5dbaf92118b0f684aaa7be87edc38e7b

Observation 4ebe96a1-4ec0-4f61-b5cb-d078b3596d3b · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:40:53.469699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T03:40:04.692279Z digest=sha256:4810dd7b6c895a27668da45bdadf6c59ad34f1f1445d7bd8df882c20b9b24fe0

Observation 4aa78386-0c4f-48fb-b006-cefa5d0a4cfd · inbound

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs cites this paper.

Benchmarking EngGPT2-16B-A3B against Comparable Italian and International Open-source LLMs Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.895325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T08:14:55.858466Z digest=sha256:27004b447bdfb60496b3d274499de1bce287b0373ffec13a079a2323fc43b0ea

Observation c9639bb0-a7b6-4a1d-86e8-14e563ea8a0a · inbound

When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction cites this paper.

When Attention Closes: How LLMs Lose the Thread in Multi-Turn Interaction Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:22:55.474903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T20:20:27.212052Z digest=sha256:4b98f3eaba29907ce9b58cab3130a25cc0390ef6cdafa29df39b4cba500e354a

Observation 9f0cbc76-5b08-40ae-b50e-c43bdbde0bf2 · inbound

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild cites this paper.

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:04:41.500899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-22T07:03:00.655352Z digest=sha256:36832744ce1b246c074dc4eb6e2d46388126d48476626b156f08e3c01a948cf1

Observation 53af5766-6f74-4442-aa83-e7be4c537bf6 · inbound

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild cites this paper.

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:44:57.876165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T17:38:21.754380Z digest=sha256:6e853b555a0bd0d61c1f037ee9095e97bc7f25df93be3fadf508093715a8fc22

Observation a4b14268-9d3e-4385-aa8a-b0213eac01bf · inbound

IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following cites this paper.

IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:33:24.273265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T12:32:00.055500Z digest=sha256:c392e925cd99f00fd0aea695afd5481bab97b3cdc8d90065cc19be456f1f586d

Observation 9ead1e01-90ec-43a3-a4b2-4732219508ef · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 149

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:17:25.642579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:eefb6d6f0da7908f19259be9b5bb8fc84fcb8c2defa720d25deef9ccdf3f61c4

Observation cab3315e-dfe5-43a9-830d-37610feac956 · inbound

In-Place Tokenizer Expansion for Pre-trained LLMs cites this paper.

In-Place Tokenizer Expansion for Pre-trained LLMs Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T23:50:15.426936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:50:15.426936Z digest=sha256:c9d52dfe0aaac976f8f7bd450d31e03e81f24cbfd5896a148d72c578af8a8f89

Observation e8d3030b-e021-4b7e-a711-47af8e9d0fd1 · inbound

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following cites this paper.

HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:36:07.636136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:36:07.636136Z digest=sha256:de73965d2281aec070989d84a19ba2597478a45d38138aa7e9b03403b1bf4b5d

Observation 3ce27f70-24a4-4c64-89b2-5b147d3a0374 · inbound

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding cites this paper.

Hy-MultiTurn: A Six-Dimensional Benchmark for Deep Multi-Turn Dialogue Understanding Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T11:54:19.499279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:54:19.499279Z digest=sha256:f70fd3574e69c267a4a53b92f0ca16e273c5600eec327706f9f50c0cf148ba80

Observation f84f4902-37d5-4d7b-9cc5-d3708ad5b61b · inbound

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+ cites this paper.

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+ Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:17:10.653129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:17:10.653129Z digest=sha256:1240443d76e1f10675d53358bfb26fd55f199fdf0caa3500b42832bbe747f1bd

Observation a5877146-a186-4571-8180-498013d79e94 · inbound

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses? cites this paper.

Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses? Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:38:43.205719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:38:43.205719Z digest=sha256:2d75d07d5814e1248718f26086dbfffe4f72b17cb4fa93ce308953df33d21235

Observation 33d0aeb1-0da7-482e-8723-1c69e8126fde · inbound

Long SKILL Compliance as Logical Reasoning: Closure-Grounded Detection with Scaling-Guided On-Policy Distillation cites this paper.

Long SKILL Compliance as Logical Reasoning: Closure-Grounded Detection with Scaling-Guided On-Policy Distillation Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:25:00.816862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:25:00.816862Z digest=sha256:0c4e3d062e8656807e72cc7954b3e96c30d21f2993d2e45837451228ac50205b

Observation 7cb0ccac-6e5e-4765-8b12-d6536c089f2e · inbound

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information cites this paper.

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T19:21:59.112477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:21:59.112477Z digest=sha256:c9732356ee7c8071ca50db089d14fd9ed551c9693b138188be578d4327f338c5

Observation 1ec3ce90-7e45-4177-ba2f-09dbbbe802f9 · inbound

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents cites this paper.

Harness-IF: Evaluating Instruction Following Across Instruction Surfaces in Coding Agents Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:35:52.896736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:35:52.896736Z digest=sha256:abdb0791bb77c4188c85967e00efabf78c71bf81792449cfef882a0ff0a5b3aa