Pith. sign in

Paper Citation Record · LEDGER

MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2501.17399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.17399 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:12.420401Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:48:56.198096Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 140c103e-9697-47e7-b78b-0a93bb3b179d · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:11:09.335166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:922bcc521821505529692a63faad9ed2431ca44f0986644d5b0d52c758f3b70e

Observation 6e46e7b3-0481-483d-8932-7b2ab6cebf05 · inbound

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models cites this paper.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.420401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.420401Z digest=sha256:11546a24f4117d39335bbcacd8ade4a2a18d8b94bd50750b29aa8695767f7c9e

Observation 8158aad3-a061-4172-a5e1-392d3e4fa3a4 · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.110876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.110876Z digest=sha256:cc457ef359346577ccccbbfd5b8c09633858c4cb5cf907b747de980a0eb97fc5

Observation 583020f1-6aaa-4059-a310-44715a23cdb8 · inbound

ImgEdit: A Unified Image Editing Dataset and Benchmark cites this paper.

ImgEdit: A Unified Image Editing Dataset and Benchmark MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:17:45.350173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T18:17:45.123690Z digest=sha256:2019f0ac30d5fa23db020df72b2600ef4c7936f960c5b64a4da987c690304303

Observation 0e6f16b8-a732-4ca0-90e5-adf7914b8582 · inbound

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention cites this paper.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.518006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:64a0a125a095414343f2b602cbc4c0a477483cdf428b1344296defd525ef7d25

Observation c76d5d24-86a7-48e5-a1ca-4b014cc39af5 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:51.856959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:51.856959Z digest=sha256:45e4c4a6bc006b417f6fed83d55b50debc0710c1f42fc2f68b4ca2a6b8ad4079

Observation b2aa6ed7-d66e-413e-b8f1-d92c97ad80f3 · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.845701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:b2e083e2158acaa0499274c9419184ef887a0600d856239dce1caf9703454f97

Observation 0a513bc0-9837-42b3-8f72-4faf142f9ccd · inbound

TextQuests: How Good are LLMs at Text-Based Video Games? cites this paper.

TextQuests: How Good are LLMs at Text-Based Video Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.669131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.669131Z digest=sha256:679b512d366e2221f38fd42e672d462a4c2e91663aa9264a8ceff5c1ab1456c5

Observation 57df649d-5060-47bc-9a32-08469a0b455a · inbound

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries cites this paper.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:48.990246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:48.990246Z digest=sha256:52d547eff0ed4c117e558a2538dea4446b15368cec1dc4af9b358cc0db3fcfee

Observation 4cfd9f18-3805-4e8e-867d-263f705e7254 · inbound

Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting cites this paper.

Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:12:10.089168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:12:10.089168Z digest=sha256:8a194330ca93979720b040908e914b6fe25b699602bf30ae0c61f23cd34a9b0a

Observation 5c091189-e56c-49de-bc03-2054b4e04159 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:39.919893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:39.919893Z digest=sha256:69ccf3a74de6da8785b3c6d57401c5289cc3bd89441d8d41327f2778488b0fe6

Observation dc68545d-532a-47d9-900e-a2d6861c90f4 · inbound

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks cites this paper.

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T09:40:46.562258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:40:46.562258Z digest=sha256:3634a0c75de8fb538b15fe11b96ca8dbef1b33969616e6c7c1df4186f357b2ef

Observation 99759f4e-bdf1-4db9-bfec-fcf4a21a6ead · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:16.105119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:bbb32b433b904a896f3d1cbde07f6f3905a5b78e8fc6ed2805cb68178c41ec7c

Observation 6d2bcf41-ade4-4328-9738-7dbb80d785dc · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:50:50.469751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:7c6b4a62217025c31ce77cc9774a454e364ff478e33a63c8ec1e584da510400e

Observation f6d4d5ec-bc87-4f69-a129-da250de9fe45 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.445640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.445640Z digest=sha256:0fc783ecba00facddd2e74cd7d2243a4c32f0653d9baa5bcd863186b4681927c

Observation f49b607a-92f2-41a3-8729-aab61af32184 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.360793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.360793Z digest=sha256:3bcdea87783c042d96e2e65538c95cd2fcd7cc800f76642b3d29f3066aa0cf52

Observation c5fb9061-7d35-49af-9232-0e9cad0e51a1 · inbound

RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation cites this paper.

RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 32

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T09:17:40.301640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:13:55.204404Z digest=sha256:8ffb8f24dfce59129e49adb72d8af04df4c56cfe91b38e0df394d27daac7f58a

Observation 2983f954-59cb-4294-81c6-ac074e7e8e25 · inbound

Evolving and Detecting Multi-Turn Deception using Geometric Signatures cites this paper.

Evolving and Detecting Multi-Turn Deception using Geometric Signatures MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T15:23:32.697954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T15:17:58.804973Z digest=sha256:a9724952d2d90cc50bf712aa29c4fa5c71600cdc26b326af46d2e21601f2d107

Observation 9b3ea34e-c4b4-41ff-bf02-b439edf5abf8 · inbound

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses cites this paper.

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:56:44.740957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T07:16:30.725358Z digest=sha256:d8a956ee8b86b260c90be595777cbdf8c86d897f9be99c0e5972b3a03dfdbdd5

Observation d3050268-e08e-479a-b75d-258cccea1f49 · inbound

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios cites this paper.

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:48:46.351978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T03:39:36.657903Z digest=sha256:aa28ab63bcdc8390c72a30d39bf35a10f8d856cb0198a7b6acdaf7545464824d

Observation 36c02bfb-a6e5-40e5-8841-42a8188303c7 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.199387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:732b3bd6c807119f3db40d88a829b01bb1abeb6f40f3b291330ffe3f81f56e55