Pith. sign in

Paper Citation Record · LEDGER

MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2501.17399.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.17399 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:12.420401Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:48:56.198096Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 140c103e-9697-47e7-b78b-0a93bb3b179d · inbound

LLMs Get Lost In Multi-Turn Conversation cites this paper.

LLMs Get Lost In Multi-Turn Conversation MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:11:09.335166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T00:57:10.262350Z digest=sha256:a3c7d4df5e955c263d4dbf55eb47e54187172e726114d8cad4abd684c7d44ce5

Observation 6e46e7b3-0481-483d-8932-7b2ab6cebf05 · inbound

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models cites this paper.

Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:12.420401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:12.420401Z digest=sha256:11546a24f4117d39335bbcacd8ade4a2a18d8b94bd50750b29aa8695767f7c9e

Observation 8158aad3-a061-4172-a5e1-392d3e4fa3a4 · inbound

lmgame-Bench: How Good are LLMs at Playing Games? cites this paper.

lmgame-Bench: How Good are LLMs at Playing Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:27:01.110876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:27:01.110876Z digest=sha256:cc457ef359346577ccccbbfd5b8c09633858c4cb5cf907b747de980a0eb97fc5

Observation 583020f1-6aaa-4059-a310-44715a23cdb8 · inbound

ImgEdit: A Unified Image Editing Dataset and Benchmark cites this paper.

ImgEdit: A Unified Image Editing Dataset and Benchmark MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T18:17:45.350173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T18:17:45.123690Z digest=sha256:7f53d13b1a42e3b33b1ede18cd2f74a2bdd840a9bc6ec42cf843df18bf1faecc

Observation 0e6f16b8-a732-4ca0-90e5-adf7914b8582 · inbound

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention cites this paper.

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:28:16.518006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T09:28:16.189617Z digest=sha256:c10db00d425ffa1d05a218d0bde51cdc4d635c4d676b7618fc4ee22541a5beb6

Observation c76d5d24-86a7-48e5-a1ca-4b014cc39af5 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:51.856959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:51.856959Z digest=sha256:45e4c4a6bc006b417f6fed83d55b50debc0710c1f42fc2f68b4ca2a6b8ad4079

Observation b2aa6ed7-d66e-413e-b8f1-d92c97ad80f3 · inbound

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains cites this paper.

Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.845701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:07:56.678339Z digest=sha256:cc48b4f027ffe7cf096005562fccfedc3d0b2cf293d0a73ed3fca36a93ee931f

Observation 0a513bc0-9837-42b3-8f72-4faf142f9ccd · inbound

TextQuests: How Good are LLMs at Text-Based Video Games? cites this paper.

TextQuests: How Good are LLMs at Text-Based Video Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.669131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.669131Z digest=sha256:7e6403dbcd570520ee0feabf82e54b370c3f132cc19de9d97111609db7bbe2a2

Observation 57df649d-5060-47bc-9a32-08469a0b455a · inbound

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries cites this paper.

OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T14:21:48.990246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:21:48.990246Z digest=sha256:70d61491601e8ecf75d88fa08867510000b0080b899b428c0984dba4567563ec

Observation 4cfd9f18-3805-4e8e-867d-263f705e7254 · inbound

Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting cites this paper.

Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T23:12:10.089168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T23:12:10.089168Z digest=sha256:00903189e229894a0f7aff32362b2a475e06fafe5c51cbfe9d587db896e5a0fb

Observation 5c091189-e56c-49de-bc03-2054b4e04159 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:39.919893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:39.919893Z digest=sha256:69ccf3a74de6da8785b3c6d57401c5289cc3bd89441d8d41327f2778488b0fe6

Observation dc68545d-532a-47d9-900e-a2d6861c90f4 · inbound

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks cites this paper.

Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T09:40:46.562258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:40:46.562258Z digest=sha256:3634a0c75de8fb538b15fe11b96ca8dbef1b33969616e6c7c1df4186f357b2ef

Observation 99759f4e-bdf1-4db9-bfec-fcf4a21a6ead · inbound

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory cites this paper.

Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:13:16.105119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-14T23:13:15.016486Z digest=sha256:f86b3a0e4ff0b06f15f7f5d4341155c66dee88adf9905c17468a0c36a3f1c157

Observation 6d2bcf41-ade4-4328-9738-7dbb80d785dc · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:50:50.469751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:5295a8f99d52c7c2306c46b6621d68d24bda8a660e21529f3044de92fa0eb9e1

Observation f6d4d5ec-bc87-4f69-a129-da250de9fe45 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.445640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.445640Z digest=sha256:0fc783ecba00facddd2e74cd7d2243a4c32f0653d9baa5bcd863186b4681927c

Observation f49b607a-92f2-41a3-8729-aab61af32184 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.360793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.360793Z digest=sha256:3bcdea87783c042d96e2e65538c95cd2fcd7cc800f76642b3d29f3066aa0cf52

Observation c5fb9061-7d35-49af-9232-0e9cad0e51a1 · inbound

RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation cites this paper.

RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 32

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T09:17:40.301640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:13:55.204404Z digest=sha256:f3f77424c17d6fadaaaf448af1a6d8f8b0c6ab24f0ebf0b4a1562cf1eef21e99

Observation 2983f954-59cb-4294-81c6-ac074e7e8e25 · inbound

Evolving and Detecting Multi-Turn Deception using Geometric Signatures cites this paper.

Evolving and Detecting Multi-Turn Deception using Geometric Signatures MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T15:23:32.697954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T15:17:58.804973Z digest=sha256:a09c8be74a4b06dff751533d5d3147d10508c20440340ae8d80af454d3354702

Observation 9b3ea34e-c4b4-41ff-bf02-b439edf5abf8 · inbound

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses cites this paper.

FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:56:44.740957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:16:30.725358Z digest=sha256:f25bfa41a022464f7d5407fe7b825bf13c4fa5fda3802e5dec1372d6d66ae944

Observation d3050268-e08e-479a-b75d-258cccea1f49 · inbound

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios cites this paper.

An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:48:46.351978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T03:39:36.657903Z digest=sha256:93ff5a30e079bc51048debef4dfab0db70a60c430aa380a3854058d141864c7c

Observation 36c02bfb-a6e5-40e5-8841-42a8188303c7 · inbound

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients cites this paper.

Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:48:56.199387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:08:52.981296Z digest=sha256:58be6c2bb7cae924ac4275c1c6cd86af946a940190571ef8c2021f73c5d3c257