Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2501.17399.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:33:12.420401Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T20:48:56.198096Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 140c103e-9697-47e7-b78b-0a93bb3b179d · inbound
LLMs Get Lost In Multi-Turn Conversation MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6e46e7b3-0481-483d-8932-7b2ab6cebf05 · inbound
Scaling Reasoning, Losing Control: Evaluating Instruction Following in Large Reasoning Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8158aad3-a061-4172-a5e1-392d3e4fa3a4 · inbound
lmgame-Bench: How Good are LLMs at Playing Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 583020f1-6aaa-4059-a310-44715a23cdb8 · inbound
ImgEdit: A Unified Image Editing Dataset and Benchmark MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0e6f16b8-a732-4ca0-90e5-adf7914b8582 · inbound
MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c76d5d24-86a7-48e5-a1ca-4b014cc39af5 · inbound
A Conceptual Framework for AI Capability Evaluations MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2aa6ed7-d66e-413e-b8f1-d92c97ad80f3 · inbound
Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0a513bc0-9837-42b3-8f72-4faf142f9ccd · inbound
TextQuests: How Good are LLMs at Text-Based Video Games? MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57df649d-5060-47bc-9a32-08469a0b455a · inbound
OpenAIs HealthBench in Action: Evaluating an LLM-Based Medical Assistant on Realistic Clinical Queries MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cfd9f18-3805-4e8e-867d-263f705e7254 · inbound
Another Turn, Better Output? A Turn-Wise Analysis of Iterative LLM Prompting MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c091189-e56c-49de-bc03-2054b4e04159 · inbound
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 152
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc68545d-532a-47d9-900e-a2d6861c90f4 · inbound
Echoes of Human Malice in Agents: Benchmarking LLMs for Multi-Turn Online Harassment Attacks MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99759f4e-bdf1-4db9-bfec-fcf4a21a6ead · inbound
Evo-Memory: Benchmarking LLM Agent Test-time Learning with Self-Evolving Memory MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 131
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6d2bcf41-ade4-4328-9738-7dbb80d785dc · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f6d4d5ec-bc87-4f69-a129-da250de9fe45 · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f49b607a-92f2-41a3-8729-aab61af32184 · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5fb9061-7d35-49af-9232-0e9cad0e51a1 · inbound
RAG-DIVE: A Dynamic Approach for Multi-Turn Dialogue Evaluation in Retrieval-Augmented Generation MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2983f954-59cb-4294-81c6-ac074e7e8e25 · inbound
Evolving and Detecting Multi-Turn Deception using Geometric Signatures MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b3ea34e-c4b4-41ff-bf02-b439edf5abf8 · inbound
FailureScope: Cross-Regime Behavioral Diagnosis of Language Model Weaknesses MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3050268-e08e-479a-b75d-258cccea1f49 · inbound
An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36c02bfb-a6e5-40e5-8841-42a8188303c7 · inbound
Zone of Proximal Policy Optimization: Teacher in Prompts, Not Gradients MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs
Reference 155
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.