Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:00:46.383924Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 7 inbound Pith citation observations for arXiv:2505.10570.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:00:46.383924Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T00:39:14.741354Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:10:05.100242Z
30 of 30 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 857def1e-ad16-44e8-a419-b52d9f00350d · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 705b8fb1-28dd-4e4b-8ead-4d5433734abc · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling DENIAHL: In-Context Features Influence LLM Needle-In-A-Haystack Abilities
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e494b8de-87c9-4c30-b4c0-f905ace11cb1 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f191ab5e-5934-4448-b23c-6214b6ddc9bc · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 88e811ce-ee0b-4ed1-883c-e8f7fb39f5b6 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Is It Really Long Context if All You Need Is Retrieval? Towards Genuinely Difficult Long Context NLP
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51ad18ae-ab93-469a-b53b-70d55f53d303 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling RULER: What's the Real Context Size of Your Long-Context Language Models?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c107c3-8267-4296-aa82-29c54857211b · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 27be5343-079f-495e-a7fb-5bb88c59fd46 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d0251190-5c29-4782-9831-1875cfc53129 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling LooGLE: Can Long-Context Language Models Understand Long Contexts?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ff31dc9-6c1b-4e87-8bb4-0d0db012f851 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b5cdf38a-bb8a-4cc2-bcd3-a2c170556dda · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; and Liang, P
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 51fa21da-7efe-4b1e-ab28-361b5f70c162 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling LongGenBench: Long-context Generation Benchmark
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a6f6e3-2253-436a-9ffe-64af38ac3585 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8d93e248-7c9e-41c2-a042-f8f162a6616d · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Toolshed: Scale Tool-Equipped Agents with Advanced RAG-Tool Fusion and Tool Knowledge Bases
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff3280ac-a549-44c2-8988-0f06420c9cfa · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 40891f62-1c8f-433e-b2f0-ac318bdf3d2d · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling NoLiMa: Long-Context Evaluation Beyond Literal Matching
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dde9595c-865f-4d40-a7ec-eee43889faa9 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c3de95e6-2c67-4f38-9475-4e8fc0061433 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling G.; Zhang, T.; Wang, X.; and Gonzalez, J
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 24dd453f-19ff-4c47-bbb1-b980cd42b369 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8f16f33f-edff-4524-9d67-a8a15f7422fc · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e18dbf84-9fc2-47d7-b834-2ca731d6a29e · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling J.; and Hashimoto, T
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dd8c6edc-57ca-4b4c-9623-533139f2911f · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6b1fa5e0-0118-494e-9098-ed4ab3e378e0 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Retrieval Models Aren't Tool-Savvy: Benchmarking Tool Retrieval for Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e428010-d7e4-47c1-9268-a70f8e363b8f · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling RestGPT: Connecting Large Language Models with Real-World RESTful APIs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bbd6cd8-da35-4fa3-97de-d84d9cfcdf40 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 28b346c6-1b9a-4e5d-9fb7-a5b1b587362e · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 316fdc99-a48d-49d4-8bfc-6758388b523e · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 64ae2a4b-c380-4064-96e4-775ae14d5668 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling C.-J.; Zhang, T.; Patil, S
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a93d73c2-c3d5-4841-a869-e210694e2934 · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling T.; Chang, K.-W.; Lee, C.-Y.; Palangi, H.; and Pfister, T
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bdfe515a-032b-41bf-a308-a7336739d78d · outbound
LongFuncEval: Measuring the effectiveness of long context models for function calling ComplexFuncBench: Exploring Multi-Step and Constrained Function Calling under Long-Context Scenario
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70870885-cd69-41ea-b6b8-cccbbebe8fe0 · inbound
Gecko: A Simulation Environment with Stateful Feedback for Refining Agent Tool Calls LongFuncEval: Measuring the effectiveness of long context models for function calling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf1952a5-4c2f-4bad-9013-2157c60342ae · inbound
FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks LongFuncEval: Measuring the effectiveness of long context models for function calling
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation dc8f1e4a-774b-4248-be39-8f6a3245cceb · inbound
How Many Tools Should an LLM Agent See? A Chance-Corrected Answer LongFuncEval: Measuring the effectiveness of long context models for function calling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8bf300bd-e075-4c1c-a356-1967bba2f63e · inbound
Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery LongFuncEval: Measuring the effectiveness of long context models for function calling
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5a9a42e0-2d43-47fa-92b5-35fe81990d19 · inbound
Buildrix: An Open Platform for Sharing and Benchmarking Agentic AI Skills in Building Engineering LongFuncEval: Measuring the effectiveness of long context models for function calling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cedc7436-85a1-4888-a455-37fd9356302d · inbound
MCP Server Architecture Patterns for LLM-Integrated Applications LongFuncEval: Measuring the effectiveness of long context models for function calling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6c87a667-9ea6-421d-9a0a-a0c5a32978ed · inbound
Don't Offer What Can't Be Done: Deterministic Executability Gating for LLM Skill Selection at Scale LongFuncEval: Measuring the effectiveness of long context models for function calling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.