Pith. sign in

Paper Citation Record · LEDGER

L-Eval: Instituting Standardized Evaluation for Long Context Language Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2307.11088.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2307.11088 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:40:27.024729Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d221e3e8-9e3c-49c4-bd7f-865d78f8f29e · inbound

LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding cites this paper.

LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:22:10.690217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T20:22:10.482509Z digest=sha256:59da7efdb70f8ca4f0e5e9d4ad201ab1e2e95fcfa8c5613840ac1bce28920952

Observation c8d343fb-e86f-43e3-9e17-4839fb1389e0 · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-15T11:44:38.169759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:e4818aa83971260fccf70b83df04ab6a17259f5ffe6e540179903f6e47697926

Observation aa329656-4e69-4293-9776-6404ff6d1aec · inbound

Jamba: A Hybrid Transformer-Mamba Language Model cites this paper.

Jamba: A Hybrid Transformer-Mamba Language Model L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T14:11:27.181188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T14:11:27.156350Z digest=sha256:ad6911924dfc80a6a4967fd13ddeeabd596ee3b7eebf3ed159a45114f054fd00

Observation 348f8151-5ca1-48bc-9cdc-d068db0934af · inbound

SnapKV: LLM Knows What You are Looking for Before Generation cites this paper.

SnapKV: LLM Knows What You are Looking for Before Generation L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T12:57:43.115018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T12:57:43.036674Z digest=sha256:72ceb055b5aa1c75ee9eb163622fb7c3b69a7e0daf7a8ec7ef0180643ce5a4bf

Observation 708c150c-2790-420c-91d6-24f836508ca8 · inbound

A Survey on LLM-as-a-Judge cites this paper.

A Survey on LLM-as-a-Judge L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T17:35:44.190233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T17:33:13.394338Z digest=sha256:1c01612d0a76501cb335ac5f398c3d6571e3c2d1db31d5ad7883a15e4eb94d14

Observation 20812d0f-3b1c-40d3-b55f-2f68fbd2ea98 · inbound

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression cites this paper.

Semantic Integrity Matters: Benchmarking and Preserving High-Density Reasoning in KV Cache Compression L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T04:17:31.131311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T04:15:36.906263Z digest=sha256:7b35729718863f4de8713dc9778aac6351264cf534150ca283021c6f7d1cf483

Observation 0b9aa49d-cfd0-4eeb-8e50-347e486a28c1 · inbound

LCIRC: A Recurrent Compression Approach for Efficient Long-form Context and Query Dependent Modeling in LLMs cites this paper.

LCIRC: A Recurrent Compression Approach for Efficient Long-form Context and Query Dependent Modeling in LLMs L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T16:40:27.024729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:40:27.024729Z digest=sha256:c4b1c7803cd94a55af4da33c7765845f3039b4b97678d04d777f065d4d946911

Observation 33655849-7d56-47b8-bcb1-57779644c1ff · inbound

RoToR: Towards More Reliable Responses for Order-Invariant Inputs cites this paper.

RoToR: Towards More Reliable Responses for Order-Invariant Inputs L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:43.095401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:43.095401Z digest=sha256:b96abd3e4b2698e6a8d8397971c009e46502893afdfe381a4690652be5ea783f

Observation 70a77ed3-9a1e-415b-9807-bd91984977d1 · inbound

SELF: Self-Extend the Context Length With Logistic Growth Function cites this paper.

SELF: Self-Extend the Context Length With Logistic Growth Function L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:52:12.226319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:52:12.226319Z digest=sha256:0ac2bf5fcab591b616b517aacb0173379ffd5e1575debd2f508c74533ed77748

Observation c728fe26-7aa8-4634-9e18-24b52d96d5b0 · inbound

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? cites this paper.

100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability? L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:24.321920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:24.321920Z digest=sha256:b7cc7d2cb830b603acbe8f0b975ef4aba8df68f9506a56c88fcedb2b93678630

Observation c2baec95-7bfa-4e92-aac5-fcc38b2ae39b · inbound

MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models cites this paper.

MiniLongBench: The Low-cost Long Context Understanding Benchmark for Large Language Models L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:08:05.311060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:08:05.311060Z digest=sha256:21d8f8b01013197bd18d18230aa934b96892bf83cb88602a026de2107dd4c09d

Observation bdf90adb-491d-4380-8de7-725ff8812b0c · inbound

NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts cites this paper.

NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative Contexts L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:31:42.196733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:31:42.196733Z digest=sha256:5032381691da721a182cdeb3690f3212325430cea80784e6f7114488dbb3a72a

Observation 56e6fd7a-77b1-4878-bb0d-11a3c679c1f8 · inbound

AbsenceBench: Language Models Can't Tell What's Missing cites this paper.

AbsenceBench: Language Models Can't Tell What's Missing L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:13:17.656859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:13:17.656859Z digest=sha256:a912561b243054e3ac47c0112596d14e85c12dfc88922b6b177e688a986b6c7e

Observation 82c4ea3c-aa9d-4bed-b09c-22bdf7a69a5f · inbound

LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions cites this paper.

LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:49:24.771230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:49:24.771230Z digest=sha256:a0f669bbc0f2c0960d094846be49a9ce7f75e9a4c5a0e6c8fa1fb4131e1ceea4

Observation 33b0b05f-9ba4-49ce-8965-5daa01945269 · inbound

MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents cites this paper.

MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:28.878915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:42:28.878915Z digest=sha256:c47249271489700553788e63cb85c53ba94042928066d6126f3add4b36c737aa

Observation f45e1cf2-8d80-4731-a64b-81c54cb0c73a · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:56:57.854965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:56:57.854965Z digest=sha256:61db8f9780cfb54d82907239209264b82465496359e0886309f11d2c98170240

Observation 8e835c7a-4d13-4570-86c9-3b8f29da5fc3 · inbound

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs cites this paper.

Not All Needles Are Found: How Fact Distribution and Don't Make It Up Prompts Shape Retrieval, Reasoning, and Hallucination in Long-Context LLMs L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T12:42:30.058516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:42:30.058516Z digest=sha256:d909561ee864790864382c1cc08285ba4ea5a1d81c967a92e081e292ffe681b6

Observation 797f398f-5c61-45d9-acde-db8426259afe · inbound

Whose Story Gets Told? Positionality and Bias in LLM Summaries of Life Narratives cites this paper.

Whose Story Gets Told? Positionality and Bias in LLM Summaries of Life Narratives L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-10T01:04:50.201027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T01:00:41.543394Z digest=sha256:7dbc25f357933de483813809675dcf127ea4c2ebdd6e01ae2e612ce458e455cc

Observation 96ae3ef3-36cd-4cf3-b0b5-c422572678f0 · inbound

Efficient Training on Multiple Consumer GPUs with RoundPipe cites this paper.

Efficient Training on Multiple Consumer GPUs with RoundPipe L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:31:26.839097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T10:37:22.251566Z digest=sha256:b1dd71d6f1373caba835314078b1f26af0f2bfaef6495601de903db41af94120

Observation d630b536-3192-4570-86ac-9e54bffbd585 · inbound

Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving cites this paper.

Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:31:10.427339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T16:19:33.613685Z digest=sha256:d6d1bbb7436f235aa392b3ed3e8a3a04b31b9376a321ac3b70556e4d834ddde3

Observation 8b982927-667f-479e-936f-73af253b2cc8 · inbound

EndPrompt: Efficient Long-Context Extension via Terminal Anchoring cites this paper.

EndPrompt: Efficient Long-Context Extension via Terminal Anchoring L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T01:33:27.224830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T01:32:32.967922Z digest=sha256:7945fe4eba1a1bcc84f9615db5f76365d488ce63cf1022af21376eafcb2a4dd4

Observation 1375033c-3ef9-4768-b005-fb1777f324a1 · inbound

EndPrompt: Efficient Long-Context Extension via Terminal Anchoring cites this paper.

EndPrompt: Efficient Long-Context Extension via Terminal Anchoring L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:15:03.983326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T21:13:23.148984Z digest=sha256:df2a04d0bd864e152b2cc9b78190319dad4c75ceb155783c96efb70bc7d6764e

Observation caa6a0ab-d8f0-4702-aa93-ccb1681042ee · inbound

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks cites this paper.

Positional Failures in Long-Context LLMs: A Blind Spot in Reasoning Benchmarks L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:00:21.921291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-25T04:58:15.184063Z digest=sha256:105925fd28a4bb924d48014d341b266ee2dd54b7705fe5e02964e094e8575388

Observation 9fb53eed-2b60-435b-b151-dde4fd5a5034 · inbound

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data cites this paper.

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-06-30T13:54:44.114006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T13:45:38.305767Z digest=sha256:ec675b60312c80889de45bcc4d9b93f56058fcfc4a2b297b5d6848bb90fbef79

Observation 2c31c983-3808-4be8-bb80-78fd0a9f8580 · inbound

NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama cites this paper.

NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T19:28:52.462812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:48:08.794210Z digest=sha256:9612cfeacaa69dc045b3bb1dca900ee90b110cf24a7e74dc5af1f898257eec7d

Observation dbaa4e11-876c-4eb7-a43f-42d4accbe3c0 · inbound

Mitigating Position Bias in Transformers via Layer-Specific Positional Embedding Scaling cites this paper.

Mitigating Position Bias in Transformers via Layer-Specific Positional Embedding Scaling L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:03:51.865567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T04:59:00.304723Z digest=sha256:b27be32dff3a808549c963ce7c9f2a6bc80ac2f7fe1f220119c64a1bb71e93cd

Observation a077d62c-96f8-46b0-b0c1-5554bd719916 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.866244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.866244Z digest=sha256:f4231dcc910eb19e8ad018cdabd63ef2be4967c73a1307c3078ac596e576d7f0

Observation 1903ada6-956f-477c-8e7f-bdaae9c022d4 · inbound

Memory for Large Language Models cites this paper.

Memory for Large Language Models L-Eval: Instituting Standardized Evaluation for Long Context Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-01T02:37:54.577072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:37:54.577072Z digest=sha256:70379e94b51f8a257f779f166787b7c0245147e37e4f6da0535e0eac6b7a3f07