Pith. sign in

Paper Citation Record · LEDGER

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

As of 13 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 6 inbound Pith citation observations for arXiv:2509.09614.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09614 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:49:58.363965Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:43:23.450401Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:19:23.892865Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31972309-3f8b-451f-95fb-dd90fe48eaf4 · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.338747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.338747Z digest=sha256:8b8143e24c37b7608f838672ea5a2788c08c4bb7064ee65c11d65e297cf1b73a

Observation ea22d1db-4028-4a29-bc75-6c222f76bd34 · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.341602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.341602Z digest=sha256:a0ffe5ce53dc21fdeae1afae8be89a73c0f103d6828009a24e7e6a24214f3f27

Observation 5820baf9-d8c7-448a-bdda-ba8a5989d429 · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.332193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.332193Z digest=sha256:5aa4f3eedb940581f79add30421f27e78790d90dba8587a27ed8a9064bb02b3b

Observation 87554561-0d91-41dc-89d0-91cf9cc53b0a · outbound

This paper cites "" def __init__(self): self._subscribers = \{\} self._async_subscribers = \{\} def subscribe(self, event_type: Type[Event], handler: EventHandler):.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering "" def __init__(self): self._subscribers = \{\} self._async_subscribers = \{\} def subscribe(self, event_type: Type[Event], handler: EventHandler):

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.335170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.335170Z digest=sha256:d1776d3428ac968a940ce38da84a3d1048aa8c41d244d9b15a7398cbf0506c6a

Observation 32b045b8-b3fb-44f7-a6e1-c607363c9379 · outbound

This paper cites APPS provides 10,000 problems from coding compe- titions, while LiveCodeBench offers contamination-free evaluation with problems collected from ongoing contests.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering APPS provides 10,000 problems from coding compe- titions, while LiveCodeBench offers contamination-free evaluation with problems collected from ongoing contests

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.319394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.319394Z digest=sha256:7005c8dccda3405c28fd9f8cf6bb06f65e2b94fe39351c0f01bf5a160631f378

Observation a814b9c7-4c01-4532-b0ef-41260f5112be · outbound

This paper cites While domain-specific, these benchmarks still primarily evaluate isolated function or script generation rather than comprehensive software development capabilities.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering While domain-specific, these benchmarks still primarily evaluate isolated function or script generation rather than comprehensive software development capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.322473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.322473Z digest=sha256:2ba5aa7e734bf5b71645d3933ce4153e230de90095bc2da0a6184c531fd491a1

Observation 2dba29e3-50f2-41fb-a81f-770d684cb2e8 · outbound

This paper cites LongICLBench (An et al., 2024) evaluates in-context learning capabilities at extreme lengths, while LongAlign (Bai et al., 2024a) addresses instruction following in long contexts.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering LongICLBench (An et al., 2024) evaluates in-context learning capabilities at extreme lengths, while LongAlign (Bai et al., 2024a) addresses instruction following in long contexts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.325923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.325923Z digest=sha256:353cd9f9a995021192a44b520367e1a5c305609a7dc6935cceb5875427468efa

Observation 02347c3b-1dad-466a-af39-59e4d75759cf · outbound

This paper cites id": "java_api_graphql_easy_007_feature_implementation_expert_01.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering id": "java_api_graphql_easy_007_feature_implementation_expert_01

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.344904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.344904Z digest=sha256:bd12ec5730cd1fc75402e383753d845e574815431e08f9ed6f794ec7e503a0e4

Observation 94a9b891-e205-4e93-bcad-286cc5e62bfd · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.348184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.348184Z digest=sha256:3526e835c48daeeec4e0c7f9adb5ab6b9edc8553f00d55c6f17864f9cddd2229

Observation 8b2af64e-2b9a-4db6-ad1c-6a30f3bfe3a3 · outbound

This paper cites GPT-4o": \{.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering GPT-4o": \{

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.351595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.351595Z digest=sha256:1cd961e31a79ddf8529b265e1425537b6ceab5f981128bd0ed7650b24533672a

Observation 1152389c-f47e-45c2-a503-f4d0739036f1 · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.354651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.354651Z digest=sha256:5df6ffebb58b759943311bbee54abaeb5003ff176fd36bc3c33b2f44cafa49a8

Observation 8f7081cd-3520-40e3-ba44-29964d933cec · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.357809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.357809Z digest=sha256:5476d787a6833c93655ef8b9df8eac1c6c16aa91537d3f053fd0260d8846d305

Observation d67603e4-9bd2-453f-81cd-f1b033d84f20 · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.360736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.360736Z digest=sha256:bc4f9f522511b8a44505d1a612c51369ec09c02a37ec9deadacdb4668c1d8ee0

Observation a832a7d4-8487-4c5c-9768-2933db128c07 · outbound

This paper cites approach.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering approach

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.363965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.363965Z digest=sha256:372fb675a1480ac94e25d8b8889019e5abb8381cbe2c2a195fa4c5dba98f36fa

Observation 82bdb0c9-8baf-4baf-a895-7bfd9ded53c8 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Long-context LLMs Struggle with Long In-context Learning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.304178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.304178Z digest=sha256:8d304d97994ab0ebb5d0bd95e224cca6783a4a1b54636a8ed81122eca7d8f224

Observation 3ebd0d67-0c64-4c58-b5c5-3616d57863e9 · outbound

This paper cites CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.308701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.308701Z digest=sha256:b4e4b2553984b491ac677be96f8a7f5932dfc7d8bee5c6b0f0eb8a1e2dd34f93

Observation ae4eccc4-520f-40c7-9b76-c1efa36ec6dc · outbound

This paper cites CoCoMIC: Code Completion By Jointly Modeling In-file and Cross-file Context.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering CoCoMIC: Code Completion By Jointly Modeling In-file and Cross-file Context

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.312286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.312286Z digest=sha256:54b2f2eeb0d5bf4175fd03e2f5927d40a88f5f24e885505054d01b3eb0af4189

Observation 5420b973-b513-4a44-9930-615871ecb629 · outbound

This paper cites Feature Representations for Automatic Meerkat Vocalization Classification.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Feature Representations for Automatic Meerkat Vocalization Classification

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.315651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.315651Z digest=sha256:ff9de1bf427adb9d06d80438e6fccb97e862fe354e1916635284a17fdada07c2

Observation afe80bab-3d70-478d-92a0-7d3725f73b08 · outbound

This paper cites unique_id.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering unique_id

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.328653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.328653Z digest=sha256:c2717827196ff0b1af3697ab9646dc25a6481839e7dd127fcd2b3521dc23d9e1

Pith citing papers

Observation 36ac7432-fe0d-4fc0-ae88-941abf859042 · inbound

Teaching LLMs Program Semantics via Symbolic Execution Traces cites this paper.

Teaching LLMs Program Semantics via Symbolic Execution Traces LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:13.431111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T08:56:25.022619Z digest=sha256:2c1da86dada83cb7893c9225f296fab998cb9cf6f01af82475bd60cd9e0646cd

Observation 2f3bccff-abd0-4bfc-8f9e-235cc02ea999 · inbound

Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems cites this paper.

Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.100767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T00:05:31.780655Z digest=sha256:e7b8ab6e3dc6e52061ea0a754570d5d091714152ab701343b531390984d07ca4

Observation 50d8dcb4-9a02-4d1b-9532-237f5fc10dd5 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.894455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:287ae9d24ef8ee3a3728ba0d98251035df30696ae3b26b6f834b472432d72b9f

Observation 33cea02d-7582-4710-bc45-154befe7621f · inbound

LLM4CAD-Editor: An Intent-Aware Large Language Model Framework for Multi-Level Computer-Aided Design Editing cites this paper.

LLM4CAD-Editor: An Intent-Aware Large Language Model Framework for Multi-Level Computer-Aided Design Editing LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.857778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T15:33:46.125268Z digest=sha256:866e64ff17bfcb9739da992edcac2cb3304d2760e1e2e4c8241b84c65198a6d8

Observation 0533a6b0-74ee-4240-9432-ce915d1fde6d · inbound

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists cites this paper.

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:20.185074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:45:20.185074Z digest=sha256:3bd86a14e17185ed0c6c244efe5ecb5cf03753ba1207178ad2bd80876b925c41

Observation 766a219c-b512-493a-88dd-24965e3e3de0 · inbound

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists cites this paper.

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T17:43:23.450401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:43:23.450401Z digest=sha256:93b0b1a2350a5c4bb36c3c25c84c45a70b825deff410709820192b32e1d4504d