Pith. sign in

Paper Citation Record · LEDGER

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

As of 10 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 6 inbound Pith citation observations for arXiv:2509.09614.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09614 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:49:58.363965Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:43:23.450401Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T02:19:23.892865Z

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 31972309-3f8b-451f-95fb-dd90fe48eaf4 · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.338747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.338747Z digest=sha256:190af7b411356c47b4a0d6682374e6b84e25cc030616653144124fb762a05021

Observation ea22d1db-4028-4a29-bc75-6c222f76bd34 · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.341602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.341602Z digest=sha256:e6d1a82bf9ffef4c1d7a7b7199a55a754e1111301df23d013f1153d9a45a56ac

Observation 5820baf9-d8c7-448a-bdda-ba8a5989d429 · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.332193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.332193Z digest=sha256:7b841fbb4dd610fd1466bac4086141c1a0a8d789c1780c33269550455a7bd384

Observation 87554561-0d91-41dc-89d0-91cf9cc53b0a · outbound

This paper cites "" def __init__(self): self._subscribers = \{\} self._async_subscribers = \{\} def subscribe(self, event_type: Type[Event], handler: EventHandler):.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering "" def __init__(self): self._subscribers = \{\} self._async_subscribers = \{\} def subscribe(self, event_type: Type[Event], handler: EventHandler):

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.335170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.335170Z digest=sha256:8a5601559e994ab7075deb4b6612db68f822dcef8b3f68f293ad7523dfae5dc6

Observation 32b045b8-b3fb-44f7-a6e1-c607363c9379 · outbound

This paper cites APPS provides 10,000 problems from coding compe- titions, while LiveCodeBench offers contamination-free evaluation with problems collected from ongoing contests.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering APPS provides 10,000 problems from coding compe- titions, while LiveCodeBench offers contamination-free evaluation with problems collected from ongoing contests

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.319394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.319394Z digest=sha256:ad554249d8238dd5b78633292419307bc233487fd092abb5044f8e64964a31b8

Observation a814b9c7-4c01-4532-b0ef-41260f5112be · outbound

This paper cites While domain-specific, these benchmarks still primarily evaluate isolated function or script generation rather than comprehensive software development capabilities.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering While domain-specific, these benchmarks still primarily evaluate isolated function or script generation rather than comprehensive software development capabilities

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.322473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.322473Z digest=sha256:d501d030a94aa345a549986e678e1f8ebe3ac43932319142e352f44d1acede27

Observation 2dba29e3-50f2-41fb-a81f-770d684cb2e8 · outbound

This paper cites LongICLBench (An et al., 2024) evaluates in-context learning capabilities at extreme lengths, while LongAlign (Bai et al., 2024a) addresses instruction following in long contexts.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering LongICLBench (An et al., 2024) evaluates in-context learning capabilities at extreme lengths, while LongAlign (Bai et al., 2024a) addresses instruction following in long contexts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.325923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.325923Z digest=sha256:c3f0f39ab651279fc99269fd143afecff730a72e4e5015936e44b54f33668218

Observation 02347c3b-1dad-466a-af39-59e4d75759cf · outbound

This paper cites id": "java_api_graphql_easy_007_feature_implementation_expert_01.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering id": "java_api_graphql_easy_007_feature_implementation_expert_01

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.344904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.344904Z digest=sha256:94db9703fe4b6efefa0c191983f8c1efce6f2cc1f057c791795943277ea3530e

Observation 94a9b891-e205-4e93-bcad-286cc5e62bfd · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.348184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.348184Z digest=sha256:646c80d4fa6fb921f6a5a2cfc72f6f8ba4e5981d6b7f1ceb07fab245a26afea8

Observation 8b2af64e-2b9a-4db6-ad1c-6a30f3bfe3a3 · outbound

This paper cites GPT-4o": \{.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering GPT-4o": \{

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.351595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.351595Z digest=sha256:d9e071097fbfa5ba676fa9ae2fe97c82fae466c55d61e8511f01440e950c66ee

Observation 1152389c-f47e-45c2-a503-f4d0739036f1 · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.354651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.354651Z digest=sha256:d33f35ef2dba44310a122e19551fc366b7f185548ee722ee1a5d6c7e48a0a853

Observation 8f7081cd-3520-40e3-ba44-29964d933cec · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.357809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.357809Z digest=sha256:2398136716037fe8ebd9e937e73e2350fc2a68400dc52327556f1f3891aefce1

Observation d67603e4-9bd2-453f-81cd-f1b033d84f20 · outbound

This paper cites an unresolved cited work.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.360736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.360736Z digest=sha256:88b0669e207ff328e3df8a5d0ec2569c88f3a2fa3daf44bb5a716254ae219956

Observation a832a7d4-8487-4c5c-9768-2933db128c07 · outbound

This paper cites approach.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering approach

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.363965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.363965Z digest=sha256:58f0b36c0efd6d6da6e7f68ecd67261c72b21c31bfd6285b7be09b3b55a0dd48

Observation 82bdb0c9-8baf-4baf-a895-7bfd9ded53c8 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Long-context LLMs Struggle with Long In-context Learning

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.304178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.304178Z digest=sha256:68168a27d2975f568e6eb51c035a2f08dfd21a30fca25b0bd9c126fd38741a1a

Observation 3ebd0d67-0c64-4c58-b5c5-3616d57863e9 · outbound

This paper cites CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.308701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.308701Z digest=sha256:b398c4f9fd4731ae6857733e490f28d0978710ccbf5f5f96764a20f4bca56761

Observation ae4eccc4-520f-40c7-9b76-c1efa36ec6dc · outbound

This paper cites CoCoMIC: Code Completion By Jointly Modeling In-file and Cross-file Context.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering CoCoMIC: Code Completion By Jointly Modeling In-file and Cross-file Context

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.312286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.312286Z digest=sha256:2ef2f37c6045d2f198764000dc928ab3659c80ac652e10f132857dc344448e5c

Observation 5420b973-b513-4a44-9930-615871ecb629 · outbound

This paper cites Feature Representations for Automatic Meerkat Vocalization Classification.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering Feature Representations for Automatic Meerkat Vocalization Classification

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.315651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.315651Z digest=sha256:fd53c6edf30bddc5aafaab61f2e6931ce7ff3e616d4f75d39b7a820dafaf7902

Observation afe80bab-3d70-478d-92a0-7d3725f73b08 · outbound

This paper cites unique_id.

LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering unique_id

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:58.328653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:58.328653Z digest=sha256:881744b28ca28c823cd74f9c07dba833587324607b61765a73ebc69cfe2f1065

Pith citing papers

Observation 36ac7432-fe0d-4fc0-ae88-941abf859042 · inbound

Teaching LLMs Program Semantics via Symbolic Execution Traces cites this paper.

Teaching LLMs Program Semantics via Symbolic Execution Traces LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:26:13.431111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T08:56:25.022619Z digest=sha256:ae57832b9618f2f9d7ad52b6b8d47a4554dca7d592b17c7674448cdcf9674032

Observation 2f3bccff-abd0-4bfc-8f9e-235cc02ea999 · inbound

Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems cites this paper.

Evolve as a Team: Collaborative Self-Evolution for LLM-based Multi-Agent Systems LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:12:50.100767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T00:05:31.780655Z digest=sha256:3537a2a7313250c46dafd3ca572640048798a24e6805cf15dcd58384c1849dd5

Observation 50d8dcb4-9a02-4d1b-9532-237f5fc10dd5 · inbound

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns cites this paper.

StaminaBench: Stress-Testing Coding Agents over 100 Interaction Turns LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:19:23.894455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T19:47:59.090249Z digest=sha256:2a6389c434a38e90ca4db2ca8c2d76911b298bf035851eb6d58b489733bb7c42

Observation 33cea02d-7582-4710-bc45-154befe7621f · inbound

LLM4CAD-Editor: An Intent-Aware Large Language Model Framework for Multi-Level Computer-Aided Design Editing cites this paper.

LLM4CAD-Editor: An Intent-Aware Large Language Model Framework for Multi-Level Computer-Aided Design Editing LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.857778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T15:33:46.125268Z digest=sha256:46ef7e017faaeb74d9227728b6a71d1bd134e16c671e47ddda5cea414f7b146c

Observation 0533a6b0-74ee-4240-9432-ce915d1fde6d · inbound

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists cites this paper.

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:45:20.185074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:45:20.185074Z digest=sha256:4fc83a4bfe060350f12677dd9bda88f56f8d7b96b58b296d1e933a035e63c036

Observation 766a219c-b512-493a-88dd-24965e3e3de0 · inbound

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists cites this paper.

RepoProbe: Benchmarking Architecture-Aware Repository Comprehension with Checklists LoCoBench: A Benchmark for Long-Context Large Language Models in Complex Software Engineering

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T17:43:23.450401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:43:23.450401Z digest=sha256:4e9aae69db4f52c8d1fb9ffcb867622d6bb128f190dab84a39721b8a7b85aa39