Pith. sign in

Paper Citation Record · LEDGER

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models

As of 13 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 2 inbound Pith citation observations for arXiv:2411.18019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.18019 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:39:33.837756Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T16:27:43.489582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T08:26:01.477245Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 09da401d-680c-41cb-9c88-4677228200fa · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.556342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.556342Z digest=sha256:5ae65ef22eea6a482881935d83ec36e7f9b25b5bf8ad4cdbd371a7953b250c8e

Observation 218b4b6e-83fb-454b-90ab-4f796936a9cd · outbound

This paper cites Program Synthesis with Large Language Models.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.561080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.561080Z digest=sha256:8085c1bc7f76e4a7ddafebe332ec237294f65bf0a0bf90b7e305ac66025f0697

Observation ff833103-d0a0-4aef-90a7-92706b999a49 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.330935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.565553Z digest=sha256:e56bcd16e0e25da401d76eff07fec6f5c6be9d5d3d295b870bb728af4afb6f2e

Observation 7c550b34-48bc-4c13-90f6-7edb3122a433 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.569353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.569353Z digest=sha256:737f894d49da57611abf6629237dd48f04afd7d902acb83ae56675683ee0f02e

Observation c385cfc2-44c4-4d5b-aa5d-3efc4dc463ad · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.573403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.573403Z digest=sha256:4f2baed5aca29d6d19405213f531eab4d92ecc04024605743fb70740190c3061

Observation f6e0fc25-e88b-492b-a1cc-8fcd40f42907 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.312866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.577316Z digest=sha256:947868ec0c07d9256dc6a15a41c2db2b957e72b6c9034cbcba9cf58a697b6cdc

Observation d36b637b-25cd-40f7-a11c-f28f7f5f6516 · outbound

This paper cites Generalization-Enhanced Code Vulnerability Detection via Multi-Task Instruction Fine-Tuning.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Generalization-Enhanced Code Vulnerability Detection via Multi-Task Instruction Fine-Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.581251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.581251Z digest=sha256:1afd4a4508961a2c6dcb96c62077584b4e7cd71a040896b8cf2f36f9bf5aea3b

Observation 0e32bc69-8746-42ce-84c1-3048914e325c · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.301704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.585128Z digest=sha256:409d8ceba1db19747de2f03445fcb4d6c0b9e7ebda28ddc6b7ca392a22f42023

Observation 8993847b-e09b-4c02-a348-b535ba8a88ff · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.588766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.588766Z digest=sha256:948010e0330079b137c55a28696859b4eca8839d3aae6527c9576e4470e835c9

Observation 3c90e457-6e5d-4901-8970-c3ae95b5f194 · outbound

This paper cites CodeApex: A Bilingual Programming Evaluation Benchmark for Large Language Models.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models CodeApex: A Bilingual Programming Evaluation Benchmark for Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.593222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.593222Z digest=sha256:4afe2d9a13cc889069e92453b2d8c17c054c92b47c891bf0597c3f4fe9833e74

Observation b66db029-2a51-42cc-9860-3601c0dda63f · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.712047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.712047Z digest=sha256:5b480fc55d47d433fc139f47c29b9b11e93dd70d701ff6cf30d9f1522a424cc9

Observation 603ebf6a-a91d-4a0e-8569-ee18c40581ee · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.270351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.719649Z digest=sha256:1a28a18ce26feb411f0051f6a6d6194b5d556662db8777892fc28e2574bd8ff0

Observation f9094f8b-fd9f-4cbd-b51f-8f7bd58ae0bc · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Measuring Coding Challenge Competence With APPS

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.723218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.723218Z digest=sha256:2f5ec86162ab51f0a025b4f82ac3f484dab19f078ea4933b58862e8d1af3ffba

Observation 3d4846ff-6186-436e-8bcb-25942094c996 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.257902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.727099Z digest=sha256:421dd12f5809edd6528ae42fe4383ca8488b9f1cbe838a1f1c30ec5c6f3099f0

Observation f91bd4fa-8d70-4197-8546-c6aa0eb04328 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.730593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.730593Z digest=sha256:3bc792ad98261bd1de1f409d49ccbbbf9d296cdb6641c828ebb0845bc9f291e3

Observation 138d75c1-91f8-4f0e-93b1-dfaa283e39ab · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.237292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.738942Z digest=sha256:e274dacba2bb30459bc29aef01aa439a0123f23c0f56ed27a071a67b04f982de

Observation 4e0c54d3-92cd-429a-bede-39ee02206601 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.742992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.742992Z digest=sha256:d567ba73ec97b4b6570fab67c879fe91f04b1a01af03015f4c2f886e9141b256

Observation 3cb517ee-ad6b-48d5-a850-253c1869d346 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.746615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.746615Z digest=sha256:e9a99c530e33fc08a1be53286e93512c259d265bcf22023bf3698f8c8ffaca28

Observation 2841977a-9a7e-455b-b836-2c4e3972167a · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.216995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.750070Z digest=sha256:c661840fb7922e69e6e4699ee757e9c58718f428cc059a11c96e6bc2c2ecd900

Observation 7794c252-cdcd-4f02-9bc9-8412265b1c88 · outbound

This paper cites CS1QA: A Dataset for Assisting Code-based Question Answering in an Introductory Programming Course.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models CS1QA: A Dataset for Assisting Code-based Question Answering in an Introductory Programming Course

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.753114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.753114Z digest=sha256:449380d0b26fc204a00c371e8a7f73612cbc4acbdd72600ff8806877a9fe5962

Observation 69294333-50c2-4ef3-b636-1bf4817087f7 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.204800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.756280Z digest=sha256:04b51f6a11d45d0bb7d37eadf02345c20e36d0d59a1e4747eea9dbd3351ccbe3

Observation b4c27d15-aa5c-44db-8fb6-0b7301f71394 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.192800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.759303Z digest=sha256:c91b73a3c839edbcb42f2728060ccc42595bb9a6e609df1ca7f4df552ee00f8d

Observation 33acb935-17d1-4061-8383-46031e2b2343 · outbound

This paper cites CodeQA: A Question Answering Dataset for Source Code Comprehension.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models CodeQA: A Question Answering Dataset for Source Code Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.762383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.762383Z digest=sha256:669f840082f95d03493d310965c364203180d8f38fee0fedb50fd811ecd666f1

Observation b18bc12a-fdfd-4b29-9623-ea71967496e3 · outbound

This paper cites Large Language Model-Based Agents for Software Engineering: A Survey.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Large Language Model-Based Agents for Software Engineering: A Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.766173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.766173Z digest=sha256:dccee15d76988723af9303f531ebdd81913065c0a5bbe98819702f71c3edd7e8

Observation 31f7ff52-f071-4e90-a145-3178924804fd · outbound

This paper cites MarsCode Agent: AI-native Automated Bug Fixing.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models MarsCode Agent: AI-native Automated Bug Fixing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.769898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.769898Z digest=sha256:933a780ba9cd4de896ee5ec1f3c10e41e73c5ca6a606561c27908f0abd0b0696

Observation b47db412-079e-494d-a814-8a89fa9c97d6 · outbound

This paper cites Alibaba LingmaAgent: Improving Automated Issue Resolution via Comprehensive Repository Exploration.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Alibaba LingmaAgent: Improving Automated Issue Resolution via Comprehensive Repository Exploration

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.773884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.773884Z digest=sha256:a295acbe7ed9c3fad5bf7c0f5dc2f7c03aabe7d9d1b79ac7aff845c031c4f3ca

Observation 66788ad5-5064-4a54-b1d9-5af72ea5cec7 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.778162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.778162Z digest=sha256:9a69dc99d1a0f6497dde6c49bb392d4c352446a3185a8795fb33d75da9abf5f6

Observation 671f294d-ae94-453c-a163-5aeb427b7e48 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.159665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.786237Z digest=sha256:074c359f706be99c082a056c580965de7a8a1c5af2cd7500c5a6af2e4f575db9

Observation aa3ec18d-c639-4023-9d2c-99dc5375fb90 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.790435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.790435Z digest=sha256:3221b23357092c2008ccc990983c13e66ef1b15cbf27124c3a0fe87a9e9147c7

Observation 1fdc7f2c-5d5b-4eb1-9228-2795fcff54a8 · outbound

This paper cites In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models In Proceedings of the 54th ACM Technical Symposium on Computer Science Education V

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T11:39:34.172802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.782273Z digest=sha256:bebc5baff3659ad45f325f75bc72fad85f156463c3b4afdb83f00c84f59bd758

Observation 0133477d-ebf4-49d4-b6a6-4910dcec5908 · outbound

This paper cites StarCoder: may the source be with you!.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models StarCoder: may the source be with you!

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.797995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.797995Z digest=sha256:bf7a5a8dce3dd38799a397e02fa3ada2a7b5bde500ff874c4204a0768b505437

Observation 27c3fe17-275b-4b4e-880c-c764f2c3589b · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.802325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.802325Z digest=sha256:4a588ab7296b621403a9d2fbc9500ab3e270087d2c8c1ddc8425f9351b932816

Observation f09e517f-a269-43da-a9fa-f52a74cc2dbd · outbound

This paper cites AgentFL: Scaling LLM-based Fault Localization to Project-Level Context.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models AgentFL: Scaling LLM-based Fault Localization to Project-Level Context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.793973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.793973Z digest=sha256:1213b0785782c1da220164e2b4083a712ef4142591a00c64f3c453061e116e0a

Observation 446a64cd-193c-4482-a7cc-e40d4b44be73 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.118431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.810551Z digest=sha256:f17b33bfb7685698545688460e143a1761894dd0ec2664d42bfe38be91b8d689

Observation 2f37515f-e2c9-4d98-9dee-2af068903482 · outbound

This paper cites MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models MAGIS: LLM-Based Multi-Agent Framework for GitHub Issue Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.814213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.814213Z digest=sha256:09e711905917f8aad8c8c5a1d79fe3e137b7e3632cb3bed0591c00a2b595f367

Observation 45ce0934-704d-4f33-9a3a-323af91f0e8d · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.130653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.806747Z digest=sha256:ae5c450f4cdf329d47fa1246baa04fd4a89be32560608e68adae01d454096d41

Observation 496d94a6-a0c0-4c82-a9a9-25c5ece1129e · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.821566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.821566Z digest=sha256:b801a7aacdbdd233586fe3fc88b290e31950fa34bcc6d92ed8d8a75826a47adb

Observation 51e64f52-0f11-41cc-b9b2-2539d36e359f · outbound

This paper cites Agentless: Demystifying LLM-based Software Engineering Agents.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Agentless: Demystifying LLM-based Software Engineering Agents

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.825804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.825804Z digest=sha256:89d3eee455bf7020a7bee3be0561fc6152b791601396f30a6769a8cae4fdd849

Observation a524a3c0-1beb-4cf6-ac80-750ab6937de7 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-12T11:39:34.106655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T11:39:33.818246Z digest=sha256:8c2e7b50bd9917a0797afd825e831d0e2233ab3904f0f520519772930f7d1205

Observation b1235a58-3ac0-4cf2-bb7c-6467c8b9fea3 · outbound

This paper cites an unresolved cited work.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.834103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.834103Z digest=sha256:44013053e7a5fc8046f35f6730388600a0e7c241f3aae64942fc2b456588a881

Observation 00923c45-29b0-42ac-85a4-417bf0225f83 · outbound

This paper cites AutoCodeRover: Autonomous Program Improvement.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models AutoCodeRover: Autonomous Program Improvement

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.837756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.837756Z digest=sha256:1513fb9be57396e0b3f8b708492bcf0d2d66ee939611d9a9a776d72bfaff0dd0

Observation 1be7cee2-c66c-424e-95eb-752600efba8c · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.829678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.829678Z digest=sha256:d8242a05e757f3f16a2cb284ac401be8b5d09e564c8276b808825af92187a94e

Observation 5c952517-67fa-4ecd-a0d6-1a4c14c51d23 · outbound

This paper cites Large Language Models for Software Engineering: A Systematic Literature Review.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models Large Language Models for Software Engineering: A Systematic Literature Review

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.734856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.734856Z digest=sha256:b40be0919fce81407bd1876dcdfc7061fb7f34c76ba21b1108d3ddfdbb1a7a44

Observation 24de3a1a-f3f7-46ff-8bf9-9a7ea1c3be3b · outbound

This paper cites In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering.

A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T11:39:33.715985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:39:33.715985Z digest=sha256:16641f792cd1546306fe4f8add6d0c846df8acdbac345b5784c9c73dfb981796

Pith citing papers

Observation 399dd694-103d-440f-80a1-257581758b6a · inbound

Evaluating LLM Agents on Automated Software Analysis Tasks cites this paper.

Evaluating LLM Agents on Automated Software Analysis Tasks A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:26:01.479107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:38:12.144252Z digest=sha256:17559002f0e9936d5cab0f27aad3d9fbc4ff16454007cdd06bc499fb1e562907

Observation f9bd4156-3400-4ce0-b5f5-ffdad8f9f162 · inbound

Evaluating LLM Agents on Automated Software Analysis Tasks cites this paper.

Evaluating LLM Agents on Automated Software Analysis Tasks A Real-World Benchmark for Evaluating Fine-Grained Issue Solving Capabilities of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T16:27:43.489582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:27:43.489582Z digest=sha256:28a54947b108ff66b19333b92f6b67c0e3f376d1e591b5b486e1dbbe11447f38