Pith. sign in

Paper Citation Record · LEDGER

OSS-Bench: Benchmark Generator for Coding LLMs

As of 22 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.12331.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12331 v2

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:41:09.087338Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 421c77fd-dd77-4d41-8c94-8255de3985fa · outbound

This paper cites Github copilot.https://copilot.github.com, 2021.

OSS-Bench: Benchmark Generator for Coding LLMs Github copilot.https://copilot.github.com, 2021

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.765896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.882885Z digest=sha256:188dd9ab0466e4ccd1fec4c9ea3af2d35f699ac039642d3a915d17dd10a17c14

Observation 9420a2ab-6a59-4c0a-ad9c-199c6e3e0c0d · outbound

This paper cites Cursor: The ai-powered code editor.https://cursor.so, 2023.

OSS-Bench: Benchmark Generator for Coding LLMs Cursor: The ai-powered code editor.https://cursor.so, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.755646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.887126Z digest=sha256:ae05610b9c22d0f9f504a1318aaab2ae1c8cdd059d35e88d696836afc047dbb8

Observation 37a4dd52-72cb-4904-9535-2acb97562cc9 · outbound

This paper cites an unresolved cited work.

OSS-Bench: Benchmark Generator for Coding LLMs Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.890792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.890792Z digest=sha256:e652a2b32807a436f872f18bb4949292ba1a788060c1dd4f7ef7f199dee0b713

Observation ffd97942-0d28-4e3f-8692-fa5f96fc972c · outbound

This paper cites Codelmsec benchmark: Systematically evaluating and finding security vulnerabilities in black-box code language models.

OSS-Bench: Benchmark Generator for Coding LLMs Codelmsec benchmark: Systematically evaluating and finding security vulnerabilities in black-box code language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.738087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.894506Z digest=sha256:58f005a92621dfd592a848de4420d53eb404081134ec2fc96f9859debd5fa6bc

Observation f4362cce-2d41-45cc-9720-d8e55ab24d87 · outbound

This paper cites DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation.

OSS-Bench: Benchmark Generator for Coding LLMs DynaCode: A Dynamic Complexity-Aware Code Benchmark for Evaluating Large Language Models in Code Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.898224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.898224Z digest=sha256:69e0fb040678621d304f1410d2ed07a9aeae06a191c296be879741060175436e

Observation 92931cb0-3c53-4536-bab2-27d63c5c177c · outbound

This paper cites Humanevo: An evolution-aware benchmark for more realistic evaluation of repository-level code generation.

OSS-Bench: Benchmark Generator for Coding LLMs Humanevo: An evolution-aware benchmark for more realistic evaluation of repository-level code generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.728038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.902726Z digest=sha256:95696d29a562cec4d843ce0d9ee7832a205dae789a7540360d30099543e2f944

Observation be0cf7b6-15bc-4966-9283-e7b95304764b · outbound

This paper cites Codeif-bench: Evaluating instruction-following capabilities of large language models in interactive code generation.arXiv preprint arXiv:2503.22688, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Codeif-bench: Evaluating instruction-following capabilities of large language models in interactive code generation.arXiv preprint arXiv:2503.22688, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.906493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.906493Z digest=sha256:87f5df045d255e11a39dc004e74b52567294e3c3c147e0f0cf12cf50a62bd63b

Observation 19ebd076-1d70-427d-912f-77193fece6a3 · outbound

This paper cites ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code.

OSS-Bench: Benchmark Generator for Coding LLMs ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.909862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.909862Z digest=sha256:723ff996bfb6f40e90cd5c5fab44f0d584fd56a4bf687d4f3bd947927876adaf

Observation 0fcf29d9-48f7-44a6-bf52-a8bd0d4fd83b · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

OSS-Bench: Benchmark Generator for Coding LLMs SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.913882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.913882Z digest=sha256:d3d8895add91afac0f342afb5bd2e54bfa1e42ecf4f1fe6645afae0c4a35d58f

Observation be14a1d2-d442-401c-a987-89953bee9903 · outbound

This paper cites Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica.

OSS-Bench: Benchmark Generator for Coding LLMs Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.718194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.917414Z digest=sha256:f95d781db2631e73b80f4761f2570975f8fd75415a84b0fc571d7aae11976be6

Observation 1f926510-652a-45e2-81e3-6297ad342f41 · outbound

This paper cites an unresolved cited work.

OSS-Bench: Benchmark Generator for Coding LLMs Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:41:09.708228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.920715Z digest=sha256:d18299ef51f2ea0c6ded94a046b97d04134241421cd47266c131c907d987c6cc

Observation 5aabb163-a1e4-4c76-8171-de64910dd20d · outbound

This paper cites Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving.

OSS-Bench: Benchmark Generator for Coding LLMs Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.924377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.924377Z digest=sha256:b38433ce69a1692bdb3594197895f1cfa629285ebcc4c4bb167954d5e767f81d

Observation 4f899914-f0ea-4ffe-be68-ef5008893aec · outbound

This paper cites SecRepoBench: Benchmarking LLMs for secure code generation in real-world repositories, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs SecRepoBench: Benchmarking LLMs for secure code generation in real-world repositories, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.698118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.928133Z digest=sha256:1b7e72c91ccb2661f94fc6bd0399bf5504a25009d5657a8d0b51f9a09cc9a36c

Observation 175ab0dc-7966-411f-b6c0-7ddeb40f37a1 · outbound

This paper cites CWEval: Outcome-driven evaluation on functionality and security of LLM code generation, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs CWEval: Outcome-driven evaluation on functionality and security of LLM code generation, 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.687983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.931450Z digest=sha256:25e71b21d947e0978fdce570a3d8d0dccaf61a14b50f83cf9b19bd4f6248c62d

Observation d267b0db-c1a3-4ef8-b916-cecb179222a6 · outbound

This paper cites Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models.

OSS-Bench: Benchmark Generator for Coding LLMs Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.934788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.934788Z digest=sha256:b392067a813c9902da4cf3da55336a0be28a7e07fcfcdf549c9bbd31511dd8fa

Observation 8097f06e-e716-4baa-a9f5-f247dd6ddd16 · outbound

This paper cites Codearena: Inspecting and improving code quality metrics using minecraft.

OSS-Bench: Benchmark Generator for Coding LLMs Codearena: Inspecting and improving code quality metrics using minecraft

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.677609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.938516Z digest=sha256:920488cbd163fc29cb1a01730c06c9496373dd6c01cc58c371d41d41581b18c2

Observation 838d148b-e976-4a6a-8eae-01547b8c5a40 · outbound

This paper cites CodeElo: Benchmarking competition-level code generation of LLMs with human-comparable elo ratings, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs CodeElo: Benchmarking competition-level code generation of LLMs with human-comparable elo ratings, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.667110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.942026Z digest=sha256:c5bcdcb2d0d0a2e7f92c26952c9dcfd9da53b5b2f6ee8e4f0a1aefcfd3bed88e

Observation f7e97651-5964-4935-a92f-b6670019008f · outbound

This paper cites ComplexCodeEval: A benchmark for evaluating large code models on more complex code.

OSS-Bench: Benchmark Generator for Coding LLMs ComplexCodeEval: A benchmark for evaluating large code models on more complex code

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.656575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.945410Z digest=sha256:f63368a31c98b1479efd7d6866b73ea0106c22f618822c84ed8045bab159e2e9

Observation d54edff5-10a6-4393-a1e0-24f00abd278f · outbound

This paper cites PythonSaga: Redefining the benchmark for code generating LLMs, 2024.

OSS-Bench: Benchmark Generator for Coding LLMs PythonSaga: Redefining the benchmark for code generating LLMs, 2024

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.646182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.948865Z digest=sha256:2feb95fa8d7f981b688d5fa2311ca4d8df6adacd9bd59292d7d4cf9e8aa5c0d8

Observation b56dff38-2269-4311-9d09-2d1988dba53b · outbound

This paper cites Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark.

OSS-Bench: Benchmark Generator for Coding LLMs Top General Performance = Top Domain Performance? DomainCodeBench: A Multi-domain Code Generation Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.952533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.952533Z digest=sha256:a98bffc7258af1f87f2e04102ddf60eb167e228b014e39c3a6a44f4ac46ca15c

Observation 5d8a4b47-36f1-4ff6-8f67-cc13055f3cd5 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

OSS-Bench: Benchmark Generator for Coding LLMs Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.635334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.956417Z digest=sha256:fb8bfb00f77f2d70f49f1e835fb0975bbbaee01ca3cbb75debbba5ddf65f8b2e

Observation f99279ee-a255-4536-9367-6778d381a976 · outbound

This paper cites ClassEval: A manually-crafted benchmark for evaluating LLMs on class-level code generation, 2023.

OSS-Bench: Benchmark Generator for Coding LLMs ClassEval: A manually-crafted benchmark for evaluating LLMs on class-level code generation, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.625092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.959822Z digest=sha256:b0ad53ffae94e34e826361440a65cb1164d9b896268f638f6d9ce9cf43b63cde

Observation 2cbeb47d-f80b-4c4f-86fb-534d76cc5f87 · outbound

This paper cites HumanEval-XL: A multilingual code generation benchmark for cross-lingual natural language generalization, 2024.

OSS-Bench: Benchmark Generator for Coding LLMs HumanEval-XL: A multilingual code generation benchmark for cross-lingual natural language generalization, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.613243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.963333Z digest=sha256:8e4e84c967dcb2f090cb02fef99323d9e3d7a795d915a3e945744b593dadc5ce

Observation 996c0b14-7950-4375-9405-0055fdad5d9e · outbound

This paper cites Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x.

OSS-Bench: Benchmark Generator for Coding LLMs Codegeex: A pre-trained model for code generation with multilingual benchmarking on humaneval-x

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.966684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.966684Z digest=sha256:46db9646d5d2be499c15877f6c49a5392b159e7d39afc8fde40c6e1b6b348374

Observation 25f43123-634e-4ac1-998e-991732e8da59 · outbound

This paper cites {AddressSanitizer}: A fast address sanity checker.

OSS-Bench: Benchmark Generator for Coding LLMs {AddressSanitizer}: A fast address sanity checker

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.970147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.970147Z digest=sha256:4b93a34c62c44514dae54a236e3bed89398b2b6f84cd33a36cad346c49caafe3

Observation 74e819d9-d3f8-47ba-9843-c946e1a8f32b · outbound

This paper cites php-src: The php interpreter.https://github.com/php/php-src, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs php-src: The php interpreter.https://github.com/php/php-src, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.596794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.973421Z digest=sha256:acdfc6b3fbeb3afdeb35ff8c576339eb3a7dfdadb32e97a02aca41065e80b958

Observation b09990d6-ebf5-434e-9d9b-0c5d93593900 · outbound

This paper cites Richard Hipp.

OSS-Bench: Benchmark Generator for Coding LLMs Richard Hipp

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.586385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.976672Z digest=sha256:43a77e8f9cfa4afcc5ec314a796bbb52ef30d4a645aa82cb14553c25c9adf7ca

Observation 6e377465-2058-4c79-ac59-b801173648b5 · outbound

This paper cites https://testing.googleblog.com/2020/08/ code-coverage-best-practices.html, 2020.

OSS-Bench: Benchmark Generator for Coding LLMs https://testing.googleblog.com/2020/08/ code-coverage-best-practices.html, 2020

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.574840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.979969Z digest=sha256:0763e1d62db1bbb667d49731760f57123b1211f512b4026fd9bb6a69d8dbfccc

Observation 88dc70a5-32bc-440f-abeb-a04a7fac21bc · outbound

This paper cites libclang: C interface to the clang library.

OSS-Bench: Benchmark Generator for Coding LLMs libclang: C interface to the clang library

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.564370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.983338Z digest=sha256:6d7e48dc0f4ec2d3410185c596b1b991e86065da196838e0d0ea9c50fc010c1c

Observation d1450324-a4f6-45df-a943-c1b8d6b4b472 · outbound

This paper cites an unresolved cited work.

OSS-Bench: Benchmark Generator for Coding LLMs Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.987369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.987369Z digest=sha256:d572b9208a601ad75d46ac4cdd79ceb8222e46f812b50866c16c1a0b5966f72e

Observation 1978e6ba-1cb1-4ecc-bc52-25c65abf6ff3 · outbound

This paper cites difflib — helpers for computing deltas between objects.

OSS-Bench: Benchmark Generator for Coding LLMs difflib — helpers for computing deltas between objects

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.548044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.990624Z digest=sha256:d34e7ba746e96e20ebc597bf690ceca65ed7e90f7431ede2dd17526aeaa504fa

Observation cd2d49f0-5dc8-4fae-98e3-42c38be67156 · outbound

This paper cites gpt-o1 model.https://platform.openai.com/docs/models/o1, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs gpt-o1 model.https://platform.openai.com/docs/models/o1, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.538208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.993964Z digest=sha256:66ff2cf719298271285e545269a163caef5284e675bd8e1695d5446a5c9e3b7e

Observation ab3f1842-84bd-444f-b18d-d960ecb00322 · outbound

This paper cites o3-mini model.https://platform.openai.com/docs/models/o3-mini, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs o3-mini model.https://platform.openai.com/docs/models/o3-mini, 2025

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.528623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:08.997232Z digest=sha256:70bf5e7d68fb1c14c15e22ac20b043aabf6c1d9f299fec4b30d5a47f722c5e85

Observation c5d584d7-b454-482d-bd96-1006e253fffb · outbound

This paper cites Claude 3.7 sonnet.https://www.anthropic.com/claude/sonnet, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Claude 3.7 sonnet.https://www.anthropic.com/claude/sonnet, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.518845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.000387Z digest=sha256:6d55d2a0a73fca10d0b85ab403f1d205b19bc84075ab93c4491e1280d67e3687

Observation 5ac39b2b-67e2-42b1-a827-d6712bfb99da · outbound

This paper cites Claude 3.5 haiku.https://www.anthropic.com/claude/haiku, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Claude 3.5 haiku.https://www.anthropic.com/claude/haiku, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.509216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.003876Z digest=sha256:37ac304e22fce2786b9ccab9814e0b19f2ff6a73a844beafe9786ab0d3260577

Observation ecb39fce-e3dc-4376-b46c-2dc6b95967f7 · outbound

This paper cites Gemini 2.5 flash.https://developers.generativelanguage.google/, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Gemini 2.5 flash.https://developers.generativelanguage.google/, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.499364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.007301Z digest=sha256:4efa641e18c22da95bf6346da8675db9c100bea3c50ed7126feffd7e34f3c6c8

Observation 614a4c67-f17b-462e-afc7-18d9722d4bb9 · outbound

This paper cites Llama 3.3 70b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Llama 3.3 70b instruct (fp16)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.489480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.010502Z digest=sha256:8c7a9c62dabba4005f1189201d904ada7695619ee9bc7d85e2401b71f72e2c46

Observation 3e19ca2f-c7eb-4839-9edb-39d90a5268ed · outbound

This paper cites Codellama 70b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Codellama 70b instruct (fp16)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.479853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.013703Z digest=sha256:dd36efc0ed9040568553ee7129e06ccee0e928401dd00af960ca199c0e74857c

Observation d97379d9-4aec-40a6-af08-2b98a80f12bc · outbound

This paper cites Qwen 2.5 coder 32b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Qwen 2.5 coder 32b instruct (fp16)

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.470476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.016914Z digest=sha256:507cd3954f2e29f31768c179f99fd57a71ef8ad5f032d1519b7b063422160bdd

Observation 16c0acef-3449-42ba-8439-7306efa9a235 · outbound

This paper cites Qwen 3.0 30b-a3b fp16.https://ollama.com/library/qwen3:30b-a3b-fp16, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Qwen 3.0 30b-a3b fp16.https://ollama.com/library/qwen3:30b-a3b-fp16, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.460304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.020235Z digest=sha256:4a79b96f2d9d3f478c7629589c85791088c183ce1fd5270581ddf278c138e2ae

Observation 36fdb130-8cbe-47e1-ad88-2d20377819a9 · outbound

This paper cites Qwen 3 8b fp16.https://ollama.com/library/qwen3:8b-fp16, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Qwen 3 8b fp16.https://ollama.com/library/qwen3:8b-fp16, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.450714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.023486Z digest=sha256:428497e4f86d30d291a1b8bbd65b6f813c2b7328d68b487ce5e89462584cff01

Observation ed415a7d-841a-4574-9f59-df739406e666 · outbound

This paper cites Gemma 3 27b-it fp16.https://ollama.com/library/gemma3:27b-it-fp16, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Gemma 3 27b-it fp16.https://ollama.com/library/gemma3:27b-it-fp16, 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.440770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.026695Z digest=sha256:0a056a68e4dd9ee42caf92405ded6560337b4386b4e6b108448e9505d0378801

Observation f14f87fd-c5ed-4fa4-9171-ec76af6b46ed · outbound

This paper cites Qwen 2.5 coder 14b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Qwen 2.5 coder 14b instruct (fp16)

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.430956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.029928Z digest=sha256:9f6f1a4bf07c1597081d9d34ce86214e102c09a21b4a7e43e2560e41cdba3aa3

Observation 2e09b457-c428-4822-8ef3-01b67b7c49c2 · outbound

This paper cites Deepseek coder v2 16b lite instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Deepseek coder v2 16b lite instruct (fp16)

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.421381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.033335Z digest=sha256:f66b88a642dc29f0a625a7d9ef5ac5afbd6212d8570646b593c65cd2e672e231

Observation 2651a8ef-fb93-47ae-a2dc-546876a5fc10 · outbound

This paper cites Starcoder2-15b-instruct-v0.1 (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Starcoder2-15b-instruct-v0.1 (fp16)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.411806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.036683Z digest=sha256:e8726357f33770092ffadbe1f50f567e1bb63bcd2a65b217c47048ee38d26c5d

Observation 065a418e-8f61-4bad-9e28-8c86c4b37fce · outbound

This paper cites Phi-4 14b fp16.https://ollama.com/library/phi4:14b-fp16, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Phi-4 14b fp16.https://ollama.com/library/phi4:14b-fp16, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.402729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.039967Z digest=sha256:a69ee783965558ee42d596fad5f4e56d8b81eb901b4618506c7962afc39d4463

Observation e8a8fc19-035d-41bc-9fb8-27aa8ccfdf02 · outbound

This paper cites Mistral 7b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Mistral 7b instruct (fp16)

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.392942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.043072Z digest=sha256:456d3e104ae8dd155d7765e53f8b2666821f40b70c95a8400c4cd0b77dd87e14

Observation 56c62f9f-203a-46fc-bb32-8424f361b066 · outbound

This paper cites Codegemma 7b instruct (fp16).

OSS-Bench: Benchmark Generator for Coding LLMs Codegemma 7b instruct (fp16)

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.382442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.045955Z digest=sha256:c33177fb1a81c0d2e8bd2dce4758c16af9cd98f84ac753d9afcc1d675196faf6

Observation 51c00e5f-a141-4348-b0c7-1541b8d6090b · outbound

This paper cites Openai: Advances in safe and beneficial ai.https://openai.com, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Openai: Advances in safe and beneficial ai.https://openai.com, 2025

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.372657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.048955Z digest=sha256:1ed5c3d701664fac47d9f15caeb78dc6fe7dc746109c4d4ff409878990e39cb4

Observation 1074e2ff-727f-496b-800a-30dcd2e6b1b3 · outbound

This paper cites Anthropic: Building reliable, steerable ai systems.

OSS-Bench: Benchmark Generator for Coding LLMs Anthropic: Building reliable, steerable ai systems

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.362447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.052241Z digest=sha256:994f76d5ba0ebd1a12c800707a7ea3f0c0744213e0c898b2b210aba70e1e7a66

Observation faa5be50-d86b-462e-97bc-7f4e4e873a70 · outbound

This paper cites Google: Organizing the world’s information.https://www.google.com, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Google: Organizing the world’s information.https://www.google.com, 2025

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.352265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.055397Z digest=sha256:3ad050946c00a5734d9495929c7a4ff9520af5d78a806e2d242995bed0c7d19b

Observation b029a42d-0535-4919-b5d9-939351530fc5 · outbound

This paper cites Deepseek: Developer of high-performance open-source llms.

OSS-Bench: Benchmark Generator for Coding LLMs Deepseek: Developer of high-performance open-source llms

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.342614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.058832Z digest=sha256:6805b34cc3011f321f7de711356aa9654dbc516ab8d053b2faae26e4036ad528

Observation c12238fa-e33f-4420-ac49-bcf5605bd33e · outbound

This paper cites Alibaba group: Global trade and technology.https://www.alibaba.com, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Alibaba group: Global trade and technology.https://www.alibaba.com, 2025

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.332122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.061970Z digest=sha256:3b1d096b8d75e0d9b6305c04cc29de64793ba4c9fff0b90adf007934983abb06

Observation c9f9aa06-2519-4215-acfd-e96bd8cb4b54 · outbound

This paper cites Meta: Bringing the metaverse and social technology together.

OSS-Bench: Benchmark Generator for Coding LLMs Meta: Bringing the metaverse and social technology together

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.321033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.065031Z digest=sha256:77928c511a7ca56f2e012a82ee7fd42d467ceb029b6dcab83b2181b17a63aea9

Observation cded8149-9ca2-4f3e-80e8-5a24d16e8f85 · outbound

This paper cites Microsoft: Empowering every person and organization.

OSS-Bench: Benchmark Generator for Coding LLMs Microsoft: Empowering every person and organization

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.310513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.068195Z digest=sha256:ef8b9a6ff77981e2520e79d22ecb1a0898cb93e65eefc93b797ce39423feb1c4

Observation 1ab6e7d1-79b9-4dc6-904f-a9092340bc36 · outbound

This paper cites Bigcode: Open and responsible development of code llms.

OSS-Bench: Benchmark Generator for Coding LLMs Bigcode: Open and responsible development of code llms

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.299952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.071519Z digest=sha256:a9383d14d2efa3734f52c6bcd95864cebd4e5e0bf1ea9f9922e90ec20a276332

Observation 2c5f0684-a7fa-4f47-ab91-988aacd0d958 · outbound

This paper cites Mistral ai: Frontier ai in your hands.https://mistral.ai, 2025.

OSS-Bench: Benchmark Generator for Coding LLMs Mistral ai: Frontier ai in your hands.https://mistral.ai, 2025

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.289957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.074554Z digest=sha256:4444dc665c67718e8e93676307c9a2bbcb83edf8e218a71dc4c70b3ac68186d4

Observation 77e66e68-6bcb-4ffd-a89b-8398e909a334 · outbound

This paper cites Ollama: Get up and running with large language models locally.

OSS-Bench: Benchmark Generator for Coding LLMs Ollama: Get up and running with large language models locally

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:41:09.279340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.077778Z digest=sha256:f866afe6ee855909953a229cb8866d7839fbc8fd84667ae15c17045523d87f3f

Observation 6376ec4b-63b3-4a79-bb6d-4e07616eb100 · outbound

This paper cites SPoC: Search-based Pseudocode to Code.

OSS-Bench: Benchmark Generator for Coding LLMs SPoC: Search-based Pseudocode to Code

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:09.080649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:09.080649Z digest=sha256:62c7814588bed830141481b7f5142f38252dadee7f7e57113ea62d7be323c9b4

Observation c79ca41e-c891-4d6a-b74f-17d8593de675 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

OSS-Bench: Benchmark Generator for Coding LLMs Evaluating Large Language Models Trained on Code

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:09.084092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:09.084092Z digest=sha256:97015d11dc395c6e326ec9b6d64c0c7b17a15bd5b9371c8a2623aaacdf923d4e

Observation ca45235d-09f5-492e-9102-3bbbf33dedfd · outbound

This paper cites Fuzzing the PHP Interpreter via Dataflow Fusion.

OSS-Bench: Benchmark Generator for Coding LLMs Fuzzing the PHP Interpreter via Dataflow Fusion

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T20:41:09.129077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T20:41:09.087338Z digest=sha256:0299a4ed7379c9dd80e5ccd8ada4fe8d9a6eb53e1370f1866ec7fc44ebd26160

Pith citing papers

No inbound Pith citation observations are available.