Pith. sign in

Paper Citation Record · LEDGER

Rethinking Code Performance Benchmarks for LLMs

As of 18 August 2026, this Paper Citation Record lists 100 of 161 outbound references and 1 inbound Pith citation observation for arXiv:2607.07619.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.07619 v1

Coverage vector

measured 100 of 161 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-09T05:16:58.549058Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T10:53:08.900571Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 161 outbound references displayed

  • verified exact23
  • verified fuzzy56
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch11

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3dc05d99-6545-4cb5-8cb8-6cb59d2262ad · outbound

This paper cites Breakthroughs in statistics: Methodology and distribution , pages=.

Rethinking Code Performance Benchmarks for LLMs Breakthroughs in statistics: Methodology and distribution , pages=

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.148720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:41cf17a974473fa4590408a635496c8f2702872b520e49643023eeb7a81c2a87

Observation acfb784e-67c3-47a4-8f9a-c293c511fe68 · outbound

This paper cites author Ralph, P.

Rethinking Code Performance Benchmarks for LLMs author Ralph, P

Reference 2

Resolution
verified exact
doi, observed 2026-07-09T05:26:01.109354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:d8f0a5101cf7b1bab2c3193419c03138028427e2eb192ae8127e3447822c6d06

Observation d34e65bf-9ed2-471e-9233-913e05179d33 · outbound

This paper cites Educational and psychological measurement , volume=.

Rethinking Code Performance Benchmarks for LLMs Educational and psychological measurement , volume=

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.216623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:41bf82027bc03d51539b4df95a8b51988f19260ce46492393ef3a3274a5fce84

Observation fe24e569-5235-495e-8778-02cb696b51be · outbound

This paper cites Biometrika , pages=.

Rethinking Code Performance Benchmarks for LLMs Biometrika , pages=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.163104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:b0535e6ed6697e295449b955693b28dd8b20c39554d0e0acda7b899f2f2df736

Observation fd78b4db-1995-47ae-9a75-a254d04906e1 · outbound

This paper cites Breakthroughs in statistics: Methodology and distribution , pages=.

Rethinking Code Performance Benchmarks for LLMs Breakthroughs in statistics: Methodology and distribution , pages=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.223370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:f36ae0a324a394c780ce99448ebd802ed748037ed40fe1e3b175435722cde9e0

Observation dca09601-3147-446f-80d2-1c4acc0cdea6 · outbound

This paper cites an unresolved cited work.

Rethinking Code Performance Benchmarks for LLMs Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-07-09T05:46:02.122374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:77b9f00354fb108aadffbeca6e937d86602b13e82479fe3170691d147e4ab931

Observation f9d0b908-ac58-463c-8952-ae9b9214289b · outbound

This paper cites IOWA , author=.

Rethinking Code Performance Benchmarks for LLMs IOWA , author=

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.267905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:e73d996650fa0569a9c8189d4cf6cebf34e94bf7f8417a52c4b05c630070c006

Observation bec1610c-c54d-4c47-8e4f-25402746903e · outbound

This paper cites Biometrika , volume=.

Rethinking Code Performance Benchmarks for LLMs Biometrika , volume=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.647556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:2054f61d661bd76c3fce264948d61193ee0c006eee276d8a7666cb23a8684f2d

Observation 8b607199-d697-40fe-b552-57c8dc90bda7 · outbound

This paper cites 2009 , publisher=.

Rethinking Code Performance Benchmarks for LLMs 2009 , publisher=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.649714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:b4c16b8489e04c52f0e84b8931b43f7ccf087e5bc3ab33bb95e3ca3f7e9de781

Observation 1e8f7f04-8804-41ce-9993-36935cbda212 · outbound

This paper cites Information and Software Technology , volume=.

Rethinking Code Performance Benchmarks for LLMs Information and Software Technology , volume=

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.636423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:9b1284d8dda065fde49d79f18737e44044ec4d63f2e670c68bfd196b77102b28

Observation 002f5173-df22-4d6c-93a8-ba36247a9163 · outbound

This paper cites , author=.

Rethinking Code Performance Benchmarks for LLMs , author=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.654341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:f5a8ba7865693d7cd48d23a44ff422889673bcc7a8d9e223e33b29c69fa498bb

Observation ce30ab09-cb96-419f-8948-5d2a941fd5aa · outbound

This paper cites Biometrics bulletin , volume=.

Rethinking Code Performance Benchmarks for LLMs Biometrics bulletin , volume=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:46:02.088696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:dcd6188d80721deae284fcbbe55b27adb6a09d75f60add3716bd949ddf5ad4d0

Observation c70f330e-d95a-4f5f-8be0-065ba12e66f3 · outbound

This paper cites The Journal of Nervous and Mental Disease , volume=.

Rethinking Code Performance Benchmarks for LLMs The Journal of Nervous and Mental Disease , volume=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.611608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:488a3e91aea192af39276b7a7454b4674a59277416ae11d2ece97ee956722403

Observation da3d9738-4d0e-4556-a515-442a6ea6a0e6 · outbound

This paper cites Devanbu, Christoph Treude, and Michael Pradel.

Rethinking Code Performance Benchmarks for LLMs Devanbu, Christoph Treude, and Michael Pradel

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T05:26:01.136701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:6b0f8c8559f8cb6f0bffa1e1fd40b26d01c6785eb55c19dd8bf864c2cbad40e3

Observation da2efc50-d426-43e3-b0ad-b580e25bcdc7 · outbound

This paper cites From generation to judg- ment: Opportunities and challenges of llm-as-a-judge.

Rethinking Code Performance Benchmarks for LLMs From generation to judg- ment: Opportunities and challenges of llm-as-a-judge

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:26:01.089108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:e6d14a6daccf292b6420a71006a9fe841f31799a1416818422177642a2eb3825

Observation d63cd483-321c-4a1e-8dc5-6f159ee94db9 · outbound

This paper cites Proceedings of the 18th.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the 18th

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T05:26:01.174380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:fb342059f111caa9a96657e1449039f96d3a12d3d215bed70593faabe20b82c1

Observation 75bd5282-35ce-4f69-845b-02fd26065bbb · outbound

This paper cites DeepSeek Chat , howpublished =.

Rethinking Code Performance Benchmarks for LLMs DeepSeek Chat , howpublished =

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.151418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:11708111608e5682d033123ef91e8e26be17153f1fa022443b9d65f43da41029

Observation 09eaa7ce-a107-4eb9-a789-4b29ab00c273 · outbound

This paper cites DeepSeek Model , howpublished =.

Rethinking Code Performance Benchmarks for LLMs DeepSeek Model , howpublished =

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.239298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:b850888f6bfa7aac4e19817420f4e77b1061bc761c7fe50121cb279c200c0d4f

Observation 3593a4a1-49ef-4a1c-9e79-9e4a92269d0b · outbound

This paper cites ChatGPT , howpublished =.

Rethinking Code Performance Benchmarks for LLMs ChatGPT , howpublished =

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.213591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:99017a91696077467d8dce09e4ba9f5cb4beded36dceace478e29ceb167a8773

Observation bf392ddf-d045-42a4-b749-b5adfa973bad · outbound

This paper cites Gemini25 , howpublished =.

Rethinking Code Performance Benchmarks for LLMs Gemini25 , howpublished =

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.234230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:dbc3aa161788e8dacda430034b3c3b64be7ff3e19366eb45a37535e06005a56b

Observation 67b7d851-4c75-4fd4-a8bb-392987f63ccd · outbound

This paper cites claudesonnet45 , howpublished =.

Rethinking Code Performance Benchmarks for LLMs claudesonnet45 , howpublished =

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.172238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:c72ac60512acadb5d2acf2d38171b54ab34ac685fc8462f9bdbd55693962bc51

Observation cd2ee14b-3790-4f90-8531-b80ed1616774 · outbound

This paper cites , author =.

Rethinking Code Performance Benchmarks for LLMs , author =

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.165779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:8778f64d4887773359f643c5b52ac7e48e1492c6b0066505557d6500c0a7a3f0

Observation fcc426f8-6dde-42e8-a42c-5f4e608b4cb8 · outbound

This paper cites Journal of Machine learning research , volume=.

Rethinking Code Performance Benchmarks for LLMs Journal of Machine learning research , volume=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.184980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:d118bd2176e3998605c40c688e7de6193d5484a3181fc33d3d6c346be88ee740

Observation 1232e5c0-4553-4a5c-94c3-179905eb8afc · outbound

This paper cites The annals of mathematical statistics , pages=.

Rethinking Code Performance Benchmarks for LLMs The annals of mathematical statistics , pages=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.160417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:0dbed0a4ff1e1e3905ae5b3501cf1735af453905b0590487a503d3d47414f542

Observation 131881fe-6e7f-40a8-b6fe-ed8246c091b4 · outbound

This paper cites an unresolved cited work.

Rethinking Code Performance Benchmarks for LLMs Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-07-09T05:46:02.105787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:3c09cd3483b1491abf1dd075d242c731924fa74107343d4e3d8a81d418e97420

Observation 9ecce98e-2188-4ac6-bf3c-97461a7e38b4 · outbound

This paper cites an unresolved cited work.

Rethinking Code Performance Benchmarks for LLMs Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-07-09T05:36:02.590323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:9bb30d1ecef1b05ed8b4500793bbe8d636cea6fd2abd678f8cfab034fa9f85b3

Observation 0c2f5bdd-b131-4300-b97b-c493b13a97e9 · outbound

This paper cites TestEval : Benchmarking Large Language Models for Test Case Generation.

Rethinking Code Performance Benchmarks for LLMs TestEval : Benchmarking Large Language Models for Test Case Generation

Reference 27

Resolution
metadata mismatch
doi, observed 2026-07-09T05:26:01.182655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:baefee6680929003961f77caf72ea3e57c984ce6432b3d9e0d7cd22ec8ef7d0d

Observation f3a21cfc-654f-4d83-9484-3509ef71d1d8 · outbound

This paper cites Robust benchmarking in noisy environments.

Rethinking Code Performance Benchmarks for LLMs Robust benchmarking in noisy environments

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:26:02.127013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:a5ced9e0d2e5a7d42fc3759f09628914542ac7a38e77c4ca397c48ba88a6706e

Observation 175216fa-9799-409d-8d00-9b4c47baa496 · outbound

This paper cites Proceedings of the 19th Annual International Conference on Supercomputing,.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the 19th Annual International Conference on Supercomputing,

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:26:01.118433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:fead4ff0728816e2acb8a67bb2c512b4486601d5906fdc911ac4a6aea46a9234

Observation a3e96f8d-0b26-4b8f-8f9e-fc2af429bed4 · outbound

This paper cites Mogul and Anita Borg , title =.

Rethinking Code Performance Benchmarks for LLMs Mogul and Anita Borg , title =

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:26:01.179288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:b95714e6db825497494e8b69ddf1903755a446d7cd0dc974b797c9ecadde91c8

Observation 2c1fe6bc-282b-45b7-9167-17afa9b8e7b7 · outbound

This paper cites Tyson and Matthew K.

Rethinking Code Performance Benchmarks for LLMs Tyson and Matthew K

Reference 31

Resolution
verified exact
doi, observed 2026-07-09T05:26:01.101433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:9fa51ccd61068124c9f826d4f0fb952218a4d062fd1bf462b54f415f40819294

Observation d410037b-3811-4b88-a02a-e2f338c29737 · outbound

This paper cites Tyson and others , title =.

Rethinking Code Performance Benchmarks for LLMs Tyson and others , title =

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.605587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:99bb6aebac36cee51a7e9b004e105ec2eb1d84294223025507c6ddd4e7a1fe2b

Observation 9a758395-9f83-4893-a6fa-41b9df04ceff · outbound

This paper cites 1999 , url =.

Rethinking Code Performance Benchmarks for LLMs 1999 , url =

Reference 33

Resolution
verified exact
doi, observed 2026-07-09T05:26:01.081703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:bd9de5ea1fffca35369e64271ffacf1f050346c0f42f430c8408c9a5afe1ed77

Observation c07b46fd-a812-47a1-a574-a2cc46901fe1 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

Rethinking Code Performance Benchmarks for LLMs The Twelfth International Conference on Learning Representations,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.575326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:d5c649878d98fc348ad8e27417164a4d40a56985ce79533e91c0b20633711bbb

Observation 3eddb141-00ad-4841-b88b-3a2aed730b3e · outbound

This paper cites an unresolved cited work.

Rethinking Code Performance Benchmarks for LLMs Unresolved cited work

Reference 35

Resolution
verified exact
doi, observed 2026-07-09T05:26:01.018502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:45aa723256a9ab97b9a4c7ac01d350c9c67d81270c04fab239b85040a13118a3

Observation 6f251ed2-c4c3-4a5f-ace8-0d11013bde83 · outbound

This paper cites NoFunEval: Funny How Code LMs Falter on Requirements Beyond Functional Correctness.

Rethinking Code Performance Benchmarks for LLMs NoFunEval: Funny How Code LMs Falter on Requirements Beyond Functional Correctness

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:26:01.002164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:1ce0af31c2079e2ab00500232421a46967e4ed66ec0bb2668d7eee18e9a4adbb

Observation e340eaad-bcae-43a1-8b69-8fb9a37cd780 · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.595190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:d49780701d9482c53296c56c8505f7489fbe4d99ef1bc3669a5d0a60ad53248b

Observation 8facd045-87ef-43d2-a3a5-1ab6d8c22dcc · outbound

This paper cites Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering , pages=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering , pages=

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.613991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:4751a37e792afa028b4751e943c0936e110adc9b4575fd04e0e9f84d59023041

Observation 6c939dc0-c918-492b-b3ab-5c1a90c9449a · outbound

This paper cites IEEE Transactions on Software Engineering , volume=.

Rethinking Code Performance Benchmarks for LLMs IEEE Transactions on Software Engineering , volume=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.173007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:5f05bc5a0154ec5076808aaa1ff93bd663b4a977043db1174dc90397e5fbe565

Observation 907cc97e-f9cf-46b9-8692-dca6caa951ed · outbound

This paper cites IEEE Transactions on Software Engineering , volume=.

Rethinking Code Performance Benchmarks for LLMs IEEE Transactions on Software Engineering , volume=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.180435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:a1982e06b4e34d0e046588265811cbb5ae680240e71c51cd998840d84b005626

Observation cf6fe173-407d-4b1f-9d17-894d3ca4fc8b · outbound

This paper cites Proceedings of the 38th international conference on software engineering , pages=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the 38th international conference on software engineering , pages=

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.587304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:86598b922f13985fa84b56438d3f3a07ccc878591f6dd1518091aa5ebebeee92

Observation ab3260c3-45c7-4c8f-b344-de75ed63c2f0 · outbound

This paper cites Journal of Systems and Software , volume=.

Rethinking Code Performance Benchmarks for LLMs Journal of Systems and Software , volume=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.652238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:d4f12b8669a51e91f380e6885c2f78745816b8e7904ce669c0a959bbb7a63939

Observation f604e6a7-bd4d-43f0-a540-0fc5dfe46f29 · outbound

This paper cites Proceedings of the IEEE/ACM 46th International Conference on Software Engineering , pages=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the IEEE/ACM 46th International Conference on Software Engineering , pages=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.178117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:ee3f1cc15e7a4c131b8803f45e6b9ef7de41fbc7ec626e8108b4633dc47f0039

Observation d0f53d18-cd9b-43c1-95a0-d8154c578290 · outbound

This paper cites IEEE Access , volume=.

Rethinking Code Performance Benchmarks for LLMs IEEE Access , volume=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.157088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:064e019fe67ed6f5b2dfd09c9da7f10b891faa8d4de1d068d9f2852b38282b5b

Observation 20541f37-aae0-4864-b8b7-a454d2fb2bda · outbound

This paper cites IEEE Access , volume=.

Rethinking Code Performance Benchmarks for LLMs IEEE Access , volume=

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.244350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:0eb49b9b15248a39cb13392155113d6a43fffe236ebeae3db7fb085900a2bcaf

Observation 2c420f38-9fab-47dc-80c3-f24f43f6dcd9 · outbound

This paper cites 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) , pages=.

Rethinking Code Performance Benchmarks for LLMs 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE) , pages=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.197483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:de618cd0ac17207d1557ed3181ffed35a4dfdb3e65d136ca36ff9af738e9aa43

Observation a56f1596-5f88-430a-b335-4129261f59bd · outbound

This paper cites Information and Software Technology , volume=.

Rethinking Code Performance Benchmarks for LLMs Information and Software Technology , volume=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.175079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:8462a3e76954bc06ed9d2fc229a00fc7fd333135e6e259b1f8047de690141de4

Observation 8cb649b0-311a-4c5b-ba59-8b49f5673e4d · outbound

This paper cites 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) , pages=.

Rethinking Code Performance Benchmarks for LLMs 2023 38th IEEE/ACM International Conference on Automated Software Engineering (ASE) , pages=

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:46:02.085408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:b1531597caa19578c35399fdedb41c49121b7feff87f8f86791ca1091cee13af

Observation 7c37771a-cd44-45bf-942e-d1f762178006 · outbound

This paper cites Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis , pages=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis , pages=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.181727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:6cbf99721d53fa4570c7d8f5298ac34d432684c5be8c53331437e24e3934c4aa

Observation f547b1ae-5aab-4844-adab-55af4b30c67e · outbound

This paper cites IEEE Transactions on Software Engineering , volume=.

Rethinking Code Performance Benchmarks for LLMs IEEE Transactions on Software Engineering , volume=

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.194201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:f077a936af10ab36646f8d4b65d6d0e6953804995f83fcb1d23d524020c6735d

Observation e18e37f9-3f1f-445c-8073-72f0a1eb9358 · outbound

This paper cites Proceedings of the ACM on Software Engineering , volume=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the ACM on Software Engineering , volume=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.210964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:6fb7a4e83a490039d2c274ef320037001cdbc44de89245b48648b8c20b3cc8ca

Observation 01a0de30-6fd4-4288-aea6-18978835a55c · outbound

This paper cites Proceedings of the 27th ACM SIGSOFT international symposium on software testing and analysis , pages=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the 27th ACM SIGSOFT international symposium on software testing and analysis , pages=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.220351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:721862d091fdef0b020e47a31ee440b37127cbd95e30e5a41823f770b68dbb0c

Observation 337665cb-881f-468d-adae-1c59ca0a4425 · outbound

This paper cites Proceedings of the 2017 ACM SIGSAC conference on computer and communications security , pages=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the 2017 ACM SIGSAC conference on computer and communications security , pages=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.592773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:50f863ab9276ebdba3187d953bc9c0fe670c56247c6a9bc1980e44d229fbd205

Observation 7c554f65-41fd-4dfa-9a5e-9c05b416b4b7 · outbound

This paper cites Proceedings of the ACM/IEEE 44th International Conference on Software Engineering: Companion Proceedings , pages=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the ACM/IEEE 44th International Conference on Software Engineering: Companion Proceedings , pages=

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.169194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:ab3c8847c3e3856e41ef9c17c6ac59c81fb54ee3224e57875ea64c66414c1fb3

Observation 4edbe300-8d74-452e-b799-8b5c334ab5e4 · outbound

This paper cites an unresolved cited work.

Rethinking Code Performance Benchmarks for LLMs Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-07-09T05:36:01.251943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:5784ccb0f1aac88a1d18e72a3b8bd764b190751fdf5468c28b626796d14dc6cd

Observation ed2840f6-bdff-48a1-8555-d6f1f08d9e4f · outbound

This paper cites Forty-second International Conference on Machine Learning,.

Rethinking Code Performance Benchmarks for LLMs Forty-second International Conference on Machine Learning,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.274104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:16a58c9519d8ccc088d1fe0c9c1e203da0aa960ca47c1387c6b94055dbfd511a

Observation 4b7d47f7-e172-4a6e-8beb-f6c2375c95b2 · outbound

This paper cites LLM4EFFI: Leveraging Large Language Models to Enhance Code Efficiency and Correctness.

Rethinking Code Performance Benchmarks for LLMs LLM4EFFI: Leveraging Large Language Models to Enhance Code Efficiency and Correctness

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:26:01.122754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:d25f627b085f9ecc3efed5e91c35987571c7a1b3ed656165f03716c15b969bf1

Observation 6bd5dfd4-3990-4fbe-bb1a-65a125f79080 · outbound

This paper cites Towards Better Correctness and Efficiency in Code Generation.

Rethinking Code Performance Benchmarks for LLMs Towards Better Correctness and Efficiency in Code Generation

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:26:01.063164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:6742cf0e1ffde76f42099d4ca5e0c65d1af4e51fb858874729681252400c2abe

Observation 222379f6-1f7c-441e-a3bb-67954cf7fff7 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.231582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:88abc5cf92ab08cff3bcfe8eea7c95b399d86ee83b0fff982d6d0d565b0bd765

Observation 363b706f-b518-43ef-9cac-f1d1d6850e6a · outbound

This paper cites 2025 , url =.

Rethinking Code Performance Benchmarks for LLMs 2025 , url =

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:26:01.131122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:d69df53199453106e72c42efdeffa4e2b1e17b1468361584a3b15f42ecd4a1b6

Observation b7632066-662c-4420-b287-2253f7fe7a85 · outbound

This paper cites an unresolved cited work.

Rethinking Code Performance Benchmarks for LLMs Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-07-09T05:46:02.108334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:2f3efa8cf66e65cf8e044ceb882ffe683469b61b9e0375f2889ce7f2adefab59

Observation 06895d48-b862-4717-8233-abc4ca6d2d21 · outbound

This paper cites an unresolved cited work.

Rethinking Code Performance Benchmarks for LLMs Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-07-09T05:46:02.052193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:f71febc1722c3790ff935a81c338107cebdb1cecc80d33b03ed74b2bea525ba9

Observation 70a49bd6-6160-4d79-946c-c35cf4b57846 · outbound

This paper cites 2025 , booktitle=.

Rethinking Code Performance Benchmarks for LLMs 2025 , booktitle=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:46:02.059319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:343e5766e8067bbe3438aaed65c0c5bc71475f9c7c53c0d3158cab6a3d65655e

Observation 9ef1df8e-ec1e-4bee-9e8d-20252471d1f1 · outbound

This paper cites Simon , title =.

Rethinking Code Performance Benchmarks for LLMs Simon , title =

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:26:01.114097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:d678d50da972b12c9856b40d5bb98fcda37c57b2de94b445e67822cca52f12eb

Observation 6e91cc5f-f75a-46df-93a2-85c12be4269c · outbound

This paper cites Shaw and others , title =.

Rethinking Code Performance Benchmarks for LLMs Shaw and others , title =

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.254539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:1f47c5d7443bbd64099b966c0023c6b12972f9704edb65d3fdf8fd5a9c498f0a

Observation a4ce6a93-8519-483c-890c-d82f54fd5027 · outbound

This paper cites In: POPL (2011).

Rethinking Code Performance Benchmarks for LLMs In: POPL (2011)

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T05:26:01.039755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:650595acce3838cb4c3ea72d9c30289ddec1e7da6bdd3b7aa901b26cdb9b0cfc

Observation 3539dbfa-2e26-4f8d-b88e-affce2135793 · outbound

This paper cites Waldinger and Richard C.

Rethinking Code Performance Benchmarks for LLMs Waldinger and Richard C

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.249679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:60cf82206de0484c558f9f5563e823852f8719124b0e31a12d6b811635d515d2

Observation aa4eb889-b347-463d-b1e1-154e5eacfd53 · outbound

This paper cites Waldinger.

Rethinking Code Performance Benchmarks for LLMs Waldinger

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:26:01.054980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:a33a4493229f1314e855a661e2aa93da7890185b9874194141edb7cf96c41dbe

Observation 05bdd701-2eb0-4ab0-a4e6-3ebecddc6bf2 · outbound

This paper cites Cordell Green , title =.

Rethinking Code Performance Benchmarks for LLMs Cordell Green , title =

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:01.208084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:8305e238cf75bb417a86ed4c122716438828a67000825def1e9e4c6ff090ff5b

Observation 8d6accf6-d999-4443-981a-5e3bd4ade773 · outbound

This paper cites 6th International Conference on Learning Representations,.

Rethinking Code Performance Benchmarks for LLMs 6th International Conference on Learning Representations,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:46:02.096776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:74899bd397a7808f975822df87bde256e15a60b24a51f75802c6b7f31fa117b0

Observation b1611c7e-97d8-4a1d-9878-af02dc5d75e1 · outbound

This paper cites an unresolved cited work.

Rethinking Code Performance Benchmarks for LLMs Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-07-09T05:36:01.191262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:f741c89f03a6b21cef48c0e336f81a93b6d3d15994c88c0f36d19bff7b95fffd

Observation d026c53f-b64e-4bc8-a598-531dcee893f8 · outbound

This paper cites Foundations and Trends.

Rethinking Code Performance Benchmarks for LLMs Foundations and Trends

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:46:02.079578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:38e9d0cc4f4a4ec5b5f272ffa439558b15e27cf534f8ff2375fb298e1f2553a1

Observation 77e5cd76-c43f-48e6-868a-e6446d139c28 · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

Rethinking Code Performance Benchmarks for LLMs Competition-Level Code Generation with AlphaCode

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:26:01.144105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:e98b9354977717fb7900dd9d1186af802776c0cbb21c0d0ccc169ddf05e6cab6

Observation 1247bc9e-46de-4ad8-bb8b-64a88aa3b93d · outbound

This paper cites CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis.

Rethinking Code Performance Benchmarks for LLMs CodeGen: An Open Large Language Model for Code with Multi-Turn Program Synthesis

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T05:26:02.119353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:52ecab57cd3f2cd1c1c617eea3b8b85da5993adbce4da3b03cba28c2803fa7ef

Observation 0cbfa65f-c166-427b-8a5a-03d1dc53dcfa · outbound

This paper cites In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, December, 2023, Singapore.

Rethinking Code Performance Benchmarks for LLMs In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, December, 2023, Singapore

Reference 75

Resolution
verified exact
doi, observed 2026-07-09T05:26:01.066720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:8bb122b8e08375c6f25ae1506fb7a286ef2706ba1b015f19cd92afd6e4277dba

Observation 1cdf3c40-124e-4554-839a-8ac2652d59bf · outbound

This paper cites an unresolved cited work.

Rethinking Code Performance Benchmarks for LLMs Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-07-09T05:46:02.102413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:86f540be59f444ca268cb0e7a2053425c8755b92704bfeb4491671dcb3f67b9a

Observation bbe5679d-44b9-43ca-b627-2af5fb9ba813 · outbound

This paper cites an unresolved cited work.

Rethinking Code Performance Benchmarks for LLMs Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-07-09T05:36:01.276676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:c15afad21f33d4253ca0940a8211095d117b17b5cfcf2fb69d39f82586b2fb52

Observation b4645939-b944-485e-9a44-f72731931aab · outbound

This paper cites InCoder: A Generative Model for Code Infilling and Synthesis.

Rethinking Code Performance Benchmarks for LLMs InCoder: A Generative Model for Code Infilling and Synthesis

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T05:26:02.096721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:44d73a48cc2fa474dd0f66e852d05440ae36bbf2aa22bcbd982f2f67a4ad23b8

Observation 70853840-a8d6-487a-ab3d-982586435e2d · outbound

This paper cites SantaCoder: don't reach for the stars!.

Rethinking Code Performance Benchmarks for LLMs SantaCoder: don't reach for the stars!

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:26:01.048861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:03c7214986b600d15a12b2967a18ad7169bab5722c82b81e39acf0685656fc9e

Observation bc90ff6c-ec32-4bb3-96a4-d14ecf167a69 · outbound

This paper cites Language Models are Few-Shot Learners.

Rethinking Code Performance Benchmarks for LLMs Language Models are Few-Shot Learners

Reference 80

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T05:26:02.115697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:0e82d1a4e21d84216fc7cf6f98e03ca71c885accda1049f4b63f8ad3e8139232

Observation 0cb8bd0c-01d1-42e8-873a-6479ea1e19a0 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Rethinking Code Performance Benchmarks for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 81

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T05:26:01.126601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:d15f48429b89ef73d4b97ec7b253b25b49d5a4455ddfe86692b3ede3d96a1461

Observation 654025cc-6061-4682-bf08-567956349037 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

Rethinking Code Performance Benchmarks for LLMs Code Llama: Open Foundation Models for Code

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:26:00.996749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:20d2c98c5e3f8c913181ff1904673a0ddba1cb345bbf0a7da9987c77f35bc204

Observation 63158b33-9d39-4c6d-b802-5857273c2c2d · outbound

This paper cites Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models.

Rethinking Code Performance Benchmarks for LLMs Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:26:01.075601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:cb21b830f75c369429fbc4a3d29354bc96ed0fc07a75525fe69116bb3a17a2a6

Observation c2a040a2-0cf0-41b5-9465-f474eb124094 · outbound

This paper cites GPT-4 Technical Report.

Rethinking Code Performance Benchmarks for LLMs GPT-4 Technical Report

Reference 84

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T05:26:01.148862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:d2998d2e5791cb08613933291569d2e260f7a9ba8a2475c4d12f3c24e1ea8416

Observation 8558cbea-3730-4633-b8c1-e8a1a81d565f · outbound

This paper cites Llama , howpublished =.

Rethinking Code Performance Benchmarks for LLMs Llama , howpublished =

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.638824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:fce072a1713933bdd09e0cf320c04050c8bffc49abbe0fe873c0f7e9c854023f

Observation 857b7a33-6bc3-421a-b3af-88b9b30255b7 · outbound

This paper cites Anthropic , howpublished =.

Rethinking Code Performance Benchmarks for LLMs Anthropic , howpublished =

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.631833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:7606d1986c85b9450b25c4803574865b460a906c46e8c3a72354b2c4d77feeac

Observation b642fd0b-5eb1-4abd-93ad-ad2cf96c0cbc · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Rethinking Code Performance Benchmarks for LLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:26:01.015668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:e2052677ba11d9c37de74ceb96b4d75ac733e9b42d8479538d3f5a1403f9080f

Observation 55bd9041-6d7d-4cca-8e49-ec24a0677a41 · outbound

This paper cites Mixtral of Experts.

Rethinking Code Performance Benchmarks for LLMs Mixtral of Experts

Reference 88

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T05:26:01.010795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:2c4f192b1175cf70e71372132afc18093181e0beaac97ee185b0015b59f47269

Observation e164568e-c3d3-4981-91e2-85f1d44d3d52 · outbound

This paper cites 2025.IRFuzzer: Specialized Fuzzing for LLVM Backend Code Generation.

Rethinking Code Performance Benchmarks for LLMs 2025.IRFuzzer: Specialized Fuzzing for LLVM Backend Code Generation

Reference 89

Resolution
metadata mismatch
arxiv_id, observed 2026-07-09T05:26:01.105667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:254c75bf11c46d85a8b3afe40b7c257037e69bd8b80b7d3145eedd3d194cb57a

Observation 8d769afd-e257-4eac-977b-2723e0160da7 · outbound

This paper cites Crash report enhancement with large language models: An empirical study,.

Rethinking Code Performance Benchmarks for LLMs Crash report enhancement with large language models: An empirical study,

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-07-09T05:26:00.990580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:eb2742ca7cfe6013e43d2ff2f1e9bee5be1d053a90365a5c0ce271fb2e3ce297

Observation 99af4452-8485-45f8-a4f7-43892cf8e401 · outbound

This paper cites IEEE Transactions on Software Engineering , volume=.

Rethinking Code Performance Benchmarks for LLMs IEEE Transactions on Software Engineering , volume=

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.634196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:a7d972712c546a2403dfd932f528c210f2f9ba9f4d7d8a217578ba8a21e4ce0e

Observation 1382c003-154e-4081-8fdd-0f9514fdc0a8 · outbound

This paper cites Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering , pages=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering , pages=

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.641152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:c1a75c5ac86285cc4dacc725169f10916e60819e29bd897a004157c1d389ee25

Observation b6d9865a-4ec0-49c2-ab5e-0d9244f2a7ba · outbound

This paper cites Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering , pages=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the 29th International Conference on Evaluation and Assessment in Software Engineering , pages=

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.627088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:9f21049e8856d27d77f78a9ab38fdbc8e04637ecb058807afe5bb88a523b591c

Observation cac72df3-d98a-4078-a8e4-1c51ad67112e · outbound

This paper cites Proceedings of the 2024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineering , pages=.

Rethinking Code Performance Benchmarks for LLMs Proceedings of the 2024 IEEE/ACM First International Conference on AI Foundation Models and Software Engineering , pages=

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.629457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:b2b1d58aa0e7a7d54e5e94b401be1c4af663b75df0898c6cb6df0cd4846bad47

Observation b4b288dc-4b9e-4073-bd4b-64da2a3d7d92 · outbound

This paper cites Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes.

Rethinking Code Performance Benchmarks for LLMs Is Functional Correctness Enough to Evaluate Code Language Models? Exploring Diversity of Generated Codes

Reference 95

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:26:02.133748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:c06bd54bd1735ed847b7b29e519bd25bec8d20ab3227cb93f9a0948caa427513

Observation b6792b08-cc59-40e8-bedd-e0c0deab4fb0 · outbound

This paper cites ACM Sigplan Notices , volume=.

Rethinking Code Performance Benchmarks for LLMs ACM Sigplan Notices , volume=

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.645017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:89e6e1d735cd122a2de10c647856a22805b6620f9a36a05ded9dc65474bf24bb

Observation e544fd1b-0ea2-4ed6-95f3-3c842f4804c7 · outbound

This paper cites ACM SIGPLAN Notices , volume=.

Rethinking Code Performance Benchmarks for LLMs ACM SIGPLAN Notices , volume=

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.656462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:ea35a9384f5b1298de56d0a5381c2fcdbb044610f381c63020bf1fe75a83ef3a

Observation 7c261bed-8a12-4839-bacf-c4da5422c85a · outbound

This paper cites ACM SIGARCH Computer Architecture News , volume=.

Rethinking Code Performance Benchmarks for LLMs ACM SIGARCH Computer Architecture News , volume=

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-07-09T05:36:02.609232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:b02e916ee7cba7c771d7f668dbb55443692a8fd2e44c66e6f22be291e1073e72

Observation 2c284cb9-cfad-49e7-a13b-aa1c28eee140 · outbound

This paper cites Evaluating Language Models for Efficient Code Generation.

Rethinking Code Performance Benchmarks for LLMs Evaluating Language Models for Efficient Code Generation

Reference 99

Resolution
verified exact
local_arxiv, observed 2026-07-09T05:26:01.186731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:274204d0690b478ef14df434209fe6f632c0d3a16f67c9eece2e68b29df1e1a7

Observation ef2278a7-81c0-4b47-b429-ec0ac5164477 · outbound

This paper cites an unresolved cited work.

Rethinking Code Performance Benchmarks for LLMs Unresolved cited work

Reference 100

Resolution
unresolved
raw_fallback, observed 2026-07-09T05:36:02.569738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-09T05:16:58.549058Z digest=sha256:4b231ca6b4536270203ded9766d44d44c6d2ac9a2b554957ae9b78c9333044ca

Pith citing papers

Observation 1b53eb0d-ebe6-4be9-9aee-8b7a819913ef · inbound

RLPF: Reinforcement Learning from Performance Feedback for Code Generation cites this paper.

RLPF: Reinforcement Learning from Performance Feedback for Code Generation Rethinking Code Performance Benchmarks for LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T10:53:08.900571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:53:08.900571Z digest=sha256:f93a53c553ae845e64b59b113f4ee06a6960f320a68e9ebcd7bb6d77a5d20e86