Pith. sign in

Paper Citation Record · LEDGER

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2506.00777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00777 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:47.526084Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:20:33.617934Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbc6415e-83f8-4705-a460-ea733bb3876e · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.279759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.279759Z digest=sha256:f22738d07224ec2f338807a3794b1d38bb375a8e3953a7685d85543c6254fbe8

Observation 7106c982-2c18-40da-b7fc-c50c17e8645b · outbound

This paper cites GPT-4 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.374841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.374841Z digest=sha256:e4da540f8814bf81f7421478b0db2a000151c6553a94b06d50abc7299b0771ef

Observation 77d10d3b-a33b-47cb-84f2-7a33f045acd9 · outbound

This paper cites PaLM 2 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge PaLM 2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.455923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.455923Z digest=sha256:6cf085a71502aa41904215ace455fcd9bf16d08fadfe255a16d8fdc71289ac92

Observation c8b80b04-3649-4f7b-a8e1-866fb883057d · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:52.076468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:43.506723Z digest=sha256:b259da99c370af532a4cb90311674b8020c65cc0396de5015d950ef4ae848036

Observation 2b483b88-b595-4566-a055-62d9d73d343a · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.570013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.570013Z digest=sha256:033ca853885caa7e7b5ccbe5c447eed194581ae78ccdccb89602d4a32c68a17d

Observation ae6a5b49-c271-4d04-bbc5-f107955debdb · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Do, Yan Xu, and Pascale Fung

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.667380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.667380Z digest=sha256:25cb902923f615c3a64b92c98d800b8aa8b1f5015238d5291dd3c0099d5a7cae

Observation 14591c23-b534-4fe1-a311-27e02139c2e3 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.774395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.774395Z digest=sha256:12bb72d5b6d1c1179a8fbad4647f73392a9c7c9f1fcc82ee2cb6cfe4ae4f854f

Observation dbe47d2d-c4e9-4e09-94f4-5926e32a1963 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.811421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.811421Z digest=sha256:8d17922563c2a16518413482c08e25ef8f026e371a3c3ade90d987465e8ef8b1

Observation fdc7ff7e-8cd7-4932-bb5f-d44032f25b2e · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.799121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:43.908302Z digest=sha256:1e017d807ae716486ac22d95da19f91c1bb95d8f60549700262ad20fa438aa65

Observation e46f93ed-c5e8-4406-8463-fbfd8a3d8f3f · outbound

This paper cites Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.971142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.971142Z digest=sha256:a63ab11e020f83adbe8a39f2c0898ef86b5c8380b9441fd4454cfa1e76eaa59e

Observation 6ec81de2-7bac-4235-9bde-7f605eb4f692 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.515495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.045874Z digest=sha256:2e8e8f8c5130fe8e47fffb92f0470787cb442b552e8fda5667e86024a584ff45

Observation c16909de-1bcd-45b2-a1ba-703ef69ca796 · outbound

This paper cites The Llama 3 Herd of Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.092977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.092977Z digest=sha256:c7802b724b6215f80f92beaba6fe8bc7026f3b9bfbd0a161267e250006d96e41

Observation 50cf506a-f7c1-4a37-a1bf-2367ccb2b5fb · outbound

This paper cites A Survey on LLM-as-a-Judge.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge A Survey on LLM-as-a-Judge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.191039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.191039Z digest=sha256:cf7155f062608a73266c81796ce6c140b38e2fa970f4d1ca77ac1d57e5948bf0

Observation 5c4fc371-c7ee-459a-a98d-c2389ef9501d · outbound

This paper cites Textbooks Are All You Need.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Textbooks Are All You Need

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.278334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.278334Z digest=sha256:6a52507aaabfbcb67a6c5ec520e0198afba168690f5d7249c213e2e763bf3891

Observation 273682e8-0e02-4749-83a5-286b9146a2da · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.337098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.343234Z digest=sha256:680ad53b50658ccd66c588e956ffa36449c149c57c8ae92230db7d8ec90c6c71

Observation cbae3a98-ca16-4dd6-bbdd-44dca82ce701 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.158245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.450715Z digest=sha256:4f90100bd301cd6e2034c590267a2a74506aa5398ab5626d5921ed87fd7116a2

Observation 318fbf0a-8b9e-46ba-b30e-1bc1ba27c0f1 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.973844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.579781Z digest=sha256:c2005f42f944c9464e221e6d3f57064284f480ad7656bba3c774e2372e946dc5

Observation 0acd8fa1-ae60-42a3-a3ab-4e4fcb1837a8 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.819956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.624950Z digest=sha256:089d42ac4590bbd2aebb395689ccc643778d688f3f7e9b20a77b17870e665ab6

Observation 750697ef-7b14-416b-8c6e-a9e46d38f0a4 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.701200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.680467Z digest=sha256:ba82691385747257f1a206634c1e4b6cb579e420ded13ced2e78023a25c594cd

Observation a37e638d-10f7-4dbd-a7d6-7fc50ad630ae · outbound

This paper cites Mistral 7B.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.774585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.774585Z digest=sha256:0cd96c9eb265494db8d52d3e4b65864e80c9ed4cd1badbbfa0fd8ab2fd652c2a

Observation a5c456ea-1544-4d02-b8f0-e6aad9adf624 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.846828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.846828Z digest=sha256:708c761b5f908fa7eb2fcb75705183fcde0802eb8b889dddf6dbd458d3b7fde6

Observation 9435f877-6007-4322-a81e-c5810a4c9458 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.551364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.901982Z digest=sha256:c0175386adbd9174047026a0d3ff6f49afd29598c0b919826beeca59433f948c

Observation 81544950-e252-4a74-9359-af0383cc8fb3 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.386188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.985167Z digest=sha256:0aa86d036ccf4ca16c76dc0c6cf6390cb280f8e9235145baf7a018909bb2e17d

Observation 26eb1209-bb7b-4000-a877-b69af6dcc843 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.234920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.060211Z digest=sha256:80f1f1a1b9bd1a3ba5dd043408a4d8c1443ca4ba8a592a9943ee9c0608076549

Observation 2560fe2a-3da9-4493-8ab7-659330931aed · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.062140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.123893Z digest=sha256:a1061e85dd8ca4ec135b1c6b02e212c2f21479e8976ff4a5ff4e2066695170ce

Observation 298d0589-20bc-4e25-85c1-96296dd9b862 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.164069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.164069Z digest=sha256:e69b444ae3e2085df29f8d905a4a41fb3670f22e6159d9d6dc63a475b55b4f41

Observation 06bf94d4-9c2f-48d7-9bff-76396c9df471 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.844122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.199548Z digest=sha256:aeb52d03467926e4e171bf684c5ade83a005468623fa588f40ef31bcca40ef5a

Observation 4a2ef83a-53bf-4b90-80a8-e7d056c11299 · outbound

This paper cites Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:02:48.093541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.290988Z digest=sha256:4f3806d376e334531c2182109cbc5ef5478a872afab9d184173b6316ed39dc12

Observation 5cadf5ed-49d7-46e9-bd80-4030446042c0 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.699229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.386523Z digest=sha256:e6243860223420f556a5e1a9c0afc9bad703d388db1673a168d566aea03f2176

Observation aff110c5-0e2c-4e02-a339-daba80ea5fbf · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.447467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.447467Z digest=sha256:377e89a881a9b804ba495d823995096d2e5e057ff8b299a3307dc7020aa6b331

Observation 83333015-ee06-4a8d-b60c-dfcf78e74003 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.488172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.522250Z digest=sha256:418aa12e9cf0b006a6bc863b207cd19490f5942d922365948d3d851c5d38a84f

Observation 42e80012-10ba-4cc7-8113-db1af6a125ff · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.662507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.662507Z digest=sha256:1e1dbec8cbfd2a5033046b37eb4e8b5735a3346f7b9935825ac322329939b17c

Observation 2095bc52-8ce1-4095-8456-d2006553892e · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.258384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.735422Z digest=sha256:2b8451c51189c26ca963915df3c34f28c5fb8d4ae504201be555c163ab440e1b

Observation 92117b81-12ff-469f-8f40-7013e7db6479 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.070708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.811541Z digest=sha256:4750df65632a43d3ec2754ce380759abbc19f5ac68f64dc085bdec5c33e540ec

Observation ce3edb10-ee22-421e-b374-a70e25b9e77d · outbound

This paper cites Large Language Models: A Survey.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Large Language Models: A Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.923985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.923985Z digest=sha256:383d66eed7f2ddb97ecf57a69180b6098ed00788131fca381a0640d006449b85

Observation 97fa6cba-6fe9-4112-b633-bd51a6587bc9 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.028316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.028316Z digest=sha256:be9c8681b29abde7d6c7401ca1946cbc85ab5dd90c097936ab07912dcb0da851

Observation 833e2177-030b-4af2-bc95-7632073d094d · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.903055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.066169Z digest=sha256:8ebbb20b058b26d4370d5bb5471cd99e5851be8ea670ea567acc8e746613b9f4

Observation 26328e55-814b-46e0-b17c-7f2f17807602 · outbound

This paper cites Capabilities of Gemini Models in Medicine.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Capabilities of Gemini Models in Medicine

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.105692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.105692Z digest=sha256:5dde64b6605529f5730d21851b61fcc0bee18d9f1a9cf50ebd050f893d5629d7

Observation e1eb2724-5824-4238-ae81-2151f47b44b0 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.696990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.175013Z digest=sha256:f0d0d76e40354805a494ff7fad325e41a280447173cb4f87677775d56c36dc82

Observation a1c00059-8291-4ba1-8775-412fa7dbda72 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.272794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.272794Z digest=sha256:e517d01f263c8af08fa5e1f4726b22fc1c81e37c1b00944bd230592152b9e3a5

Observation ae4979be-f501-4ceb-9281-b6eb49557a47 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.559413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.351173Z digest=sha256:cd2297d8520587a44afd3438728b42066455b5974f8402d788b327b442e12d16

Observation 314fb2e5-19f0-4af3-9c87-50d91fa7f7fd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge LLaMA: Open and Efficient Foundation Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.447687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.447687Z digest=sha256:44154845fe7ebd47c26fc6c1f4ca53c7cb4bbf22c5e4709ae24a86ddcecc8d92

Observation e742cba8-03c6-46fa-b04f-7a509c554f3c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.559562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.559562Z digest=sha256:a4d0aa37ae8c9dd750d2ba58b29dd9489ea8080efe7a372781c1030dfb2aeeed

Observation 68bca6ec-97b6-4e23-84de-f71a635099d1 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.635872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.635872Z digest=sha256:b3da3bbc20a1fa302116d332602294e14ba89a5c91f921aa80d87bf4fe44cb6f

Observation a894cab8-c0f5-4d16-a157-c7161b72cab6 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.689377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.689377Z digest=sha256:94a55ca6065a98ce57e3b7f6175608511fe65b5f09fe8f3f7bacfd86436abb8b

Observation 4663eb2b-2686-445f-a5d7-a8a7b0df7032 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 46

Resolution
verified exact
doi, observed 2026-08-07T12:02:47.762727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.767706Z digest=sha256:d677e1746ac7ba9dfc28c1f99a410e857abc155b3f17a719c9b369b16e21f244

Observation 30c32547-5fa8-4a28-bec9-fa48dd432be1 · outbound

This paper cites Qwen2 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Qwen2 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.842724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.842724Z digest=sha256:3e2d6b9c4e5c5863e80e7a196b8be82da7726960872da5c88684835209129e5c

Observation 95adc132-2783-4092-a826-74124f853fc3 · outbound

This paper cites Qwen2.5 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.912360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.912360Z digest=sha256:bd2e77f4c28b4d9b0ab76d26de9838bd0721cf63647c7567c9524d3611edd97a

Observation 58ebfe33-87a9-4cf2-acb7-27659c5495b8 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.010722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.010722Z digest=sha256:889c1ee52e533c0d641f2ef6755821ddf2526fc254aa7baa5b9061565cadd862

Observation 04287b94-8cbe-4d13-a478-4cbb1f7b122d · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.084587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.084587Z digest=sha256:f5af7b44abd2398b7d0d76c9468d058533d963c99ddf3cd12361d10e1a5e2790

Observation 82a3b566-caff-4747-837c-9698baec91c2 · outbound

This paper cites A Survey of Large Language Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge A Survey of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.192473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.192473Z digest=sha256:b126a3401aa5fe19db82a50e79ba0a919191633ec907b1e0adef48ea92a2b6ca

Observation b3160186-47ab-43f5-aa75-3981d65cb121 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.241849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.241849Z digest=sha256:95104450eb254fe29b2c049044c14c0b016a289767db9ec91a68b66d3bb0e380

Observation 29b03e43-fffa-4182-a8a2-54e8643a8ea0 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.359311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T12:02:47.356554Z digest=sha256:19c354c0a378f97d63dd64be4e72b15ab03c32c2fb73d48bc080966339d277ae

Observation fa99093f-eff0-48e5-8bd8-ad20e2821bd3 · outbound

This paper cites online" 'onlinestring :=.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge online" 'onlinestring :=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.459081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.459081Z digest=sha256:0e317004038567c310e7e9a4bc9d3f0563a5b42ae319726f60a21c670a1f83f7

Observation 65b484b4-22af-470f-9f3e-ec5db528a1a3 · outbound

This paper cites write newline.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge write newline

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.526084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.526084Z digest=sha256:6204edd5cb7316191ee409be48ca8e675938263ea64d5b62710c5c79cfa26f60

Pith citing papers

Observation b2697847-5dea-482a-82da-5a066fb655a6 · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T11:49:31.595873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:49:31.595873Z digest=sha256:3d68ae22dc3e09f5eaada312b884c80dada66eed9bf5395643e72f75946de3c0

Observation 3d01d9e4-23ce-496d-8e1e-f52f7e92bc81 · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T07:20:33.617934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:20:33.617934Z digest=sha256:e91792461f8ff8b4b33b7cff5d5a28e73fddaf8f3d69e4cfc1bbd76274aef789