Pith. sign in

Paper Citation Record · LEDGER

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge

As of 23 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 2 inbound Pith citation observations for arXiv:2506.00777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00777 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:47.526084Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T07:20:33.617934Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cbc6415e-83f8-4705-a460-ea733bb3876e · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.279759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.279759Z digest=sha256:77bfb59e6c7fb4f7b4aa9ec5895565f85f20c14d5c3d9045acff661fad337124

Observation 7106c982-2c18-40da-b7fc-c50c17e8645b · outbound

This paper cites GPT-4 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.374841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.374841Z digest=sha256:f755adbf6881058d53263d8b05f1dad0a0a111f901deda60f525f1bd8ce4da55

Observation 77d10d3b-a33b-47cb-84f2-7a33f045acd9 · outbound

This paper cites PaLM 2 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge PaLM 2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.455923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.455923Z digest=sha256:b81efb257c37068326158975dd9cee10029e48e2116da73f00358e0a853b28dc

Observation c8b80b04-3649-4f7b-a8e1-866fb883057d · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:52.076468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:43.506723Z digest=sha256:c6336964561cee9f8fe5e80b4d576fdd0bced25b0b9b537cef88d76c0f17e9ca

Observation 2b483b88-b595-4566-a055-62d9d73d343a · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.570013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.570013Z digest=sha256:113f9ca285c9788719f1f8bb36655aae3f08d23cfdfe37620b8b19e0ba945ab3

Observation ae6a5b49-c271-4d04-bbc5-f107955debdb · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Do, Yan Xu, and Pascale Fung

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.667380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.667380Z digest=sha256:83ac7532828cb9e55d461371a993be76892392865c887e8e4494d95c14b34627

Observation 14591c23-b534-4fe1-a311-27e02139c2e3 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.774395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.774395Z digest=sha256:0cae9073c2e266313e2f976fd49ee5b8bf21cc265853a7ced912c3c7b750ea53

Observation dbe47d2d-c4e9-4e09-94f4-5926e32a1963 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.811421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.811421Z digest=sha256:49955d74a527165972e54fa1bfce1c284d026b82cde63715d88cbdb387c9676e

Observation fdc7ff7e-8cd7-4932-bb5f-d44032f25b2e · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.799121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:43.908302Z digest=sha256:e84fd786db7082f587558bccc9efbe128079b7567e3964123e977a9b2d853850

Observation e46f93ed-c5e8-4406-8463-fbfd8a3d8f3f · outbound

This paper cites Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Tiny Titans: Can Smaller Large Language Models Punch Above Their Weight in the Real World for Meeting Summarization?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:43.971142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:43.971142Z digest=sha256:5bfc6163afb294d16adf9a2f86b9cd62184263fb2ec519cfec63d3393ef6df21

Observation 6ec81de2-7bac-4235-9bde-7f605eb4f692 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.515495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.045874Z digest=sha256:6c13aa617b00ee469152f9f85a09600c46cd882478fd807dd13c065e7c906d44

Observation c16909de-1bcd-45b2-a1ba-703ef69ca796 · outbound

This paper cites The Llama 3 Herd of Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.092977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.092977Z digest=sha256:e7384fb8d3bebeff769290b91cc53d534df7153e75ef3cec8e0d5821027bc338

Observation 50cf506a-f7c1-4a37-a1bf-2367ccb2b5fb · outbound

This paper cites A Survey on LLM-as-a-Judge.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge A Survey on LLM-as-a-Judge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.191039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.191039Z digest=sha256:5865b14c5340bd46346d79dd89a11858452a3f69e5064c83e3d549369506fa25

Observation 5c4fc371-c7ee-459a-a98d-c2389ef9501d · outbound

This paper cites Textbooks Are All You Need.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Textbooks Are All You Need

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.278334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.278334Z digest=sha256:cd51968f181c951981e5156a4d3afebee0d2717fa18022d2fe650c78def36538

Observation 273682e8-0e02-4749-83a5-286b9146a2da · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.337098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.343234Z digest=sha256:ec0f3915d2cd96175a5b637bc76b007278b374d6b3e2d3dd79aa152ff8aab605

Observation cbae3a98-ca16-4dd6-bbdd-44dca82ce701 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:51.158245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.450715Z digest=sha256:de75965da9af31e3ab03b13ffa922fec930a707d0059707586070dd6e946960e

Observation 318fbf0a-8b9e-46ba-b30e-1bc1ba27c0f1 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.973844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.579781Z digest=sha256:69afcf8f18646659018ea38c62ceef0c3433b3d53ee8a1ff58b2cb277f426f94

Observation 0acd8fa1-ae60-42a3-a3ab-4e4fcb1837a8 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.819956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.624950Z digest=sha256:13fb01a249186f2c25a6122c321377e065fabbd3c2f3f2a6f720887d1129f84f

Observation 750697ef-7b14-416b-8c6e-a9e46d38f0a4 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.701200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.680467Z digest=sha256:29d3045e356489c6b45e415f036ef01e779131099d5ccd0830386e2227d4b607

Observation a37e638d-10f7-4dbd-a7d6-7fc50ad630ae · outbound

This paper cites Mistral 7B.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Mistral 7B

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.774585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.774585Z digest=sha256:e95d83cf4a6a19a155437d6fa48a713f3eee3fc19939d2cb14b2814065e1ac79

Observation a5c456ea-1544-4d02-b8f0-e6aad9adf624 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:44.846828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:44.846828Z digest=sha256:a3397601ee1736ea9ab4669f3a8ef569d62b477ad1b1c09c220fe063dc6da129

Observation 9435f877-6007-4322-a81e-c5810a4c9458 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.551364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.901982Z digest=sha256:8072bf3f986e993879ecf56c18d2bbf042acf466e375c5b7fdc9aed4a6169422

Observation 81544950-e252-4a74-9359-af0383cc8fb3 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.386188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:44.985167Z digest=sha256:452ce674dea3e16ce53be3060f697cb855468750b68c4b683f8c8f51eb0c1102

Observation 26eb1209-bb7b-4000-a877-b69af6dcc843 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.234920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.060211Z digest=sha256:3d49a88baac4837002748a0ca9d1b2948835d36cb9f361dc50421516b28405e3

Observation 2560fe2a-3da9-4493-8ab7-659330931aed · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:50.062140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.123893Z digest=sha256:cd76ec95f5390b733dd5ebc5bab71c5cefbcd89b435b5da3bdbfd6958b3280e9

Observation 298d0589-20bc-4e25-85c1-96296dd9b862 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.164069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.164069Z digest=sha256:4671f9dc2d39727a50403cdbccb5770751495883042ad640af7de3e34b8ac32f

Observation 06bf94d4-9c2f-48d7-9bff-76396c9df471 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.844122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.199548Z digest=sha256:b9ba0ec626a184753576098fe6ec86f1f4878e6e4f775577d2d81c352e5da415

Observation 4a2ef83a-53bf-4b90-80a8-e7d056c11299 · outbound

This paper cites Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:02:48.093541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.290988Z digest=sha256:dc807285d6ee0a16b7540502ba866684823d82660f7cedb8f50472029428b4f1

Observation 5cadf5ed-49d7-46e9-bd80-4030446042c0 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.699229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.386523Z digest=sha256:5037387d18882a446aabad049aa4a123a72fd281e53bd65798dc495b0988bf97

Observation aff110c5-0e2c-4e02-a339-daba80ea5fbf · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.447467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.447467Z digest=sha256:8a0b7ef6980ecc2d0f8af5b4c24c74a5af4eab48984e204db784e35423ba6ffa

Observation 83333015-ee06-4a8d-b60c-dfcf78e74003 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.488172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.522250Z digest=sha256:f5249582ef2c8a6ca99d4a5fb071515a8bb2a76f5e8ae7bb290851f2e204b3c7

Observation 42e80012-10ba-4cc7-8113-db1af6a125ff · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.662507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.662507Z digest=sha256:5ec8ea7b853fe2769fdd8a3ec496fadf6045fe34b959dcd9bd52565ac31da306

Observation 2095bc52-8ce1-4095-8456-d2006553892e · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.258384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.735422Z digest=sha256:7fe7187c1f57fe33059cb0035381d1780c366df5e7c816b00e38f94b5eee8b0b

Observation 92117b81-12ff-469f-8f40-7013e7db6479 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:49.070708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:45.811541Z digest=sha256:23e81b2461f479299b71720fb9de0cca6617c8c9c62f7a72618b9f385576769a

Observation ce3edb10-ee22-421e-b374-a70e25b9e77d · outbound

This paper cites Large Language Models: A Survey.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Large Language Models: A Survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:45.923985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:45.923985Z digest=sha256:49c6a265dab54e35f07b074cb6950a96677171540deb2fda0d064d0add14f93d

Observation 97fa6cba-6fe9-4112-b633-bd51a6587bc9 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.028316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.028316Z digest=sha256:1abcd1f07dc9e6aecd7add2b6315c9095e065b56fcf7f9d7b9196ca9b358dc23

Observation 833e2177-030b-4af2-bc95-7632073d094d · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.903055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.066169Z digest=sha256:6a9c23585c2637e0fd6fec8ad4258e39a845869829885dc8854fa0a25f617447

Observation 26328e55-814b-46e0-b17c-7f2f17807602 · outbound

This paper cites Capabilities of Gemini Models in Medicine.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Capabilities of Gemini Models in Medicine

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.105692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.105692Z digest=sha256:d335a776b90636b2695677166aa58a4a606b7000de90e1de83026ebe1464d38b

Observation e1eb2724-5824-4238-ae81-2151f47b44b0 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.696990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.175013Z digest=sha256:6d2e08a8c9ac5add747a5b4164c046c1e8e6d12648399e8939b15683e76854bd

Observation a1c00059-8291-4ba1-8775-412fa7dbda72 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Gemini: A Family of Highly Capable Multimodal Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.272794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.272794Z digest=sha256:a9eab1ef5189c8034a4210808dfa597e341c3a6446d4c061949418e677dea119

Observation ae4979be-f501-4ceb-9281-b6eb49557a47 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.559413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.351173Z digest=sha256:b1ad8e58a379fafced675ce46db0e043559be3459df176a2aff8fb74d9f4c017

Observation 314fb2e5-19f0-4af3-9c87-50d91fa7f7fd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge LLaMA: Open and Efficient Foundation Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.447687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.447687Z digest=sha256:38172928ea085baf4784145ab51a4ac46283c69ff022e532d2d04a0d201629dc

Observation e742cba8-03c6-46fa-b04f-7a509c554f3c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.559562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.559562Z digest=sha256:53a6154bb0b1af0bf1ee612099c5310fab295dd93b51401a5df966bb0b7b66c8

Observation 68bca6ec-97b6-4e23-84de-f71a635099d1 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.635872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.635872Z digest=sha256:7dbf97fb3b90bd77e365772efde83cb2b2ad1e67c1e134371c59b9ddb2a22cc4

Observation a894cab8-c0f5-4d16-a157-c7161b72cab6 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.689377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.689377Z digest=sha256:56347c4d070b076d7e8a15305cb18ef9a09fc3b0d6aef6271debf28c219994c6

Observation 4663eb2b-2686-445f-a5d7-a8a7b0df7032 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 46

Resolution
verified exact
doi, observed 2026-08-07T12:02:47.762727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:46.767706Z digest=sha256:c3150944cb483d79721e821d3de3e1b5cd3b0f70f23229b3077cc43d3e1d4a67

Observation 30c32547-5fa8-4a28-bec9-fa48dd432be1 · outbound

This paper cites Qwen2 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Qwen2 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.842724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.842724Z digest=sha256:677fa9085281f0f90e606ff93c9dc7f0639222f2a08bcfe26b25ab5cb9ba5d5c

Observation 95adc132-2783-4092-a826-74124f853fc3 · outbound

This paper cites Qwen2.5 Technical Report.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:46.912360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:46.912360Z digest=sha256:f76b0e9189f4253bff31268137289faa6ee826bd37102cdd4e336de6b659782d

Observation 58ebfe33-87a9-4cf2-acb7-27659c5495b8 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.010722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.010722Z digest=sha256:45b6426a16bf3432535314c975a0309e8a59fdd5f08a41494540711f43745d6e

Observation 04287b94-8cbe-4d13-a478-4cbb1f7b122d · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.084587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.084587Z digest=sha256:3a8d607bee96f5643b45e6b4113eb4cd2370bb785e98d9033a3a40aa17949b1b

Observation 82a3b566-caff-4747-837c-9698baec91c2 · outbound

This paper cites A Survey of Large Language Models.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge A Survey of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.192473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.192473Z digest=sha256:914842b3647252a39437160da6009bb0436945f5b344208d4637d97d5495b535

Observation b3160186-47ab-43f5-aa75-3981d65cb121 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.241849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.241849Z digest=sha256:b3e5031c1c5447978d7b6327195f42d5cf619c09bea12170bd4ae9d12cfcf673

Observation 29b03e43-fffa-4182-a8a2-54e8643a8ea0 · outbound

This paper cites an unresolved cited work.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:02:48.359311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:02:47.356554Z digest=sha256:86d0d0f11882a2ca932e20a890b342fda775d2c0d2f3da5338eda5ddb2764f5c

Observation fa99093f-eff0-48e5-8bd8-ad20e2821bd3 · outbound

This paper cites online" 'onlinestring :=.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge online" 'onlinestring :=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.459081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.459081Z digest=sha256:55efee4069f5f6a8d1a502a84f7c282e40891e2415ca2ec96e18d4a1057caa92

Observation 65b484b4-22af-470f-9f3e-ec5db528a1a3 · outbound

This paper cites write newline.

Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge write newline

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:47.526084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:47.526084Z digest=sha256:0f8986d73d81e875edaf5000bb99201847943a5a90e59971969b4f820e08e24e

Pith citing papers

Observation b2697847-5dea-482a-82da-5a066fb655a6 · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T11:49:31.595873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T11:49:31.595873Z digest=sha256:27165bcdad9441099857fb9f91a09ba6363d2a4e2b57914a05bb26e7467d5e9d

Observation 3d01d9e4-23ce-496d-8e1e-f52f7e92bc81 · inbound

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging cites this paper.

Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging Improving Automatic Evaluation of Large Language Models (LLMs) in Biomedical Relation Extraction via LLMs-as-the-Judge

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T07:20:33.617934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:20:33.617934Z digest=sha256:9f7b823df033f5c44faa7cef9a6ba6e9b7c916c2558f113b3cfb51ec5a46d5b1