Pith. sign in

Paper Citation Record · LEDGER

The Science of Evaluating Foundation Models

As of 9 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 3 inbound Pith citation observations for arXiv:2502.09670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09670 v1

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:35:42.932998Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:59:00.294806Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T13:35:26.411399Z

Reference resolution

100 of 109 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved92
  • parse uncertain0
  • malformed identifier4
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 42c9f172-4f43-45bb-a54b-f79576fea6e9 · outbound

This paper cites GPT-4 Technical Report.

The Science of Evaluating Foundation Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.614892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.614892Z digest=sha256:fcc1f08a5daa49ddbe8438ea8149dfa279e1ea5fcd42e5bc428f9a93a260f190

Observation 7a60ccf5-011b-40cb-a6bf-a6017ed74b60 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.619032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.619032Z digest=sha256:5cfb9804d76ec9de74f7dac27c53c306a7e163cfd977cc1f2f6d68a239546fe8

Observation 6f7afbaa-a360-4dd2-ad22-312df632e1e6 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

The Science of Evaluating Foundation Models AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.621868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.621868Z digest=sha256:e3d640bc05affc4c1ba2472f62a1f1cb08ae845cba165befdc09b978f931d2fd

Observation 8f75078f-c123-401a-a525-63783e7c7e83 · outbound

This paper cites Qwen Technical Report.

The Science of Evaluating Foundation Models Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.625010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.625010Z digest=sha256:11259e3a271f8bdf7287dc241276aa70d95cf3ac79de55908232ef880dda6e08

Observation 1e7b9739-cbf3-4c4b-9f8b-764e250d0aae · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

The Science of Evaluating Foundation Models MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.628193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.628193Z digest=sha256:ad6314f3e72cf1ace9d05eaa2f16e1bfa6bfe3fd199af58feba0fec28dec66e6

Observation 54d987c0-95d4-4044-9d51-e53b33a6e710 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.631123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.631123Z digest=sha256:8e1079493951f5d49291ca78f0ebad8e898039769ae9a8bce2d2dda30a0c6b40

Observation 41511bae-4253-4804-b53a-48a73a38451c · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.634066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.634066Z digest=sha256:dc5e58b88054e15d8430c64a3ac5459680642676b900aacd6fb31fe7c827c7ee

Observation b85ccb95-4a73-48d2-b1e6-a9281c26e127 · outbound

This paper cites A large annotated corpus for learning natural language inference.

The Science of Evaluating Foundation Models A large annotated corpus for learning natural language inference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.637213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.637213Z digest=sha256:8479f7d651dfd489db9708c8124344af8692df9ecb4852422a60c0110117cfd7

Observation 5967fedb-279c-4f5d-b6e8-245121c56191 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.640482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.640482Z digest=sha256:0ccdf921c058d824f7c6f4f1bf14cc09c2d4eb7199dddae22c7092f991acf078

Observation b9a4d036-ebc9-4ecd-9edd-e3393beb80e4 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.643255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.643255Z digest=sha256:b893d3b1d5eb24d4e5a9c77872224be7e58648d911316f8ad9ff8f33885616f8

Observation 29df79b0-d079-4839-bd57-ba240b292bfa · outbound

This paper cites Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations.

The Science of Evaluating Foundation Models Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.646158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.646158Z digest=sha256:0abb0b8994ad7c4d70a4ab33d3f71158145cfafa56df76443e83bedc68000939

Observation 7e87e0c8-0110-4424-b767-beb1980c8365 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

The Science of Evaluating Foundation Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.649410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.649410Z digest=sha256:4045dd46aab1475d3c33a1e8ee8155585640fe9da74149db056a6ecac9767dd9

Observation 8ab6763d-8366-497c-8cf7-6fd8c8c54b69 · outbound

This paper cites Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text Classification.

The Science of Evaluating Foundation Models Improving the Faithfulness of Attention-based Explanations with Task-specific Information for Text Classification

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.652396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.652396Z digest=sha256:5b09993abeeb6f81e3db095c0b77bcd1a47b24dabdebef004824f55734c50237

Observation fe5d6907-4908-4224-bf26-be7280651b14 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.655458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.655458Z digest=sha256:7c68dd52f11dfaa3b7404daaae941d9f771a8f37cb39e6aacfedb7ccb1a02a9d

Observation 1596f007-4aa1-4066-9a88-8c96bc8e9316 · outbound

This paper cites RobustBench: a standardized adversarial robustness benchmark.

The Science of Evaluating Foundation Models RobustBench: a standardized adversarial robustness benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.658408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.658408Z digest=sha256:d9d15ce5275460c6179bd761f4d139674e55e60d224ace51de447b6c211835c1

Observation 80be2df7-f961-48de-a6bf-b87100cf0ffb · outbound

This paper cites Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems.

The Science of Evaluating Foundation Models Risk Taxonomy, Mitigation, and Assessment Benchmarks of Large Language Model Systems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.661836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.661836Z digest=sha256:a2408a23df8eb27d5b8bd4bddf436dd43a5ef20e59b697b2fe4810d1e9f40f4f

Observation 3537a145-0e41-4fe3-80e4-182d8abd2937 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.664945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.664945Z digest=sha256:73e5b9de4e6ee82bb7101b508cfe439601a71ed7f6dd0530cba2b67004fa3521

Observation 45b93dc5-4227-475a-9f92-ee852864aeb3 · outbound

This paper cites ERASER: A Benchmark to Evaluate Rationalized NLP Models.

The Science of Evaluating Foundation Models ERASER: A Benchmark to Evaluate Rationalized NLP Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.667856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.667856Z digest=sha256:188fd3947fceaf18049c51641420830612e9ab2181b353215c49c44936a8ed0e

Observation e5204879-ae56-4307-a2e4-b57e582ad2d3 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.671333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.671333Z digest=sha256:451482d2dd9a9d937c4068c5d47dac3f60a953d7fcc67821fb7286c253866197

Observation 5a6fbcf9-4b9d-4bf2-856e-f4cb37c65970 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.674183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.674183Z digest=sha256:7b6949d11598bd53127df6b295795d35cf58fad68cefdaac515e8e74da72662e

Observation 706442f5-54ec-4af2-82af-ac355f421310 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.677278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.677278Z digest=sha256:ba425b80a3d34d517e1a950dafb403de657a2778cd9b662a3345a00b53813e8c

Observation b512d5ff-cb96-4239-a582-af833ebb7998 · outbound

This paper cites An Intersectional Definition of Fairness.

The Science of Evaluating Foundation Models An Intersectional Definition of Fairness

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.680314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.680314Z digest=sha256:f89e14e0ca27197b90fd0ad5b97fdf9f1e074a60a326598ef5004c516e8f6933

Observation f66b685a-41b0-454b-b26f-954ce8367b2e · outbound

This paper cites Selective Classification for Deep Neural Networks.

The Science of Evaluating Foundation Models Selective Classification for Deep Neural Networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.683229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.683229Z digest=sha256:66474d95a2afc0797d75f7ef0cfaabb30b4d1f7d3d76a94c63b544e54df76995

Observation 90609972-6f82-472e-8db8-3423a1fe6404 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 24

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.686588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.686588Z digest=sha256:2078baefaab80d29a314f7d3a869874fd73251602b1c793d31eb795052160210

Observation 57fca449-93b7-4e35-9ca3-947cb0e00a98 · outbound

This paper cites On Calibration of Modern Neural Networks.

The Science of Evaluating Foundation Models On Calibration of Modern Neural Networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.689465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.689465Z digest=sha256:a0210bd365c86fb9c3068a55ca61f10760046705785e65b551ec0d676bf7a969

Observation ed300267-aa21-4039-8308-acb5efe10507 · outbound

This paper cites Large Language Model based Multi-Agents: A Survey of Progress and Challenges.

The Science of Evaluating Foundation Models Large Language Model based Multi-Agents: A Survey of Progress and Challenges

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.692496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.692496Z digest=sha256:f1879d9e3454b2d0b3ccaf9e444b98b6602cbf9d4fc82a36f0786cb713863bf5

Observation c49507bc-c841-4861-a8c9-d433c0f1f365 · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

The Science of Evaluating Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.695574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.695574Z digest=sha256:f2fffc478ae33f575c5efa7b0716d6624e81b560e70e9217b9d71467628c6ad5

Observation 9ae3ded5-0dbb-4794-accb-8e714e4015cf · outbound

This paper cites The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models.

The Science of Evaluating Foundation Models The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.698050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.698050Z digest=sha256:1ecfdcb277aa092495f819b97fae542d72729df58090acfd6bfdb0ad302a2362

Observation 12b41117-de41-4783-9142-bb499db5a010 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.700573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.700573Z digest=sha256:30b5229195573bb05e273d6d72d914ca885a5faf230909babdb1da00f75cfb48

Observation 0a328410-f232-4698-99d7-280aaf548897 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.702763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.702763Z digest=sha256:9b37c7fb3d6238364f04c10343b95b9f558663423c3e54a567316062806efb4f

Observation 8f579ee1-1005-4ee9-8589-065189ae3929 · outbound

This paper cites Mistral 7B.

The Science of Evaluating Foundation Models Mistral 7B

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.704954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.704954Z digest=sha256:055ea8d8fe1bbae45a48b5eaea9ffacccd05c8ea0be902f8dabe30fc268e707a

Observation d995213e-6350-442b-842a-42a32a09be59 · outbound

This paper cites Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment.

The Science of Evaluating Foundation Models Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.707286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.707286Z digest=sha256:cf1aab7a1af4d3997b05112a01d4cef54f8d3b2eed5b09519fc9f1440950df9f

Observation 53901a15-baae-434b-a914-f309e49e8e35 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.709720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.709720Z digest=sha256:9db9ca9de397e5becdae7e52b3b94d617b3f73cb62fa6e9c8466093f33154614

Observation 972216ea-dae4-4d33-b25f-1cb12a26cb26 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 34

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.715639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.715639Z digest=sha256:4cb092889c989276911621e8d48c91d52a9b329ec6d96b5aad097fb91b3a3db9

Observation 552ef713-bd7d-483a-9c03-a35e71a76be5 · outbound

This paper cites WILDS: A Benchmark of in-the-Wild Distribution Shifts.

The Science of Evaluating Foundation Models WILDS: A Benchmark of in-the-Wild Distribution Shifts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.718445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.718445Z digest=sha256:cec09261ffc6ad9b454a7c489dfb57a37c149e33cadd31e85f129ea5a807d9a6

Observation cad99ac7-c79a-42bf-90e9-c92e65216222 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.721987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.721987Z digest=sha256:c0efb4269fb949a21172f9c64d815e7fc12a908aeb2b6fbf7316ea01d2c402ad

Observation a59ec1e8-3887-4873-9d5f-81906eb6e7b0 · outbound

This paper cites Dai, Jakob Uszkor- eit, Quoc Le, and Slav Petrov.

The Science of Evaluating Foundation Models Dai, Jakob Uszkor- eit, Quoc Le, and Slav Petrov

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.724900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.724900Z digest=sha256:befd32fdd9fac5ff02596c9fa0ca094e7ede205988acab107bfc7168514d40db

Observation 075f5f71-150d-420a-a730-4a4277b33180 · outbound

This paper cites Measuring Faithfulness in Chain-of-Thought Reasoning.

The Science of Evaluating Foundation Models Measuring Faithfulness in Chain-of-Thought Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.727917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.727917Z digest=sha256:ee21ee7c691390b2fdd558e56312023a3d3cab24791f8363f9f5f488a74b2400

Observation a619e8b3-ff6f-4ee2-8d17-20b15e5522da · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.731195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.731195Z digest=sha256:9459b7a0f4e14f1a1911862a0256cffd4209eac0069a0e56cb53425b689d3da5

Observation 65441e79-c127-4d2f-b3c1-c41653c992cc · outbound

This paper cites Evaluating Human-Language Model Interaction.

The Science of Evaluating Foundation Models Evaluating Human-Language Model Interaction

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.734153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.734153Z digest=sha256:c1296a5eb95a671c16dff5cf2c369118db2effa1aa93923adc609020e418323d

Observation 6be26943-f8c2-421c-884b-dba5cd61f4c9 · outbound

This paper cites Can Large Language Models Capture Dissenting Human Voices?.

The Science of Evaluating Foundation Models Can Large Language Models Capture Dissenting Human Voices?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.737130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.737130Z digest=sha256:5223f1b7cf14a722f6211b54bba1aa262cf4c162a47096fa6549427bf6e9a9bd

Observation 90b11f8a-52fb-4b25-9266-8b7ae0f47d75 · outbound

This paper cites HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models.

The Science of Evaluating Foundation Models HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.740185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.740185Z digest=sha256:506aecde731249bb7665e7330687dcf75998323b59d3024584abb043f5332b87

Observation 2a33283f-1f8c-4b73-8fa4-c8bc1bec5053 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.743247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.743247Z digest=sha256:b3ad15da354bfae6459afb1127c4ad13e154a641eb04017e1fe4b4cb83d54941

Observation 88eb5930-81a1-40fd-bdd6-eebf5bcb7c84 · outbound

This paper cites Holistic Evaluation of Language Models.

The Science of Evaluating Foundation Models Holistic Evaluation of Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.752311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.752311Z digest=sha256:82756ac5b1dcecc5bc33a7b1c0e0aa8c7e7b3c990bf11d3b78a70bb49b187a36

Observation f4539750-84aa-460b-904f-863ae408ad74 · outbound

This paper cites AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents.

The Science of Evaluating Foundation Models AntEval: Evaluation of Social Interaction Competencies in LLM-Driven Agents

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.385959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:35:42.755226Z digest=sha256:f14e29b2460e48fe3ee27af61117491ef8c0acd086f614691097e4be1ddf2c8e

Observation f9d79ee3-df9a-4330-8066-1f8c920c6163 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.758625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.758625Z digest=sha256:65c0fee42703468d7397694caf3b6dd8995c4e3fc59b792e09710c9eee051e27

Observation 6338a6ff-677a-4838-86d6-58c4a007de11 · outbound

This paper cites Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models.

The Science of Evaluating Foundation Models Error Analysis Prompting Enables Human-Like Translation Evaluation in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.761353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.761353Z digest=sha256:145e485cd5edac8b54b9df7de94ae8a28b30d39a035b9e07cf4981231bc1d1a9

Observation 358695cd-1638-4b11-81ea-3e5ca8383f1b · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.764319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.764319Z digest=sha256:1370dc755d5250a46b233b33a63b5fb43563549c4367e4566f24efa94c654b02

Observation 3620ce70-def8-4509-b833-ea60b5484b25 · outbound

This paper cites Social Bias Probing: Fairness Benchmarking for Language Models.

The Science of Evaluating Foundation Models Social Bias Probing: Fairness Benchmarking for Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.767138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.767138Z digest=sha256:897abf408ec00693f363a44212f680a889aa588cb432a0f26cab5b205f249e79

Observation 9ad605d3-0278-4cc8-9617-9aee498174d5 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.770006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.770006Z digest=sha256:a65b41b32620fb66dd8067b34a961b0688e52c20270280b23ce5604e3796267d

Observation c116ac70-da36-4744-993d-1caae578a27c · outbound

This paper cites StereoSet: Measuring stereotypical bias in pretrained language models.

The Science of Evaluating Foundation Models StereoSet: Measuring stereotypical bias in pretrained language models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.774783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.774783Z digest=sha256:d72c828fc26d382718f33f5d55f05d6c946e8dce72281842ee50bf33f77d1433

Observation e110ec71-f37a-4eee-bcae-08209c44724d · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.777377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.777377Z digest=sha256:63ef1822f8bb1839d18f08900d8d5d17e25e1896648e62ab9200c9c16e79f43a

Observation d9a87917-5e9f-4fd5-9f10-1e37eab5d1a8 · outbound

This paper cites Pointer Sentinel Mixture Models.

The Science of Evaluating Foundation Models Pointer Sentinel Mixture Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.772371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.772371Z digest=sha256:ee8ab7ec96a7f1dfc705da3ede04a8d5643c8c547ff245b3f4a043691a19775f

Observation b030956a-0cfb-4015-96ce-8a20373fa5e8 · outbound

This paper cites Cohen, and Mirella Lapata.

The Science of Evaluating Foundation Models Cohen, and Mirella Lapata

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.787601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.787601Z digest=sha256:2cf66a43ee40c864f557bf1d215cad2b26207ff7f8c3482517b789cefbfb6b1a

Observation 1f715672-1eff-412e-8e4c-c29850f55252 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.790559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.790559Z digest=sha256:856adbbdb616401a54ec66ae7f71f40d7f13c2ef978e33205fc4125cce76500d

Observation fe1fa383-b0af-4bd7-befa-bbc665696e72 · outbound

This paper cites Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond.

The Science of Evaluating Foundation Models Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.779786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.779786Z digest=sha256:7bbb795ee57398340d954d07f39624f5e8052050b0f4da59124aea5067d96377

Observation 81653e57-c00c-4a4a-9b16-663d87670804 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.782282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.782282Z digest=sha256:eb312c8f0b5e6a5d22f0953a3d336765ca84caa29e982fc04aeb205a460f51ad

Observation 90319e29-dc75-449c-87b7-3c6aa7c23604 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.802455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.802455Z digest=sha256:9005cb6b33d167f8f6d537a0f48b4a8cf7ee62f774df5adc1b743670bccd557e

Observation 98b069ff-6c98-47d4-8c32-b0b209f1ef16 · outbound

This paper cites Is ChatGPT a General-Purpose Natural Language Processing Task Solver?.

The Science of Evaluating Foundation Models Is ChatGPT a General-Purpose Natural Language Processing Task Solver?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.805153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.805153Z digest=sha256:c7662d8cbaea3ae1c19ee23036ee34a2d81e97bef70cbd60030c6d6efd9e6e74

Observation bb7c04e8-9804-4129-8c8b-a96455d5984a · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.808366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.808366Z digest=sha256:5d380a6a7bdcb757c6628d02f92673e096cc655e225ede4be74b8f8c5d7d17cc

Observation 2d65a7e5-c86c-47ff-9d8a-40cf58e7980a · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.793354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.793354Z digest=sha256:cc9639cb5cdf1e5e28d56db613d04f8930889b546ec030e67ecbad5f9a326ed6

Observation 23278487-8ae8-4dd8-b04e-786112a023a7 · outbound

This paper cites Question Decomposition Improves the Faithfulness of Model-Generated Reasoning.

The Science of Evaluating Foundation Models Question Decomposition Improves the Faithfulness of Model-Generated Reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.817579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.817579Z digest=sha256:3941225879d67cd53a57eee8c05f415c535ac19a14232a948d2e6f9d4cbbdca3

Observation b750a0c5-1e45-4cf6-bb69-966979345786 · outbound

This paper cites A Survey of Useful LLM Evaluation.

The Science of Evaluating Foundation Models A Survey of Useful LLM Evaluation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.799324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.799324Z digest=sha256:7ee9f0ea6742a127f975b0791828c3994e2198a079ae22b2f68a50e8cc7f9413

Observation 2c0b10e6-86cd-480f-9519-3439a79e9530 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.823750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.823750Z digest=sha256:497550c12199fabc24a12373bbd0d87a6a80803e77717b2d1b1ee069c5ebd7fc

Observation a1f5741e-f9c0-45b3-a459-22c647d2a224 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 66

Resolution
malformed identifier
no resolver link, observed 2026-08-07T23:35:42.829679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.829679Z digest=sha256:47fd85d9938b4dcf8904c91aec2e1e01601636e3ba4ecafafddbfccbdbd75e99

Observation dec6e7cb-b5c9-48f1-8070-dfe70d40a234 · outbound

This paper cites Rush, Sumit Chopra, and Jason Weston.

The Science of Evaluating Foundation Models Rush, Sumit Chopra, and Jason Weston

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.832607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.832607Z digest=sha256:1d53e55d167e9fc0f932083e755f779bb36cd4339b96718834031002d4171ac3

Observation 13445a94-7ada-472e-be00-90bbbe398e95 · outbound

This paper cites Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition.

The Science of Evaluating Foundation Models Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.835466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.835466Z digest=sha256:313e7673fac9960b75c81a4a451afffdb2314fedbf9a29d4029575a54a0f0a6f

Observation 01e16b8e-c3f9-4db7-b028-afd108b7d7fe · outbound

This paper cites LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models.

The Science of Evaluating Foundation Models LongHalQA: Long-Context Hallucination Evaluation for MultiModal Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.814464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.814464Z digest=sha256:0b5ff1d468664f4b9749968b624d917ec3651bf336367d286a2b23fcd4bac418

Observation 5f764ef2-c647-432d-b8c1-e4e3a3f9e7d0 · outbound

This paper cites An Interpretability Evaluation Benchmark for Pre-trained Language Models.

The Science of Evaluating Foundation Models An Interpretability Evaluation Benchmark for Pre-trained Language Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.259889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:35:42.841478Z digest=sha256:e7f9f6274c0bcadce53fd5fa46ece9d0c72e3dec39dc978afa64128bfefb8e05

Observation 5e1c44a6-b37c-4e66-a242-719aa6949594 · outbound

This paper cites SQuAD: 100,000+ Questions for Machine Comprehension of Text.

The Science of Evaluating Foundation Models SQuAD: 100,000+ Questions for Machine Comprehension of Text

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.820557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.820557Z digest=sha256:4e18b654d4b1ef56065c2cd5ccbb0faa9d58c20764ee3d5e8fe423bb3c26a976

Observation 781e5b33-a901-4679-8be3-d59c8d9d1e48 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.847446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.847446Z digest=sha256:b50435fa4dd4cfccbd7897762b9270defc82cef8396118a6fb367affa2fefbd8

Observation 4bd856cc-13c9-4883-9f6b-0f6a181f49bb · outbound

This paper cites In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Car- reras (Eds.).

The Science of Evaluating Foundation Models In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing , Jian Su, Kevin Duh, and Xavier Car- reras (Eds.)

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.826508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.826508Z digest=sha256:4f5494893f5056a4ac567bda6eb49feb176289030f7036016e50b0f4871c9767

Observation 3a20dc31-a722-4aef-8497-83847d01c642 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.852413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.852413Z digest=sha256:824afbdba1a24d89149501951ede3060f00c3e675253512b43215a605b3acd1d

Observation 70765af2-18bd-473d-ab69-26c73a29e4ad · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.854802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.854802Z digest=sha256:1337808b0891adfbeccdecfe0845389289d992a6d6e2c6b5af9a9980627abdfc

Observation d7c248df-e0b5-46e8-a888-d41c432aec1b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

The Science of Evaluating Foundation Models LLaMA: Open and Efficient Foundation Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.856940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.856940Z digest=sha256:8d560a9cb38710a39b7edccb73872797ebc95ce2f2d11a5611d1759c2c7d4c77

Observation eb1aeec2-bb7d-4074-a21e-c13535ff7ecc · outbound

This paper cites Evaluating Large Language Models with fmeval.

The Science of Evaluating Foundation Models Evaluating Large Language Models with fmeval

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.272176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:35:42.838370Z digest=sha256:9cf67f193b75fd3e96926c2f85b12b54e79346b2480b549d28c6d1b910e3826e

Observation e5469b24-71e6-4831-9b20-84825d9d63c7 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

The Science of Evaluating Foundation Models GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.863898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.863898Z digest=sha256:0e4fd84e2d37a4d99cbe4b4f1bc3c467cb575b4285b49fb02a348341bb98409b

Observation 6bd40893-61f8-4de3-ae64-01645250aaba · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.844395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.844395Z digest=sha256:c9b5820ad182dba0bf5cd968f043f424a65238dbd8990c83202fdcc477aa41af

Observation 0955883b-c37d-4f8a-a42f-19e68cc01e40 · outbound

This paper cites DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models.

The Science of Evaluating Foundation Models DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.869827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.869827Z digest=sha256:63eac86b705a01217d7d4e7cf84fcec071dab676b7e3a7a9c26060b0c59d171b

Observation 9dab2ad0-c7f6-4b39-8ac3-479ce16c67f5 · outbound

This paper cites LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models.

The Science of Evaluating Foundation Models LalaEval: A Holistic Human Evaluation Framework for Domain-Specific Large Language Models

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:35:43.246503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:35:42.849901Z digest=sha256:8045b521af8d90434f5bd308e8f15d71677498d49fe7fdb2e590f96d40f42481

Observation 3e92f87e-2686-455f-83f1-4bca005cd247 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

The Science of Evaluating Foundation Models AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.876049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.876049Z digest=sha256:4698d6e88139a27af756ae651d6ee9a12d1006d864690c8048a0a09c889b7648

Observation b31ecd82-ff6f-4bce-8643-30eec429bcf5 · outbound

This paper cites Document-Level Machine Translation with Large Language Models.

The Science of Evaluating Foundation Models Document-Level Machine Translation with Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.879026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.879026Z digest=sha256:a19d726380d96eeb68b1e22f2dddafbec25a49503d03ff0768ab3cbdd20c7a3e

Observation 58fd2708-310c-406b-a025-44a2c6c04a52 · outbound

This paper cites Smith, and Teruko Mitamura.

The Science of Evaluating Foundation Models Smith, and Teruko Mitamura

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.882205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.882205Z digest=sha256:21bdc24bcb770107ca3eef769cf2e2d2d14129a42b474cfa1b91e3d65a0aba8d

Observation 5df30076-7d26-4ce6-8d0c-9d2f199cec50 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.859441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.859441Z digest=sha256:cc6ba1ba9da9d759f65068fa5bd0228594dca2144fa09c1865d3900c0df31878

Observation 6ce5459a-8839-4d48-9da7-b3cf16f2b55f · outbound

This paper cites Advances in Neural Information Processing Systems 36 (2024).

The Science of Evaluating Foundation Models Advances in Neural Information Processing Systems 36 (2024)

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.861678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.861678Z digest=sha256:b8360d78fe09fb0f70d73cd73b75555a46a4f29e0e25a7576306769f655a6ff1

Observation ba2eda73-8fe5-4972-a8bd-8add8a617f16 · outbound

This paper cites DHP Benchmark: Are LLMs Good NLG Evaluators?.

The Science of Evaluating Foundation Models DHP Benchmark: Are LLMs Good NLG Evaluators?

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.891112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.891112Z digest=sha256:6439bbbf1977bd6369a0a4a2d2466b77b269bf30a8293363deeacb8eacf1e035

Observation 43dc27af-db77-471d-9052-f793109c05f7 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.867035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.867035Z digest=sha256:a2a4328d475296be95171f60bd9cbc0abf739805d63a6ae52662713985d70d59

Observation 6fd23ec5-6389-4d9a-92ee-cd3f27240be0 · outbound

This paper cites A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference.

The Science of Evaluating Foundation Models A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.897234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.897234Z digest=sha256:01747f684c56151e30989492789bfb82fa745ce036e52a78697bb0cd0731886e

Observation 51a6cad7-7eab-4dbe-9466-5731b9e54fb6 · outbound

This paper cites Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models.

The Science of Evaluating Foundation Models Adversarial GLUE: A Multi-Task Benchmark for Robustness Evaluation of Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.872968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.872968Z digest=sha256:392dd752cab937d32a64f3b3c251dd6ef7ce75f7d0b0811a12510852418120a2

Observation 6b7b782d-e3d9-49aa-8842-b8816314797e · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.598649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:35:42.903121Z digest=sha256:b21fe2513c5884f8b4d031bc740984dfa1433eca8bbb03e0743f3707a0cb1dfc

Observation 10a31815-2774-45dd-9871-f3e508c4e89f · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.589935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:35:42.909203Z digest=sha256:a1830b0f8dd2d289ac5283312112963384bb6596172ee3ef29903fcf62642d62

Observation 2639926f-7a1a-44a4-bd2a-1327389db7fa · outbound

This paper cites R-Judge: Benchmarking Safety Risk Awareness for LLM Agents.

The Science of Evaluating Foundation Models R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.912130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.912130Z digest=sha256:44f9e46ceaf61bd1b8510bb8e4d21a8fa761ba1fe25c1417ba80cccdd39bb1be

Observation a40c18e0-4c67-4b32-bd52-1fb541fbcbcf · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.885080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.885080Z digest=sha256:c7bf55b3043c16ab751605472abb365086f859a66c441fee11d0a6cf20a1b88e

Observation ffce3c16-92b4-40d9-a5a9-e9114ee65dc5 · outbound

This paper cites A Theoretical Analysis of NDCG Type Ranking Measures.

The Science of Evaluating Foundation Models A Theoretical Analysis of NDCG Type Ranking Measures

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.887915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.887915Z digest=sha256:47254bc66da5aef6ea9a36aadc1bc8d45dfad2657ae234be7faae4a450705822

Observation 8a072e5a-e6ea-4d15-8dcd-93e24a914f35 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 96

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.581098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:35:42.922646Z digest=sha256:35c06534a651b148e3f47ad8a3fb132911e676699a2ef83f80b0539cb61f1c2c

Observation 9ea7a298-1b5d-4e66-81fa-4c2d9a97de88 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

The Science of Evaluating Foundation Models Aligning Large Language Models with Human: A Survey

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.894068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.894068Z digest=sha256:61ed7777117991b6a4bc89b75909acdc92441b1b8f6bb9ecdacb55e101e542a0

Observation defb6ea1-6aed-4bd9-b8e7-c2bba422d85b · outbound

This paper cites PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts.

The Science of Evaluating Foundation Models PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.927960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.927960Z digest=sha256:6aaf7b1930174847521cd243f52129ee025d5d8681bb1bc319324f64324dee9e

Observation 3a059a11-00b6-45aa-9180-3e42089465e4 · outbound

This paper cites an unresolved cited work.

The Science of Evaluating Foundation Models Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:35:43.606460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T23:35:42.900246Z digest=sha256:dcdc5bd3ee1818d8515e1846fbdbaabc4f70dbce979ee533138f234788b147cd

Observation 4c8835d8-bdec-40c5-981d-85d628bf4a3b · outbound

This paper cites Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books.

The Science of Evaluating Foundation Models Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.932998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.932998Z digest=sha256:f8f0e15f9eda75680556269c8bcc73201bd94df43f44b959db6b85d9fc53bce4

Observation 32c74742-06d8-4d11-b2cf-0bad6dd002e1 · outbound

This paper cites LLM as a System Service on Mobile Devices.

The Science of Evaluating Foundation Models LLM as a System Service on Mobile Devices

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T23:35:42.905965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:35:42.905965Z digest=sha256:b13d5d6d62831f7c368b77bf8c5b7dd816311c9430bbbeea0992c033de97cd13

Pith citing papers

Observation 2fabdbc0-8037-4596-bd31-d827de9ce9cf · inbound

Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch cites this paper.

Deterministic Inference across Tensor Parallel Sizes That Eliminates Training-Inference Mismatch The Science of Evaluating Foundation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T20:59:00.294806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:59:00.294806Z digest=sha256:9cd01ac861af2a15e71a312ca5090ee0c10434129a1c7d023c15f9e5db5a07ed

Observation df984a59-0850-4fc8-88b0-ca88a2f458de · inbound

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models cites this paper.

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models The Science of Evaluating Foundation Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.412843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:31:46.940449Z digest=sha256:7f154a9eb1ab3e29838f2fe46c7b678eb78a37fb251407b21fbbfa54326e9a46

Observation 817c9b0d-9628-423d-9255-866ac54a79ff · inbound

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis cites this paper.

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis The Science of Evaluating Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T20:50:53.567415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:50:53.567415Z digest=sha256:ad9b205c3bb68112d5210ecc6ac0c64d4b12bb7802bd0b710ddc4fac8d3749c9