Pith. sign in

Paper Citation Record · LEDGER

Statistical Multicriteria Evaluation of LLM-Generated Text

As of 20 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 1 inbound Pith citation observation for arXiv:2506.18082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18082 v2

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:01:23.023749Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T04:49:15.239636Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T04:52:17.135964Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact5
  • verified fuzzy39
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73934b5a-b8bd-4fad-9184-f14341d7d33b · outbound

This paper cites GPT-4 Technical Report.

Statistical Multicriteria Evaluation of LLM-Generated Text GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.612704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.612704Z digest=sha256:e62e35f03dd8e085bf18fabb1c5a154ba68c1c82b1968f7b4640272871439f81

Observation 047cc169-9748-4ac8-8dd0-afd73b65e784 · outbound

This paper cites A learning algorithm for boltzmann machines.

Statistical Multicriteria Evaluation of LLM-Generated Text A learning algorithm for boltzmann machines

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.618735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.618735Z digest=sha256:417158fa216ec6f8ad962f58705967c871c1aeb6add06e0488245d518d7072d0

Observation 3d5a76c7-339b-4488-b6d2-3deb50423093 · outbound

This paper cites Bayesian Optimization for Building Social-Influence-Free Consensus.

Statistical Multicriteria Evaluation of LLM-Generated Text Bayesian Optimization for Building Social-Influence-Free Consensus

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.623798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.623798Z digest=sha256:1b8aa99d5c0bbaf6790f6837e2ad838a3e0c1c609a12521a54440a2bff6c0796

Observation bd10c74c-bbdc-4b18-9328-500dd98a5989 · outbound

This paper cites Jointly measuring diversity and quality in text generation models.

Statistical Multicriteria Evaluation of LLM-Generated Text Jointly measuring diversity and quality in text generation models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.629098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.629098Z digest=sha256:cdf03c77d039d554b18dc3dee26eac5d5839cd7cb8cbf35ef8b51b1ad93388af

Observation 6c93fbea-aa3e-477e-b6c0-6075fae97fbc · outbound

This paper cites Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges.

Statistical Multicriteria Evaluation of LLM-Generated Text Text Generation: A Systematic Literature Review of Tasks, Evaluation, and Challenges

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.634347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.634347Z digest=sha256:eb4a8b4c5175647048230c2e5f303409fd8ec1f61964f7f024af29640254920e

Observation 6ab273dd-3369-4e8b-8c2c-519065bcb098 · outbound

This paper cites Howcroft.

Statistical Multicriteria Evaluation of LLM-Generated Text Howcroft

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.639624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.639624Z digest=sha256:e44d5c73baa36ecc643ae9b4afd51bbebd431d76f18be97f62bfad4b7935e591

Observation d80f3a6b-10a5-4745-abef-2f56854447b3 · outbound

This paper cites Time for a change: a tutorial for comparing multiple classifiers through bayesian analysis.

Statistical Multicriteria Evaluation of LLM-Generated Text Time for a change: a tutorial for comparing multiple classifiers through bayesian analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.645123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.645123Z digest=sha256:1c1dead1809938220052368cb3c48b359db38bad19fd4b00576ecbe9cc4afc3a

Observation be63f9a0-102e-48ad-ad99-ea2ee7f1b679 · outbound

This paper cites Estimating the replication probability of significant classification benchmark experiments.

Statistical Multicriteria Evaluation of LLM-Generated Text Estimating the replication probability of significant classification benchmark experiments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.649563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.649563Z digest=sha256:2cf79dd902555bcbd90f1f9e87e75f25ca35cf63bd7a9ecb0d5e048734db3493

Observation 180392b0-bcbb-4614-b21d-5cc702d4a68f · outbound

This paper cites Comparing machine learning algorithms by union-free generic depth.

Statistical Multicriteria Evaluation of LLM-Generated Text Comparing machine learning algorithms by union-free generic depth

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.654122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.654122Z digest=sha256:768ddac42c241f96477151eb628e984f6357eda5a811089a20aeb9270595720a

Observation 7fd33ca0-421c-428a-b184-435ac2a8c915 · outbound

This paper cites How to Choose a Reinforcement-Learning Algorithm.

Statistical Multicriteria Evaluation of LLM-Generated Text How to Choose a Reinforcement-Learning Algorithm

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:01:23.498287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.659145Z digest=sha256:b63e80f76c9747f21271ac2faea64dcaec0b8745453c8e789cfe779ef8ec8e20

Observation 2f2a6f9b-d7d4-4e06-85b5-1e658d827352 · outbound

This paper cites Self-learning from pairwise credal labels.

Statistical Multicriteria Evaluation of LLM-Generated Text Self-learning from pairwise credal labels

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.664456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.664456Z digest=sha256:062183a88b9db811822bc052273a64935da0bda95327c8b81ac2b505ebb04bcd

Observation 06115b3f-35d0-41a9-92d9-a41461efc11c · outbound

This paper cites Credal Bayesian Deep Learning.

Statistical Multicriteria Evaluation of LLM-Generated Text Credal Bayesian Deep Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.668977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.668977Z digest=sha256:227097b050755548f095a8696dc20d09c0a3859f9dd4fc52732a255263363266

Observation a7374a2e-e032-4ad9-b368-1b027ee2837b · outbound

This paper cites Evaluation of Text Generation: A Survey.

Statistical Multicriteria Evaluation of LLM-Generated Text Evaluation of Text Generation: A Survey

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.674048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.674048Z digest=sha256:4787228423e80fbc6f991565c1bfdbaeb9997dd34f98353feb8b153e55291eba

Observation 37a75dca-7e3d-4efb-968e-8ccfe098d594 · outbound

This paper cites Evaluating language models as risk scores.

Statistical Multicriteria Evaluation of LLM-Generated Text Evaluating language models as risk scores

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.679008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.679008Z digest=sha256:4aca2abd92746cdb26385e4806af574f6940421e80add1039f15e1f055403a7c

Observation de2d369e-ba1e-41c7-83f6-96f86b9678c6 · outbound

This paper cites Dem s ar.

Statistical Multicriteria Evaluation of LLM-Generated Text Dem s ar

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.684349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.684349Z digest=sha256:4256ee8d704a5198337a155f4fe141303acb2f647929d2a2090b5314323b9702

Observation 48152cc6-be62-424b-90c2-7df5bd442a12 · outbound

This paper cites Semi-supervised learning guided by the generalized bayes rule under soft revision.

Statistical Multicriteria Evaluation of LLM-Generated Text Semi-supervised learning guided by the generalized bayes rule under soft revision

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.370372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.688853Z digest=sha256:01fdf498b193199cac229056f1561debf19ec7f7149273a1e97bb8d6a59b57de

Observation 50948c9c-2caf-48cc-89a4-cd79f0608d6f · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:01:24.355254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.693271Z digest=sha256:62c0c74de40435a39213880a43196b6f43cf4c2a3daf20ba25950ea65c4f67aa

Observation 5d176290-587d-4e28-8a29-41249b8a12a9 · outbound

This paper cites Eugster, T.

Statistical Multicriteria Evaluation of LLM-Generated Text Eugster, T

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.340072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.699076Z digest=sha256:66c6815b5350c315c277234a4fbbf2268d71ebe99c6eb7863c2ff14b6ecaebc8

Observation 0b2b8622-edac-4319-8f46-9928f3308b49 · outbound

This paper cites Hierarchical neural story generation, 2018.

Statistical Multicriteria Evaluation of LLM-Generated Text Hierarchical neural story generation, 2018

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.323904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.703565Z digest=sha256:3d167066b5c40dc0f29c8def943bf726824351d04efb471f538f84fba6b9ad40

Observation a470a5ca-8b39-47e1-9cfa-ce2230e1f604 · outbound

This paper cites Beam search strategies for neural machine translation.

Statistical Multicriteria Evaluation of LLM-Generated Text Beam search strategies for neural machine translation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.708202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.708202Z digest=sha256:cff97a25639a7fe4955291314566382149ae3920e35aa2e47dba99cf8fa51bc4

Observation 502d7b1d-5c40-4a6f-8b77-93039ee13f03 · outbound

This paper cites Simcse: Simple contrastive learning of sentence embeddings, 2022.

Statistical Multicriteria Evaluation of LLM-Generated Text Simcse: Simple contrastive learning of sentence embeddings, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.712984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.712984Z digest=sha256:bded2f848fa54902fcaadd20ff0c95d0fc58f3982edc8c3517747b2a711f1999

Observation c70d8371-79a8-4660-aed6-08e2de291b8b · outbound

This paper cites Adaptive Contrastive Search: Uncertainty-Guided Decoding for Open-Ended Text Generation.

Statistical Multicriteria Evaluation of LLM-Generated Text Adaptive Contrastive Search: Uncertainty-Guided Decoding for Open-Ended Text Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.717493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.717493Z digest=sha256:9d531741cecaffc941e867b9d9c05b63af0f904dd5641c366ecfe4ee54f0301c

Observation 809d74bb-aed3-46c0-b02c-1d76cb139846 · outbound

This paper cites Decoding decoded: Understanding hyperparameter effects in open-ended text generation.

Statistical Multicriteria Evaluation of LLM-Generated Text Decoding decoded: Understanding hyperparameter effects in open-ended text generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.295768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.722044Z digest=sha256:7ab91ab250a863def264238faaffe2b644add96e187e6b8ed87bf3c99cb44ba9

Observation 41b2dbed-66f3-4aca-8dca-3aeb8f0d9da1 · outbound

This paper cites Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework.

Statistical Multicriteria Evaluation of LLM-Generated Text Towards Better Open-Ended Text Generation: A Multicriteria Evaluation Framework

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:01:23.406919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.726512Z digest=sha256:f86306a7c881b94bf1901f1fa338d0d392adfe795c61a63d79f1e93936d54234

Observation 47b56d44-0696-4371-ae7c-de8d090e9441 · outbound

This paper cites Garc \'i a and F.

Statistical Multicriteria Evaluation of LLM-Generated Text Garc \'i a and F

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.280452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.731330Z digest=sha256:b1e6d74511835880cbf4170f4871cf959218c2fdf2f65015c151b4837353f954

Observation 98af811d-5f3d-42e5-8310-aefcfe993338 · outbound

This paper cites García, A.

Statistical Multicriteria Evaluation of LLM-Generated Text García, A

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.266075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.736083Z digest=sha256:6217bec4e684ab622fb5df7753d8cf4b762ff495c193cac2429f7cb2e28dcf11

Observation 847f3e74-df78-4e72-835d-c74db774bf9d · outbound

This paper cites The Llama 3 Herd of Models.

Statistical Multicriteria Evaluation of LLM-Generated Text The Llama 3 Herd of Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.740637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.740637Z digest=sha256:68d68c3243fbf39d3526ad6906332e00a5827c7b7bd0536f4116d6875d517cd9

Observation 62970aa5-b0fd-448e-87e5-f8fdbb56fd96 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Statistical Multicriteria Evaluation of LLM-Generated Text DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.745434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.745434Z digest=sha256:a17354413cf385af028d1d69b13ea3143533f37bac48960ac281ccfe70399eed

Observation f06a955f-23d7-49ff-9010-f1e8e2bd8cf0 · outbound

This paper cites Unifying Human and Statistical Evaluation for Natural Language Generation.

Statistical Multicriteria Evaluation of LLM-Generated Text Unifying Human and Statistical Evaluation for Natural Language Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.749990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.749990Z digest=sha256:037710b40f96de8bdfea4935d3578f8b16fd170616c1d77628149a924cf7fc4f

Observation d9a87aa2-064c-415d-845c-7fab3f7e4763 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Statistical Multicriteria Evaluation of LLM-Generated Text The Curious Case of Neural Text Degeneration

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.755057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.755057Z digest=sha256:c9766da3ba77c32a43c06cdb47ca21afe902b616d7ac16c044103c822cfc0a15

Observation 3eb4cdf9-2109-46e9-8d61-42e0b965f71d · outbound

This paper cites Hothorn, F.

Statistical Multicriteria Evaluation of LLM-Generated Text Hothorn, F

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.250826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.760019Z digest=sha256:f9c325cf395f94458c7615126a124d839ff838cf3121bf26e110060f10cd1ba8

Observation 482aa6f5-0bbe-4eb1-950c-6434c19dd408 · outbound

This paper cites Open graph benchmark: Datasets for machine learning on graphs.

Statistical Multicriteria Evaluation of LLM-Generated Text Open graph benchmark: Datasets for machine learning on graphs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.764669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.764669Z digest=sha256:98af48e131eb3e1dad5b16da09f1bbed91b44baf72f52487b4b0779b6ab94658

Observation d4c2c78a-1f2f-4702-83eb-305930651da5 · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:01:24.225154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.768921Z digest=sha256:68389d1b6ecb981d5b1a70798ee16c621539ecbe98ae5df5a4953cfca1d8c828

Observation d6245343-16f7-468b-be3d-58c25aa7a060 · outbound

This paper cites Jansen, G.

Statistical Multicriteria Evaluation of LLM-Generated Text Jansen, G

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.208171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.773521Z digest=sha256:b5857e954405fd2f403b948dd6faeb0ab3dfab03eac29f8fcfb5b8c00732442e

Observation 7f5a2312-0e66-4ddb-a436-6c2159c36e4e · outbound

This paper cites Jansen, H.

Statistical Multicriteria Evaluation of LLM-Generated Text Jansen, H

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.191725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.777980Z digest=sha256:88f5deb27cd4a1a3c74a343b7caa0d924128e6bc31d0b057f8a357de7a5e4d8c

Observation abde3ebb-d0fc-4b16-872b-7c4b2e5a2439 · outbound

This paper cites Contributions to the Decision Theoretic Foundations of Machine Learning and Robust Statistics under Weakly Structured Information.

Statistical Multicriteria Evaluation of LLM-Generated Text Contributions to the Decision Theoretic Foundations of Machine Learning and Robust Statistics under Weakly Structured Information

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.782509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.782509Z digest=sha256:cb93991d9f4261e5e45ce8715df423daec49045d098826ac997105615a15ac15

Observation 0d9394e2-1fc2-49d7-a269-ea2bb814f8ad · outbound

This paper cites Statistical comparisons of classifiers by generalized stochastic dominance.

Statistical Multicriteria Evaluation of LLM-Generated Text Statistical comparisons of classifiers by generalized stochastic dominance

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.175823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.787435Z digest=sha256:c5f7b29e8745cf692f833717127d425d8b8e87a45567485fb78ab904c313c77e

Observation ee6a52ff-a88f-4ff5-9a52-0ef5ac3231dc · outbound

This paper cites Multi-target decision making under conditions of severe uncertainty.

Statistical Multicriteria Evaluation of LLM-Generated Text Multi-target decision making under conditions of severe uncertainty

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.158971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.792086Z digest=sha256:221e03f53b185291db0c7ada53e048a2f8a7f30baec1fc528a13e2c066425dc8

Observation 9478c78a-6f67-461b-8abc-620cc3c3064c · outbound

This paper cites Robust statistical comparison of random variables with locally varying scale of measurement.

Statistical Multicriteria Evaluation of LLM-Generated Text Robust statistical comparison of random variables with locally varying scale of measurement

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.143662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.796599Z digest=sha256:70868e3bafa4520057b662014f56b1d23de9a45b818bb0d3e8acc5e1550bd013

Observation 2e7b3895-f8a9-44d8-9276-bcecca37765b · outbound

This paper cites Statistical multicriteria benchmarking via the GSD -front.

Statistical Multicriteria Evaluation of LLM-Generated Text Statistical multicriteria benchmarking via the GSD -front

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.126866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.801334Z digest=sha256:32ab66dfe47072fcc4cb3c561f816fc8139db2d72e1abf1fedd72f6c00bcd0cd

Observation 47d7e42a-b273-4d21-9b91-4c51e05154ad · outbound

This paper cites Jelinek, R.

Statistical Multicriteria Evaluation of LLM-Generated Text Jelinek, R

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.805751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.805751Z digest=sha256:2a02f4f779292825f4f53dade86adc1d9d29210cdaec02b0ae34ea4f14eb749d

Observation 405d4636-93bb-47fd-989a-03089d3a33a9 · outbound

This paper cites The prism alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models.

Statistical Multicriteria Evaluation of LLM-Generated Text The prism alignment dataset: What participatory, representative and individualised human feedback reveals about the subjective and multicultural alignment of large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.106718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.810810Z digest=sha256:3493e0554add981ff31473e77d61b2f701f7edd1a2ab18f252ef834f9cedc353

Observation 552b9996-ea13-499c-9877-1c36bc3abfd1 · outbound

This paper cites Efficient multi-criteria optimization on noisy machine learning problems.

Statistical Multicriteria Evaluation of LLM-Generated Text Efficient multi-criteria optimization on noisy machine learning problems

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.084317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.815635Z digest=sha256:b4aa462e8ff92403c194e7f88fb82e61e44097364cfc937dbcdd6e0281f067ff

Observation ad91afc2-cd04-4249-8be3-23b5cd9e65f1 · outbound

This paper cites Towards quantifying the effect of datasets for benchmarking: A look at tabular machine learning.

Statistical Multicriteria Evaluation of LLM-Generated Text Towards quantifying the effect of datasets for benchmarking: A look at tabular machine learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.068837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.820250Z digest=sha256:e52ffcac8ababdfa54022516173ea831d237b51af2bf819f14bebee43a917514

Observation 7e220059-a974-498a-bdc9-bd69afd6e9f9 · outbound

This paper cites Accelerated experimental design using a human-ai teaming framework.

Statistical Multicriteria Evaluation of LLM-Generated Text Accelerated experimental design using a human-ai teaming framework

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.051460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.825117Z digest=sha256:a30d5087369e44fba931d6c6994a8e789f52ffa8f1ecbe3e81d3475cb9067acc

Observation 8363a457-cb5f-41f4-8e89-a1f6dc33daab · outbound

This paper cites Factuality enhanced language models for open-ended text generation.

Statistical Multicriteria Evaluation of LLM-Generated Text Factuality enhanced language models for open-ended text generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.030793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.829565Z digest=sha256:4f04f1be6241b4cb9a19ffb1c577fb535cd9e647320c0fd646d2bbcd4bb34c0b

Observation 94544a59-e947-480a-8e71-1e8d656d5c9e · outbound

This paper cites A diversity-promoting objective function for neural conversation models.

Statistical Multicriteria Evaluation of LLM-Generated Text A diversity-promoting objective function for neural conversation models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:24.010099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.834159Z digest=sha256:763bac25bdf464a41f233d1078d05577e132cb4b8061d5e847efe94fc082d2cd

Observation 57e6a772-cc2b-407d-b354-7fa7a9464712 · outbound

This paper cites Contrastive decoding: Open-ended text generation as optimization, 2023.

Statistical Multicriteria Evaluation of LLM-Generated Text Contrastive decoding: Open-ended text generation as optimization, 2023

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.839032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.839032Z digest=sha256:243f43b97f8b1a4233928047d48c9dd18011278f8f0af1aba27b4683df717cce

Observation 875b34cc-8ac6-4733-9fe6-ab580adcc110 · outbound

This paper cites Quantifying Variance in Evaluation Benchmarks.

Statistical Multicriteria Evaluation of LLM-Generated Text Quantifying Variance in Evaluation Benchmarks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.843911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.843911Z digest=sha256:6cb22f1121f27cd210ba8d4d8cf3a5f3b2e8ced2da2d9c48124b2bb1e5ea0683

Observation 8993ba60-1552-4443-8894-7bb642e477e4 · outbound

This paper cites Pointer sentinel mixture models, 2016.

Statistical Multicriteria Evaluation of LLM-Generated Text Pointer sentinel mixture models, 2016

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.849573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.849573Z digest=sha256:72fa78280cec1837ca4443970c5c28c2d7f46b780eea2abb6593e5502c6b0c78

Observation cecee8ff-15b0-4016-a5c2-1b9d7fb8f91e · outbound

This paper cites Mersmann, M.

Statistical Multicriteria Evaluation of LLM-Generated Text Mersmann, M

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.967599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.854610Z digest=sha256:1143ea33727181e579e471a75302be12cc9a62dc405e90e78759972c9f654149

Observation f3b8f4af-bcea-4658-8d44-d31cd6c37a8b · outbound

This paper cites Meyer, F.

Statistical Multicriteria Evaluation of LLM-Generated Text Meyer, F

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.953009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.859368Z digest=sha256:9d7d9ad6e941ebb4922146d712de8d5e02fce9339def5975464ca9873eb8d35b

Observation 006789af-20db-4f85-9519-c68d78d8394c · outbound

This paper cites Learning de-biased regression trees and forests from complex samples.

Statistical Multicriteria Evaluation of LLM-Generated Text Learning de-biased regression trees and forests from complex samples

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.937613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.863706Z digest=sha256:48a1e485460ebbe1f9861fb1c565e882ebeffc7adcb60ea6ca585a8b48ced6a2

Observation d32f776c-4f0f-494a-b456-b7912e58b482 · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:01:23.921306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.868285Z digest=sha256:716227cc43b5cb968171a517280be86af6307bed1a8fa111afb12538e5fdf430

Observation 5fd61fd8-57aa-4927-b217-4cdf03e943ec · outbound

This paper cites Mauve: Measuring the gap between neural text and human text using divergence frontiers.

Statistical Multicriteria Evaluation of LLM-Generated Text Mauve: Measuring the gap between neural text and human text using divergence frontiers

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.905053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.872613Z digest=sha256:fc6d1372d5a8c3459e22d9399b96271ea12985f828462022e09d13a30ae513e4

Observation 30056779-8cec-45f2-b62c-1190e981c105 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Statistical Multicriteria Evaluation of LLM-Generated Text Direct preference optimization: Your language model is secretly a reward model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.876746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.876746Z digest=sha256:2e1bcf53815cfea76e18d5d96171eeb44882662976fbd857e873ccc85ce02274

Observation dd83a02d-229d-417e-a2c9-46d02366bcf8 · outbound

This paper cites Partial rankings of optimizers.

Statistical Multicriteria Evaluation of LLM-Generated Text Partial rankings of optimizers

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.879394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.881040Z digest=sha256:bd4701456df8e7775ec19a99e4ba1f89247e496ddb6e352c68426ba09c6af801

Observation c5a11e91-667c-4885-b3d2-382221ab37b1 · outbound

This paper cites Levelwise data disambiguation by cautious superset classification.

Statistical Multicriteria Evaluation of LLM-Generated Text Levelwise data disambiguation by cautious superset classification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.862698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.885608Z digest=sha256:34eabe9a3497c3c6c04591aab43f1268172947caddb5bcc7ce98a913d58b4152

Observation e07be0ab-dc50-4d13-9a55-4209ff7d99e0 · outbound

This paper cites Approximately bayes-optimal pseudo-label selection.

Statistical Multicriteria Evaluation of LLM-Generated Text Approximately bayes-optimal pseudo-label selection

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.846974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.890326Z digest=sha256:17442257a24fbc2258a8fffc3b1a295c059c890d8a1c384dfa552d5770614371

Observation 75b9fd2e-c80a-47e0-abcb-a0e83b233fd3 · outbound

This paper cites In all likelihoods: Robust selection of pseudo-labeled data.

Statistical Multicriteria Evaluation of LLM-Generated Text In all likelihoods: Robust selection of pseudo-labeled data

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.830638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.894647Z digest=sha256:75f0631a5b47e92510874b42ae67fac46d900f6a867734478eb4bc9ff51b55cb

Observation 5b254428-beea-4c10-9932-a0170c11af7b · outbound

This paper cites Explaining Bayesian Optimization by Shapley Values Facilitates Human-AI Collaboration.

Statistical Multicriteria Evaluation of LLM-Generated Text Explaining Bayesian Optimization by Shapley Values Facilitates Human-AI Collaboration

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:01:23.280879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.899482Z digest=sha256:00e832a971fb27532b82c340e6b98a0774bc0a6b50aaf98c5e07fe3ae2092d6a

Observation 7053843b-859e-4f0e-bdc4-214a69dae1a9 · outbound

This paper cites A Statistical Case Against Empirical Human-AI Alignment.

Statistical Multicriteria Evaluation of LLM-Generated Text A Statistical Case Against Empirical Human-AI Alignment

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:01:23.257694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.904326Z digest=sha256:abc0ed797a09a14f837a281fba927d411dfeaee721fe47f228710f25ee51423d

Observation e1478825-dbcd-4ec9-86c8-772f9d969d6d · outbound

This paper cites A meta-analysis of overfitting in machine learning.

Statistical Multicriteria Evaluation of LLM-Generated Text A meta-analysis of overfitting in machine learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.815993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.909152Z digest=sha256:9a182f71c2ad8bf9e3ca2e69d3ca19c54979a5948f6235abe1fb3daf1065239e

Observation 8c2823b6-54a7-4337-b4fc-06440b25f108 · outbound

This paper cites Schneider, L.

Statistical Multicriteria Evaluation of LLM-Generated Text Schneider, L

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.800614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.913529Z digest=sha256:21960ca59a6a94eb01b30a03c95e1cf7efb110f6f0a9c5bc0aaff8afd3e2d957

Observation 96a68390-0e70-4846-8512-f23487770614 · outbound

This paper cites A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes.

Statistical Multicriteria Evaluation of LLM-Generated Text A minimax learning approach to off-policy evaluation in confounded partially observable markov decision processes

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.781777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.918228Z digest=sha256:8cd14756b552cfe8e62f433634d976731927371f8840a1d238c69abd98107834

Observation c232a105-626a-44e4-a2d6-a8d84a0dc1c4 · outbound

This paper cites Shirali, R.

Statistical Multicriteria Evaluation of LLM-Generated Text Shirali, R

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.764739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.922919Z digest=sha256:f02ff8a0b626483596417d5777545ca8f0a1a1b365d08d9c8087615e85f67a2b

Observation e0938fbb-fa31-4e29-95c0-98b183dd0140 · outbound

This paper cites An empirical study on contrastive search and contrastive decoding for open-ended text generation, 2022.

Statistical Multicriteria Evaluation of LLM-Generated Text An empirical study on contrastive search and contrastive decoding for open-ended text generation, 2022

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.746915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.927214Z digest=sha256:711ace168c21207cbf4283137cb0153a7a1821c502d989c5fbeca70b48c7b4c7

Observation 8aa759bd-0119-472c-9ecb-b54d50bca22e · outbound

This paper cites A contrastive framework for neural text generation, 2022.

Statistical Multicriteria Evaluation of LLM-Generated Text A contrastive framework for neural text generation, 2022

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.730152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.931636Z digest=sha256:902e8abca35cf6a929e7530c498f0f0eac8eb8a99a0a2222bb968b50770d61f3

Observation 8d2b3088-51e4-4876-91a8-b01909cd9031 · outbound

This paper cites Evaluating the evaluation of diversity in natural language generation.

Statistical Multicriteria Evaluation of LLM-Generated Text Evaluating the evaluation of diversity in natural language generation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.713796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.935910Z digest=sha256:2265d950b4dd25c57eb21f20bdc53351b7896e262c5fca7b39645d56907dfe5e

Observation 3f3b1e7d-5e75-4367-bc58-f7eac5178bc6 · outbound

This paper cites Scientific machine learning benchmarks.

Statistical Multicriteria Evaluation of LLM-Generated Text Scientific machine learning benchmarks

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.696300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.940493Z digest=sha256:978cae989d69117021b381707279515241eb1c149daaa0e249fe556e54daf5b8

Observation 25df4b8e-684c-49ce-ae01-76e76555ceab · outbound

This paper cites Openml: networked science in machine learning.

Statistical Multicriteria Evaluation of LLM-Generated Text Openml: networked science in machine learning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.944941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.944941Z digest=sha256:6013425a36460a68e6263bb46de08c714adf240d55ff5d08cf1feaae24dfaabc

Observation 688c6c05-f339-4aca-aea3-74751283f638 · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.949512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.949512Z digest=sha256:bccaefabf801c357bb965ceb8f18fc960239cfdbda2a2a17a05e27f1ed7c24ce

Observation 69419229-b8ae-4126-ae63-cd1742e20623 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Statistical Multicriteria Evaluation of LLM-Generated Text LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.954002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.954002Z digest=sha256:4d65c0920dff504efa3e33374cbab7dff89545a0195ec6323ca87d83b9d083a1

Observation 79b025d7-184c-4c3c-8d74-3b23bac0e1bd · outbound

This paper cites Principled Bayesian Optimisation in Collaboration with Human Experts.

Statistical Multicriteria Evaluation of LLM-Generated Text Principled Bayesian Optimisation in Collaboration with Human Experts

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.958957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.958957Z digest=sha256:10a0cf1cb8d3814aea29fa2661e488c97f8474d5fccdd474ab8f208cef4ccfd9

Observation 02fad94f-8661-4f1e-a412-cf73f8c6bb4f · outbound

This paper cites Qwen2 Technical Report.

Statistical Multicriteria Evaluation of LLM-Generated Text Qwen2 Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.964167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.964167Z digest=sha256:524debbd2c2b679e7895fc80c30348efbdd22f0201e81ca9ae8ef6da75040cb8

Observation 9dd85e22-0986-42fd-9dd0-adf64cea4b89 · outbound

This paper cites Benchmarking llms via uncertainty quantification.

Statistical Multicriteria Evaluation of LLM-Generated Text Benchmarking llms via uncertainty quantification

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.657388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.969120Z digest=sha256:9e2e69a2ec4a4d0b303542babe912553a3755c39cebb30db0494c8f8f9e3b330

Observation 80df0dbc-876b-48e6-b825-c1c79dbccd9c · outbound

This paper cites Zhang and M.

Statistical Multicriteria Evaluation of LLM-Generated Text Zhang and M

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.641910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.973816Z digest=sha256:3fc30ebf5c1e3c7d2818a21a1f5b7549117d5edf414cee13f31f3c35633c7b2e

Observation dee8534b-81d3-4281-b9c6-c88936471016 · outbound

This paper cites Inherent trade-offs between diversity and stability in multi-task benchmarks.

Statistical Multicriteria Evaluation of LLM-Generated Text Inherent trade-offs between diversity and stability in multi-task benchmarks

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.626263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.978485Z digest=sha256:ae98bf68943434f652b00f3e67391fee72c7aab245ccf8c1389d2e42f75798e6

Observation dcc7d216-632a-4d04-98b1-03c0e069dab1 · outbound

This paper cites Zhang, M.

Statistical Multicriteria Evaluation of LLM-Generated Text Zhang, M

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:01:23.607704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.982795Z digest=sha256:a1ec5480c7c532ce402698b9aa1a67f8667d88024c66bfe84f13151663095807

Observation bbbe07c5-3379-493f-b2b2-4478d742f342 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Statistical Multicriteria Evaluation of LLM-Generated Text OPT: Open Pre-trained Transformer Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.987507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.987507Z digest=sha256:97f832f22db584a37918ac916c1bfebb06d49700ca2bf3118eaa00ee3c326852

Observation ce165ab0-9b6b-4fd7-b5c3-b509f1e98262 · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

Statistical Multicriteria Evaluation of LLM-Generated Text Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:22.992886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:22.992886Z digest=sha256:36528df8649c0822a78b8ea9929ba817ae28876a83a2443728a30a7e6998b12e

Observation 84ec1c5f-a10f-4a70-aa16-4e7972dd720e · outbound

This paper cites Time-Varying Gaussian Process Bandits with Unknown Prior.

Statistical Multicriteria Evaluation of LLM-Generated Text Time-Varying Gaussian Process Bandits with Unknown Prior

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-08-15T19:01:23.127948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-15T19:01:22.998385Z digest=sha256:ec20f7d8e3dae1e39d71ec6cd5b15b9304d593ec2961de72d8719695c37a47b1

Observation f055e5b2-a437-4db0-beb4-bd59ff17b15d · outbound

This paper cites write newline.

Statistical Multicriteria Evaluation of LLM-Generated Text write newline

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:23.004618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:23.004618Z digest=sha256:bc798d3fa59f9dae4e147e60e51e1fea35d409345ca065133ec3f1565be00621

Observation 937452c5-e766-4b61-a551-55a656f248d4 · outbound

This paper cites @esa (Ref.

Statistical Multicriteria Evaluation of LLM-Generated Text @esa (Ref

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:23.010839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:23.010839Z digest=sha256:31e7eaf87465f394df9cdc8674c2e2b726a323c18d8377428603f23c155ade01

Observation 85b98448-6d2f-48b0-85f8-7259e64911d8 · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:23.017499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:23.017499Z digest=sha256:5bf527bf56018fb11b08ce36024f43fa1f7457b02773fe588a63eacc418a8739

Observation 07c14cbb-1a00-47ed-a09a-6fc1bb80fc65 · outbound

This paper cites an unresolved cited work.

Statistical Multicriteria Evaluation of LLM-Generated Text Unresolved cited work

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-15T19:01:23.023749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:01:23.023749Z digest=sha256:581ba31225d4f56056539436a0f6e9c4b5ea30fc9daae2726d215acce057830c

Pith citing papers

Observation 9b8b12fd-8aec-424e-83be-138f89dee1bc · inbound

Self-Supervised Laplace Approximation for Bayesian Uncertainty Quantification cites this paper.

Self-Supervised Laplace Approximation for Bayesian Uncertainty Quantification Statistical Multicriteria Evaluation of LLM-Generated Text

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:52:17.137374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T04:49:15.239636Z digest=sha256:3d98bb7388313635f5f3cc67ad3bf55b626823fa3d0977a9037f676ff3ebcc3c