Pith. sign in

Paper Citation Record · LEDGER

Statistical multi-metric evaluation and visualization of LLM system predictive performance

As of 11 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2501.18243.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18243 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:16:37.811235Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:45:26.654826Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:47:27.686727Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 566efa22-b0e5-4cca-b077-c4d02f84bf7f · outbound

This paper cites Using Combinatorial Optimization to Design a High quality LLM Solution.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Using Combinatorial Optimization to Design a High quality LLM Solution

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:16:37.855138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.732561Z digest=sha256:fdca57c9b585eaeba24b76ad4cc6ab31c51980989ac17fa81ab8323cfc7fe92e

Observation 04234310-eb59-4515-bb0d-50ff4bb4e5e1 · outbound

This paper cites Simultaneous confidence intervals for ranks with application to ranking institutions.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Simultaneous confidence intervals for ranks with application to ranking institutions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.068354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.736883Z digest=sha256:335202280521a57a0cf7bb7fca05e5a46d92e14caee62e26e93b8f53cb046b16

Observation dbba6ae6-80c4-4a9a-aefb-711f4223e140 · outbound

This paper cites Statistical Power Analysis for the Behavioral Sciences.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Statistical Power Analysis for the Behavioral Sciences

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.057584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.740522Z digest=sha256:f306be5a35b508df89b592bbeec4423b45d457e25ddbc479895f1b3a5af86a4c

Observation 65ecd948-64de-421c-b42e-bc66b4a359c7 · outbound

This paper cites The spotis rank reversal free method for multi-criteria decision-making support.

Statistical multi-metric evaluation and visualization of LLM system predictive performance The spotis rank reversal free method for multi-criteria decision-making support

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.047600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.743949Z digest=sha256:82d8eca3f2679ae03a9275b4111315feae593cca363a742a49788bdbf25d804b

Observation 146b089a-d610-4ffb-9ddd-3d33d9cc0478 · outbound

This paper cites CrossCodeEval : A diverse and multilingual benchmark for cross-file code completion.

Statistical multi-metric evaluation and visualization of LLM system predictive performance CrossCodeEval : A diverse and multilingual benchmark for cross-file code completion

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.037414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.747876Z digest=sha256:b90a741b3251916d29b56ed48b5a97cccdc025fa3434c1f798b5a7ee4125ff29

Observation 10d7a8cd-a604-4d69-8e87-deff4db695cc · outbound

This paper cites Multiobjective optimization in river basin development.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Multiobjective optimization in river basin development

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.027060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.751357Z digest=sha256:0618ca1abe983367da4a61de3b2322be237163edfa5510d778c38850d87aa3f1

Observation 18e66ac1-1997-42f8-af99-510f415f4646 · outbound

This paper cites Sensitivity of decisions to probability estimation errors: A reexamination.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Sensitivity of decisions to probability estimation errors: A reexamination

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.016971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.754939Z digest=sha256:bb42f40336a837d92fae08e848ff5d3fe3962fb1eb45786fd3923cb82836853a

Observation a3cb713e-bac8-4366-ba79-3b15c89ed2ac · outbound

This paper cites Performances are plateauing, let's make the leaderboard steep again, 2024.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Performances are plateauing, let's make the leaderboard steep again, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:38.006862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.758183Z digest=sha256:c282210e92f5422a0c89ab5e80fbd343a1df5d4e1724fa9a55fb669ff2ae5311

Observation 3924f555-4d03-4f13-a201-3e123cb7d0b7 · outbound

This paper cites Choosing between methods of combining-values.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Choosing between methods of combining-values

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.996655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.761668Z digest=sha256:8a03dac66bd806666fecac6a39eb4b93016c132d9d09bb00dd24dcef629ebe3f

Observation 6ead93d0-6238-4d15-90c9-47eb85b512fb · outbound

This paper cites Methods for Multiple Attribute Decision Making, pages 58--191.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Methods for Multiple Attribute Decision Making, pages 58--191

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T00:16:37.765034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:16:37.765034Z digest=sha256:f14f5a061a02a92a8c14cc4a1431940776b7a9f03fa9aa85c80982f714958be7

Observation 7d9e3fec-4b19-4467-902d-03cbffc60de0 · outbound

This paper cites pymcdm—the universal library for solving multi-criteria decision-making problems.

Statistical multi-metric evaluation and visualization of LLM system predictive performance pymcdm—the universal library for solving multi-criteria decision-making problems

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.986406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.768698Z digest=sha256:fce717c6c6763de449e61a5d727451803689371216c43d5c15d4ed2431e77a0c

Observation fb46097e-e7f7-4c1c-833a-96b88c2b8ee0 · outbound

This paper cites an unresolved cited work.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Unresolved cited work

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T00:16:37.976123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.771984Z digest=sha256:9f545a3205ce25f5872f38e6fed1a74062e08a0928749256a11680ddbbe27240

Observation 80e6f8df-2f84-4966-902d-3c5a65cb3b2e · outbound

This paper cites an unresolved cited work.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:16:37.965615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.775570Z digest=sha256:15724b7e44a0e65f6785a042483909a390d551669be74a1813704d7a25c125fa

Observation e0dde2ef-6946-42d1-a5ff-4b4f2838527c · outbound

This paper cites Mangiafico.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Mangiafico

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.955295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.778900Z digest=sha256:885b1f3543150caa1a53efc0edde504fc1df4c5a308729856d05a49a74cbbe14

Observation fea60797-81dd-4a12-b480-25545207c333 · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T00:16:37.782195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:16:37.782195Z digest=sha256:41097ab88afbeef17829becc3e40f54d774b43f9feb937814c4f507b14e26635

Observation 6cd96f46-fd17-454c-851c-7f8ef76e2f37 · outbound

This paper cites New effect size rules of thumb.

Statistical multi-metric evaluation and visualization of LLM system predictive performance New effect size rules of thumb

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.944589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.785701Z digest=sha256:94ec7d5e2b6f43519927e5e58df15efaec4faf8c0c6dfb61e077abb5b0a17bb2

Observation 38214852-9c4a-4136-8253-42363027f9b3 · outbound

This paper cites Statsmodels: Econometric and statistical modeling with python.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Statsmodels: Econometric and statistical modeling with python

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.933028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.788961Z digest=sha256:75d1eef8345e7e60db694101ea20b726173c93286742bde4426e11850b7ead01

Observation 10c01e60-8dca-4ec2-80f1-0f73ea3f7893 · outbound

This paper cites A multiple criteria decision making method based on relative value distances.

Statistical multi-metric evaluation and visualization of LLM system predictive performance A multiple criteria decision making method based on relative value distances

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.923503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.792056Z digest=sha256:7fb7a4f36ae2bcd01f86329f089826c7157e67ef88a2a922c5eb6ad47d9fc141

Observation e09f9612-cc96-4cb7-97eb-66ac0de49479 · outbound

This paper cites Comparative analysis of some prominent mcdm methods: A case of ranking serbian banks.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Comparative analysis of some prominent mcdm methods: A case of ranking serbian banks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.914021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.795146Z digest=sha256:6ad9b39a7343fd887ebbccb036b983e08eeefc6929c8958c025fefc6e71bde7a

Observation 2befde40-c5ff-4ec2-b8a0-d120decf3582 · outbound

This paper cites Using effect size—or why the p value is not enough.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Using effect size—or why the p value is not enough

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.904452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.798176Z digest=sha256:ba4969059809d4040fb04285e5b230b40928fd20fe36b6382a6ce62f8ea1a184

Observation cc56db27-137e-4e46-8edc-bc2589e04b89 · outbound

This paper cites Calculating and synthesizing effect sizes.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Calculating and synthesizing effect sizes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.894808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.802219Z digest=sha256:641d1bfa8b7b877161ed0b795065da116eaac6ef4eb59962e43de8c654349bf4

Observation ebff8b2c-07e0-4a4a-acb6-6d151fa38c95 · outbound

This paper cites Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, St \'e fan J.

Statistical multi-metric evaluation and visualization of LLM system predictive performance Oliphant, Matt Haberland, Tyler Reddy, David Cournapeau, Evgeni Burovski, Pearu Peterson, Warren Weckesser, Jonathan Bright, St \'e fan J

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T00:16:37.805163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:16:37.805163Z digest=sha256:28170bff0d2fc54903f8de58156c384b2238849f1f459fd59be0bb7e4b26a9df

Observation 616e18c3-5041-4049-9a04-c05c459df20e · outbound

This paper cites The harmonic mean p-value for combining dependent tests.

Statistical multi-metric evaluation and visualization of LLM system predictive performance The harmonic mean p-value for combining dependent tests

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.876952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.808144Z digest=sha256:d0c5aa310e976c5e4ccdf88720ae46b11b147db1701c65d87569a7c3819c2cea

Observation 651f04b1-56d5-49d7-bca7-d10048e76774 · outbound

This paper cites harmonicmeanp tutorial, 2024.

Statistical multi-metric evaluation and visualization of LLM system predictive performance harmonicmeanp tutorial, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:16:37.866262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-10T00:16:37.811235Z digest=sha256:846bbec761ecc3c3ad0865001516fd7c904651ea4fada42557a68236075db359

Pith citing papers

Observation 5e68a197-fb9c-432f-9662-1cf1e06d0e88 · inbound

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation cites this paper.

Rank Intervals for Leaderboards: A Hierarchical Framework for Model Evaluation Statistical multi-metric evaluation and visualization of LLM system predictive performance

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:47:27.688270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T17:54:22.974336Z digest=sha256:7483f35d020a3321219a28079218a800da1dc2cfa34c2bd448f7308a6647cdc0

Observation 0ce2d397-8c5e-43ab-b500-c7fbe4be67dd · inbound

Quantifying Ranking Uncertainty in LLM Benchmarks cites this paper.

Quantifying Ranking Uncertainty in LLM Benchmarks Statistical multi-metric evaluation and visualization of LLM system predictive performance

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T09:45:26.654826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:45:26.654826Z digest=sha256:e63a2af5ebc4eb7e7f8a8899d3e4f70cda04215c8a773dd51c37692435680174