Pith. sign in

Paper Citation Record · LEDGER

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data

As of 11 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.23735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23735 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:45:02.174033Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b59611fa-8696-4bb1-9e6f-0df2cd2a98a9 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.136247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.136247Z digest=sha256:c8e80c5ee62c4c4a38d4cd0d658d667164b77873ec3e356593358309b397dec0

Observation 025d6689-70c1-4393-902d-71418d1ccd7c · outbound

This paper cites Deepseek-v3: Scaling open-source language models with mixture of experts.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Deepseek-v3: Scaling open-source language models with mixture of experts

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:45:02.796646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T21:45:01.188367Z digest=sha256:b956d8bdb5ba85031e5cda59b3da5ce860028592ae852a8c7e36ec351d9c1646

Observation 2a67eaf4-1d63-4ccc-9c49-ca13d65b5496 · outbound

This paper cites Black-box generation of adversarial text sequences to evade deep learning classifiers.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Black-box generation of adversarial text sequences to evade deep learning classifiers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.255575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.255575Z digest=sha256:3d8063c4a2f284299e1a1c846b0721755819679a9e55fcfbbe94012e71b30a07

Observation e6813975-75c2-4d2f-ad10-1df82cbb37cd · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Measuring Massive Multitask Language Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.381819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.381819Z digest=sha256:e77417a88e3af38984d95e92f43331477149741711f33ad5ba7105ccd1a8377e

Observation e14cfab9-b745-4349-ba58-6caa471c5fef · outbound

This paper cites C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data C-Eval: A Multi-Level Multi-Discipline Chinese Evaluation Suite for Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.465673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.465673Z digest=sha256:2cd7b435ec7e7c645470a56a43ca2be12fcd039fcec720ba7830dae56a95c537

Observation ea7df36b-156b-492d-ab37-f536b1963162 · outbound

This paper cites Adversarial text generation by search and learning.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Adversarial text generation by search and learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.552009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.552009Z digest=sha256:e1755103a3909bcb15fa5496409402f8d63ed801f87875bb2e70d2f70edbbc26

Observation baaf731e-2d06-4501-9f47-67da1aceedf5 · outbound

This paper cites Beyond Static Datasets: A Deep Interaction Approach to LLM Evaluation.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Beyond Static Datasets: A Deep Interaction Approach to LLM Evaluation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.648196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.648196Z digest=sha256:0f4123fa44d033d8366c6d110eae78d3d8e1d3f90cfc44ed81ca7786ceb0f60e

Observation 53df4759-7c2d-4fb4-9e65-93c92e196499 · outbound

This paper cites PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.719577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.719577Z digest=sha256:2db9667b3d4194ce4d532be16741aa09406db7b9671a2067d1fabb6c71448cb5

Observation ca5d8135-ce8e-42d7-bc25-ef1220d1d270 · outbound

This paper cites Using Adversarial Attacks to Reveal the Statistical Bias in Machine Reading Comprehension Models.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Using Adversarial Attacks to Reveal the Statistical Bias in Machine Reading Comprehension Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.788221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.788221Z digest=sha256:c9ac398167a3a538f2ece19f7252ae07e8129d81f316526243e77771800ff2d2

Observation 0b0ab9bf-24d6-45b8-b827-73b7a419d91d · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.892149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.892149Z digest=sha256:26d27ec4d4f5cafa492b26b787661085a11378b339942ec1f8ce60dcad9588fd

Observation 6fab41cb-ecc3-4953-aca6-697cfc9ef00b · outbound

This paper cites MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data MedMCQA : A Large-scale Multi-Subject Multi-Choice Dataset for Medical domain Question Answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:45:01.978036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:45:01.978036Z digest=sha256:ae85e70821df73febd5d8878052a9ceee7c13619ca9e1a28014fe129bbe99188

Observation 6c0eae8f-4d19-48d9-ac76-bfaf9f4e6033 · outbound

This paper cites Alcuna: Large language models meet new knowledge.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data Alcuna: Large language models meet new knowledge

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:45:02.656878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T21:45:02.073822Z digest=sha256:d7143c8b4f32e3c32180052cc0e658769c696835f5bda476b332f5bbdc5596b5

Observation bfd065d6-b13c-4793-8dbe-5967b7222084 · outbound

This paper cites ALCUNA: Large Language Models Meet New Knowledge.

AutoEvoEval: An Automated Framework for Evolving Close-Ended LLM Evaluation Data ALCUNA: Large Language Models Meet New Knowledge

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:45:02.341402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-06T21:45:02.174033Z digest=sha256:d30ee42adedb7cd41d96fb6159d35ae6cde1d2cdf7c2da2cec56fe50a790cb56

Pith citing papers

No inbound Pith citation observations are available.