Pith. sign in

Paper Citation Record · LEDGER

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2502.01683.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01683 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:09:01.728564Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:39:36.819142Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T22:40:49.906159Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cfc5650-ba39-478e-84b4-ee36304aca78 · outbound

This paper cites Claude 3.5.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Claude 3.5

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:09:02.311909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.591091Z digest=sha256:f51a0299c316e1ce07912ada6014d98e52e383495d6f9bbb03b943b274109526

Observation 409b5d83-32f7-40d0-90ca-e91565196e3c · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.304253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.594566Z digest=sha256:701bf412577988c862d88ecba72ef82a6b7058c102d4dce7fed3d93080505a0b

Observation 36a20913-6729-453e-880c-f85fcff73202 · outbound

This paper cites Scaling Synthetic Data Creation with 1,000,000,000 Personas.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Scaling Synthetic Data Creation with 1,000,000,000 Personas

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.597866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.597866Z digest=sha256:f2bf4c242a9bac73a1e24d5d16f517774f0e50ce9b43ad1bff264d85c50e6250

Observation 92662432-f6f3-4d5e-954c-b51bd990f875 · outbound

This paper cites Yu, Qiang Yang, and Xing Xie.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Yu, Qiang Yang, and Xing Xie

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.601806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.601806Z digest=sha256:3822c03176c0caf19d1d3e4fdcbc089242536ddcce7fc132001a2dedf86cad7f

Observation 63181b01-6028-476d-a761-cf84cc0c8618 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.604922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.604922Z digest=sha256:9e237054ee8d2bc86f796179e34cee0907a5262804d5ad5a8fa0e48bbd0a6320

Observation a34b53b7-cf57-4948-83e5-602a6b6665bf · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.608303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.608303Z digest=sha256:192aaf9599b2ae184092f845113a0d15c86d299e6eebaf3511b9175eba8d99c2

Observation 0c544555-8fb4-4a4d-862e-3b86ac823563 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.296828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.611435Z digest=sha256:3f154f2e94a68fae434ef6159309c5d800ed508b7b7766d54b39674d34697ce4

Observation 98f7e8f4-cd76-4e95-acbf-8ac9a4df1e8d · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.614868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.614868Z digest=sha256:cab111c1d78a142709c878f875e5778af67107eed53e6ae1dbf09e20207c4961

Observation e87f3976-8e7a-4022-a988-58993e941b82 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.282798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.617635Z digest=sha256:0bd34baf5149daafb8f0208dc523bc5fced0a82fa136a3cf4bf4318e74c0b581

Observation 883b262a-7a1d-4d85-a5a3-b806e69c743c · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.272809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.620591Z digest=sha256:e11e4a236917da3f00778c4a50952ad9f321d740d26d233406ce205a2fc06596

Observation 46110983-3b6c-4d9d-a724-85304659df95 · outbound

This paper cites GPT-4o System Card.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient GPT-4o System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.624555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.624555Z digest=sha256:6b47b266a8894988ed0e9ecf0e55abc465c7944dce29da882e94fdea328b4a59

Observation 8048585d-e5f7-4b1b-97c7-e9d50480054f · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.627606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.627606Z digest=sha256:3fdef68f4c0c77b40db36479181c689ecd23d8ee0be32abf27b634197bf38719

Observation b41e42e1-dd22-4450-b693-b7db6bbc2283 · outbound

This paper cites Causal Machine Learning: A Survey and Open Problems.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Causal Machine Learning: A Survey and Open Problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.630239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.630239Z digest=sha256:0a51e2ddf36c559aa5af300a1fac33f1fac80cf04097400e64bbc055ed0fad07

Observation ed49413c-ef14-4cc2-9cf8-c678d0e839e4 · outbound

This paper cites S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient S3Eval: A Synthetic, Scalable, Systematic Evaluation Suite for Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.632998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.632998Z digest=sha256:6fced65da0330dab89db7322bc54c9b3b98241c3f6b91344a3b6b229f23caf9a

Observation 9dc084b1-b4bb-49ff-a534-a4ec3c054b90 · outbound

This paper cites u ttler, Mike Lewis, Wen - tau Yih, Tim Rockt \.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient u ttler, Mike Lewis, Wen - tau Yih, Tim Rockt \

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.636225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.636225Z digest=sha256:4fa6f1fa986ecc54309f7e71784b6192216cb85632e172bd7d96e5908f7a2d7b

Observation a9011def-dae2-4f1d-87ff-a979333e27b6 · outbound

This paper cites PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient PertEval: Unveiling Real Knowledge Capacity of LLMs with Knowledge-Invariant Perturbations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.639375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.639375Z digest=sha256:dd8be8a1cc46b362fef96e3c058d66e6c240ee41e23fed319e2b7a95f9efd19c

Observation 9de263b1-2e74-46f2-9b23-f5dce2ebfe53 · outbound

This paper cites SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient SciLitLLM: How to Adapt LLMs for Scientific Literature Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.642749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.642749Z digest=sha256:65da16e5b4b0969fcf50e3a5a5c9e28118dc95a9d9f9a4f8b49b697592481f2d

Observation fa841f16-f458-4531-b321-e2c6e00cf99c · outbound

This paper cites Neuro-Symbolic Data Generation for Math Reasoning.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Neuro-Symbolic Data Generation for Math Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.645942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.645942Z digest=sha256:b4d84ee97d6de5ce3ae372018ec90679f59f48fc5f9c9307d09c4820a29f48e4

Observation f34460d2-0fbb-435c-9df8-cf06c89b198e · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.259671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.649187Z digest=sha256:2c3a2233f6a20ec663ad9dd8d85f267ba35727fa9b879ac332bc9c1dc530fe3a

Observation 6bae59c9-87d6-4204-992f-178418605f0d · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.651983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.651983Z digest=sha256:4229ea2f1c66d032d316c1fb7b0a0f877aa2bde09d02b5959108b3e31c2d820e

Observation 4b83488c-1a55-4c61-99f9-e55a04e3cea2 · outbound

This paper cites Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.654904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.654904Z digest=sha256:1db77893e4713ef856511d70e62357f676337281afce523be5f2deb0aab9cc53

Observation 062255dc-8e6f-4159-a0b9-9b7359880775 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.250661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.658012Z digest=sha256:c5f1c54a08a5cb9b120ff0d40d4df03ac7f6d758f9b11a3a5a60cb8141426bcf

Observation 47e2f6de-f977-4a8f-a546-7593aeafc266 · outbound

This paper cites Efficacy of Synthetic Data as a Benchmark.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Efficacy of Synthetic Data as a Benchmark

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.660763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.660763Z digest=sha256:67a5f4b4ad6555db476adfadcfae1847284220e44cc821895cc52ca61b9d7402

Observation 813d0a62-8588-41a4-9753-68695f3865a2 · outbound

This paper cites Jointly Measuring Diversity and Quality in Text Generation Models.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Jointly Measuring Diversity and Quality in Text Generation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.663911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.663911Z digest=sha256:4af63d50f1fa07798d23b6f4b06b9e016d3eda1870e47a4ed739ecf3f7aa28f6

Observation 01342007-e778-4b7c-9999-a77c32f4bee6 · outbound

This paper cites text-embedding-ada-002.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient text-embedding-ada-002

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:09:02.241496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.667583Z digest=sha256:dbc50e016ebc65fba2247ac528af8944eb5f8b53a2b6d04bd89c80650ed2aac9

Observation 1d3daef6-8ce7-4d38-920d-10f5a683dc7b · outbound

This paper cites Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.670630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.670630Z digest=sha256:e28d98dee8f44a87624c5ca8904d4daf54297b1756baead2e0b239a531357b86

Observation c3ac89a6-e4f0-452f-9a4f-66abd3dcc4f5 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.227588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.673654Z digest=sha256:37bbdf00cf0a7a0583c9346073325babfaae9a385d7441aa4e088dc6864b4a71

Observation 322e1967-c2d6-4d9b-9e65-c7729d065468 · outbound

This paper cites A Survey on Self-Evolution of Large Language Models.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient A Survey on Self-Evolution of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.676414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.676414Z digest=sha256:fee6624985f882e8842e4aeeeb962dbdb4d496f2a28b247ba05a08982e4e6085

Observation 6338cc35-3f69-4baf-81b6-61d9ce791fa7 · outbound

This paper cites Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Judging the Judges: Evaluating Alignment and Vulnerabilities in LLMs-as-Judges

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.679147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.679147Z digest=sha256:6ef0e95676ca30dbd0815bd2090cac8476031d4058e5aa1690667f8c7af004c3

Observation 15100bd0-b637-4948-b51c-5437ba88c379 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.681746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.681746Z digest=sha256:109d5d716b66ca885fcf041149e939e68e973f7e1cb43d06fd3d7a93a21cb75b

Observation 7aadc72c-cbde-4595-9497-b379dc9755cc · outbound

This paper cites A Survey on Data Synthesis and Augmentation for Large Language Models.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient A Survey on Data Synthesis and Augmentation for Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.684146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.684146Z digest=sha256:dcea44a3db29387225b7300f46a559aedecc461fa5798ff1f2dbf64ae8d9f440

Observation 802f192e-afe5-4bc8-8a87-08459d2ff2e5 · outbound

This paper cites Le, Ed H.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Le, Ed H

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:09:02.219700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.686768Z digest=sha256:e617655b05883fb165d6f1afbaaaac57725019126b40ebfc9a6ac478f0a7da88

Observation 0dfa2173-c8c7-4ebc-843b-b17971b1f720 · outbound

This paper cites Smith, Daniel Khashabi, and Hannaneh Hajishirzi.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Smith, Daniel Khashabi, and Hannaneh Hajishirzi

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.689516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.689516Z digest=sha256:7d5313f6527d0718f5a76e2c0d9e5800e98a784e4b299785f61365569bdaa94a

Observation f201a6fa-2ab2-48c6-9b03-78a3b22f9612 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.692472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.692472Z digest=sha256:da0d9f5a7f2007b51a095b144c1c04e269fe7defef7111d1f0cffa316aba7039

Observation a00b135b-b20e-4822-b2c6-cbe659517f71 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.695833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.695833Z digest=sha256:a6c011caeffda95c4af886b1d41b9e2852ca5e4a30ba3374b3f357996e03fb05

Observation 349d573e-8f42-4ca0-8451-ab19848bfad3 · outbound

This paper cites Qwen2 Technical Report.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Qwen2 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.699197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.699197Z digest=sha256:acb8090f23b7a49919d1063c2ff17df4577ac759b00a7a96c22a6e9767f3d7ca

Observation 23275359-7005-4eb7-865c-0dfc726f1abf · outbound

This paper cites Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.702260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.702260Z digest=sha256:ee1f16e6a428af80da9144b0364e52522842506c1311f8ea57a2c05c75d2f1ca

Observation 6d891ced-ae80-4651-8f13-d79f342473d2 · outbound

This paper cites Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:09:02.206843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.705080Z digest=sha256:d0de5a17ef2a38a6651f70282887e2cc1269e1618ab0a1b4e624acf9170b101f

Observation 01b89090-28f8-44de-9f76-cf7415f5dd79 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.708002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.708002Z digest=sha256:5ba339968ac7817236605a0c3eed63714ed09492b035ed470eb56560f2f493bd

Observation 1a30f425-3980-4673-b508-1551318daebf · outbound

This paper cites Xing, Hao Zhang, Joseph E.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Xing, Hao Zhang, Joseph E

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.711080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.711080Z digest=sha256:cb597086498e8fbdf82edd36783e846fccd22827e59d57bf4bc5174d2c0ac1cc

Observation 327f84c9-0dcd-4483-98f7-2268689f33f9 · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-09T18:09:02.191757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.714267Z digest=sha256:96557d0eeb9cbdef1b5d46475712cbaebb367dd105fa1617f72a4e2a3f83e053

Observation 1ee7a73b-7456-4d9b-8939-e8246680a985 · outbound

This paper cites Dynamic Evaluation of Large Language Models by Meta Probing Agents.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Dynamic Evaluation of Large Language Models by Meta Probing Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.717256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.717256Z digest=sha256:dead46acd219d51d4beb9748ff7175ff0e25b2d53a6f4ec0f59f97083f601eb0

Observation a16acfb5-74d5-4b93-9f88-ea0107d16fcb · outbound

This paper cites an unresolved cited work.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient Unresolved cited work

Reference 43

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-09T18:09:02.157754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-09T18:09:01.720740Z digest=sha256:0aa659ce286b932a74e06796056cada7ffa3eccf713b3d36a0c938d1424c1c6f

Observation 8b6d801a-7488-4531-aaa1-6d43c4e56507 · outbound

This paper cites URL: " 'urlintro :=.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient URL: " 'urlintro :=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.724723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.724723Z digest=sha256:1d1e08cbe50a474726cbd6400114b05fa0df27bba5b644be6cf5e76a0f8188d3

Observation 26448a66-4b4f-4581-b83b-0d457d9303ed · outbound

This paper cites write newline.

LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient write newline

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T18:09:01.728564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:09:01.728564Z digest=sha256:db507265d075a7a62d9d18fae7b9bdde4b72d95ea1b9bad00b3ca23bca023e51

Pith citing papers

Observation c925eb55-b244-4ff4-a252-25304c795a85 · inbound

Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis cites this paper.

Automatically Generating Hard Math Problems from Hypothesis-Driven Error Analysis LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:40:49.911691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-10T19:39:36.819142Z digest=sha256:3cc6ae287ec3eb5ab710baccfc994a91dc408e390219fdc7b5033a0b516fb9cf