Pith. sign in

Paper Citation Record · LEDGER

Data Swarms: Optimizable Generation of Synthetic Evaluation Data

As of 15 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2506.00741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00741 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:03:42.484850Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d016e96c-6595-4339-b0cc-cb96a723844b · outbound

This paper cites Kgquiz: Evaluating the generalization of encoded knowledge in large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Kgquiz: Evaluating the generalization of encoded knowledge in large language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:34.567215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:34.567215Z digest=sha256:a778d042f829c0f08aa4e0baf79d459a873c3ea33f8a6d5f756c14423efb0c58

Observation 3d0e456c-a8c9-4635-8c4b-0b332bff53bd · outbound

This paper cites AutoEval Done Right: Using Synthetic Data for Model Evaluation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data AutoEval Done Right: Using Synthetic Data for Model Evaluation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:34.669899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:34.669899Z digest=sha256:182088ad4b591b24143a6f2dee3e1da34aef775414718c769ac9ded54bc1c91e

Observation e5d3074a-1dc2-4a2c-8ffe-8c5179553f00 · outbound

This paper cites Adaptively evaluating models with task elicitation.arXiv preprint arXiv:2503.01986, 2025.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Adaptively evaluating models with task elicitation.arXiv preprint arXiv:2503.01986, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:34.895664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:34.895664Z digest=sha256:b83b5f321c7b2d633eaf7f6381212fecb0d213d2de02cf95a5e9cbb684caa7d5

Observation b921ed93-8e24-4bca-9e25-0bce68d43405 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.075359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.075359Z digest=sha256:200941422d1247d51dd79365ffa91050479b0a0c35b0e178324715b1743f7703

Observation 55493653-a93b-4192-b056-e954f47d10a2 · outbound

This paper cites Knowledge crosswords: Geometric knowledge reasoning with large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Knowledge crosswords: Geometric knowledge reasoning with large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:50.656173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:35.188332Z digest=sha256:9335ef4e11817c2d6b020348fd0a670d34aa22611318c6bd8b3afdf1fac17e56

Observation d38a2290-f18f-47e1-a916-e0d9b553976b · outbound

This paper cites Self-Boosting Large Language Models with Synthetic Preference Data.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Self-Boosting Large Language Models with Synthetic Preference Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.341828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.341828Z digest=sha256:6b59b9640ca712b1eec6e916b17ef716a045d3cd07aebbff87cebced3bf87f4d

Observation 013d73e7-0a88-4b2d-b3f3-9d4d03f5362b · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.Advances in Neural Information Processing Systems, 36, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Alpacafarm: A simulation framework for methods that learn from human feedback.Advances in Neural Information Processing Systems, 36, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.502201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.502201Z digest=sha256:4d7faa49256d73fd1af246a2c1571c72b4bc8261b7a368d25bf7c3f6970a5981

Observation d41b1c59-fd0e-4f48-afef-853faba12af0 · outbound

This paper cites Clas- sifying the classifier: dissecting the weight space of neural networks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Clas- sifying the classifier: dissecting the weight space of neural networks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:50.381847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:35.540623Z digest=sha256:b0c6fcc3d0fd5153e6bf16c2c893d7a85a0aafaaab896c405b31b52b2676fa65

Observation e987ecb4-b7f4-4c52-b6c9-3e905dfb908e · outbound

This paper cites Heterogeneous swarms: Jointly optimizing model roles and weights for multi-llm systems.arXiv preprint arXiv:2502.04510, 2025.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Heterogeneous swarms: Jointly optimizing model roles and weights for multi-llm systems.arXiv preprint arXiv:2502.04510, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.613142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.613142Z digest=sha256:aa32dbaf33062df7aa4379812e9b1b35f4980227e8add54bfe722b09a49c59f7

Observation fd4fb5fa-c7bd-4b36-85ce-d6907ad443ec · outbound

This paper cites Model swarms: Collaborative search to adapt LLM experts via swarm intelligence.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Model swarms: Collaborative search to adapt LLM experts via swarm intelligence

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:50.181618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:35.762723Z digest=sha256:6ffc3699de639503f8216c149ba87b24e00b4d1feb18c1359fd1770062c315cf

Observation 6301967c-1eb0-41d1-8804-d4b16fd9704a · outbound

This paper cites Promptbreeder: Self-referential self-improvement via prompt evolution.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Promptbreeder: Self-referential self-improvement via prompt evolution

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.915732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:35.916500Z digest=sha256:e74324804c68d6fa2e17024e385eb43826954643d51800807a5b3b2273ef9744

Observation 07f41ea0-74e2-4f30-ae48-8840954b843f · outbound

This paper cites Open llm leaderboard v2.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Open llm leaderboard v2

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:36.045712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:36.045712Z digest=sha256:ebc324d59f71cdfed3fcc3b86522dde2b04015001af0d1763bfd7278cefc1055

Observation e1da0c0b-27ac-47ea-bb3c-1a6b8142652a · outbound

This paper cites Time travel in llms: Tracing data contamination in large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Time travel in llms: Tracing data contamination in large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.649629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:36.184289Z digest=sha256:7bcebc541f790ba346cbccea43155a7297f0a5ec0c16b87bd6b19a0ab199a5e7

Observation 8178c412-ca1b-45f4-8c1f-4987319c36c6 · outbound

This paper cites Automated evaluation of retrieval-augmented language models with task-specific exam generation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Automated evaluation of retrieval-augmented language models with task-specific exam generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.382389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:36.332679Z digest=sha256:20e532359c0f84de3d9a6d291d891e1e79e2e42ce26a3fd7583fd933c5fe7f38

Observation 81453eb1-4329-4599-b071-66bf8072cbe2 · outbound

This paper cites Measuring massive multitask language understanding.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Measuring massive multitask language understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:36.464445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:36.464445Z digest=sha256:c5f7c09371412787a2e0c4d5eb69bb46676360e01b8615ed87c74bd1400d234b

Observation 5b49f598-0ae1-451f-a5fa-5464a234093a · outbound

This paper cites Datagen: Unified synthetic dataset generation via large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Datagen: Unified synthetic dataset generation via large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.134596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:36.603676Z digest=sha256:d99bc7b2d4f12c1bf2548172164fdea6ef67613abb4f4a14ad85af77ab972a6d

Observation 4f27019b-e3df-4b9c-8f00-91f772bb388f · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:36.765030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:36.765030Z digest=sha256:97b984216516c96aa3f052d91ada4b3c7ed2266146585b040b7051be72cda43a

Observation e55dc512-29eb-492a-8c5c-652e67ef1c96 · outbound

This paper cites Wildteaming at scale: From in- the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Wildteaming at scale: From in- the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.937768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:36.978535Z digest=sha256:84552dac448c35ba57d919dc62ce97c1d008c56355488f9fa7825dc00956f3d0

Observation 7002d98f-da27-43be-829b-d8edf4490dcd · outbound

This paper cites Teaching language models to hallucinate less with synthetic tasks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Teaching language models to hallucinate less with synthetic tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.645883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:37.116286Z digest=sha256:4056ea85e46ece1529ea993a9f1d1680a3f674cea8461650f6d0a86814b77cd4

Observation 34ed43f9-c088-4ec6-9a3f-ff7c1b7f7e9b · outbound

This paper cites Au- tonomous evaluation of llms for truth maintenance and reasoning tasks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Au- tonomous evaluation of llms for truth maintenance and reasoning tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.458711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:37.251988Z digest=sha256:a76b5e0002278b92bc7a027d8f2f2b13b137a6244599cc79870efe99d9bdc52b

Observation 736059c6-e70f-43bd-9a8a-6dc29d8e1b03 · outbound

This paper cites Realtime qa: What’s the answer right now? Advances in neural information processing systems, 36:49025–49043, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Realtime qa: What’s the answer right now? Advances in neural information processing systems, 36:49025–49043, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:37.403279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:37.403279Z digest=sha256:9772c516fe1b9acd16dc4e14fd55beb11a4f55bf335833144a0b992dc896640c

Observation e30f5e7f-8934-4ad5-8a4c-84496bd00fd0 · outbound

This paper cites Particle swarm optimization.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Particle swarm optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:37.581123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:37.581123Z digest=sha256:06cdae6468b037d97cc813d509c17a8deebb559458497d95f1de00e9f27bbff1

Observation e7b23351-9ef6-4c55-8752-f02f950bcf67 · outbound

This paper cites Openassistant conversations-democratizing large language model alignment.Advances in Neural Information Processing Systems, 36:47669–47681, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Openassistant conversations-democratizing large language model alignment.Advances in Neural Information Processing Systems, 36:47669–47681, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.194529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:37.767108Z digest=sha256:1ec93bbf6bd3bd9369020815fa19565e237b17f8d7df9dcd2245e6356ca27f3f

Observation e40cc949-e6b2-43c0-b393-e6d8927bf100 · outbound

This paper cites Eliciting Language Model Behaviors with Investigator Agents.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Eliciting Language Model Behaviors with Investigator Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:37.921485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:37.921485Z digest=sha256:3054af6730e0d2c4069d0ed6fca6cfa1a91b94bd3360356ff62f4358ff6d9da3

Observation f3b21320-0d42-4975-ba23-a63db57a5890 · outbound

This paper cites Autobencher: Towards declarative benchmark construction.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Autobencher: Towards declarative benchmark construction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:38.098878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:38.098878Z digest=sha256:9a3187348583270e6fe4b45440a705e4b9bc67f725efb972dafa2d4dace53c62

Observation d244d9dd-3773-4826-ad02-5471a91be604 · outbound

This paper cites Gen- dataagent: On-the-fly dataset augmentation with synthetic data.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Gen- dataagent: On-the-fly dataset augmentation with synthetic data

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.979881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:38.227102Z digest=sha256:95e5ab0330834a94e4ea994d163c6e2865fe9fdeb4304d0006b5958b14233256

Observation 8da932b7-31ef-4c50-8c38-e62277f8c458 · outbound

This paper cites Hemm: Holistic evaluation of multimodal foundation models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Hemm: Holistic evaluation of multimodal foundation models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.710055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:38.373411Z digest=sha256:2e2141486ec8b11ad49367afb2f76582ded3f36b30ddf41ea18d54b52e368ddc

Observation a89be644-4aab-44c0-92d4-34e813ff1c33 · outbound

This paper cites Holistic evaluation of language models.Transactions on Machine Learning Research, 2022.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Holistic evaluation of language models.Transactions on Machine Learning Research, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.397511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:38.517906Z digest=sha256:369455fc6ad921ce45c2436a474c7e80d4808073f1a907fd78fd46f3d257090c

Observation 50e6751b-e4e9-4fd8-b1ca-b844c8542f33 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Truthfulqa: Measuring how models mimic human falsehoods

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:38.666791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:38.666791Z digest=sha256:1b08da274baa42c4ec7b824dd49bd5c8704a428efe2c645f45a726f3ef353c0a

Observation 34406dc8-cbb5-4a5b-a9e9-46bc8580c585 · outbound

This paper cites Best practices and lessons learned on synthetic data.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Best practices and lessons learned on synthetic data

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.205529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:38.854417Z digest=sha256:585f5e40f8107028dbe763c5ba1528ffee9c38994c51ec1860cbf6db4d9cd123

Observation 080387ab-ccae-423b-adc7-7451384d0e5d · outbound

This paper cites Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.014964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.014964Z digest=sha256:b679100bd6c77c84818afd07720306ac45f8ae1144749a977b9ee268bb19cb35

Observation ec6e65cc-1c8c-4b78-b252-8c1ea0f7441f · outbound

This paper cites Adaptive labeling for efficient out-of-distribution model evaluation.Advances in Neural Information Processing Systems, 37:70981–71003, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Adaptive labeling for efficient out-of-distribution model evaluation.Advances in Neural Information Processing Systems, 37:70981–71003, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:46.853533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:39.157235Z digest=sha256:64b8199589ce512ea91943e3e267bc64e20ad9af63512b0885d85d55721bffbd

Observation 93bb72f4-c315-4cbb-9533-827ac636cf70 · outbound

This paper cites Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.313256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.313256Z digest=sha256:eb576841ca9e7808c9833203a6a0e1b8274044f26a258f29f03efbc00d40bb94

Observation 4c594c94-5c1a-42e8-8e81-dcd7a06e0923 · outbound

This paper cites Enhancing reason- ing capabilities of llms via principled synthetic logic corpus.Advances in Neural Information Processing Systems, 37:73572–73604, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Enhancing reason- ing capabilities of llms via principled synthetic logic corpus.Advances in Neural Information Processing Systems, 37:73572–73604, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:46.425652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:39.451503Z digest=sha256:c0d0c14680496dabd62074260b095b129e143d743fa43af401c40ccfb2e8a921

Observation 68588ff6-2b66-43e7-8058-4dc7072c5046 · outbound

This paper cites Caps: Collaborative and private synthetic data generation from distributed sources.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Caps: Collaborative and private synthetic data generation from distributed sources

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:46.206363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:39.570494Z digest=sha256:10718d7412522e21fc8331450076af8cc0a9ecab17ab3805bfd776aa5cd6cb5e

Observation c185db6c-5de2-4728-b983-20224c9e5697 · outbound

This paper cites Humanity's Last Exam.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Humanity's Last Exam

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.704349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.704349Z digest=sha256:4a6c2b5e8b07b9e8ed4f210449cccf961a6d785a3ec8479182ef2c6347e1f080

Observation 1440dcac-f8d7-4683-9148-2dba3f03f29a · outbound

This paper cites Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.843361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.843361Z digest=sha256:975425c3807d92f4bfc70ff7267419fb38262deb7e72dfbcaa3c7f73b4f02ff2

Observation 4b304ac3-1d89-486f-a811-dd6974a27604 · outbound

This paper cites Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.949056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.949056Z digest=sha256:41872e2f86a27ae64b2c863c8cb3deb594d3c4a097243d4028b107f3d3f026db

Observation a82e3d99-1f46-439c-93e6-25414521ab33 · outbound

This paper cites Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold.Advances in Neural Information Processing Systems, 37:43000–43031, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold.Advances in Neural Information Processing Systems, 37:43000–43031, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.030301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.030301Z digest=sha256:b23cf65f3c0cabec0abe72aea30f308450125b836baa600ca5fbc45ec9e2e17a

Observation 2d635b6f-aa87-4f42-b0eb-f251f2f356b5 · outbound

This paper cites Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:03:42.954865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:40.130964Z digest=sha256:0b5daac9f504843f5f2d2873277f19b71e77fe4c5dff3606aaf21f0d85c2f9c6

Observation a4ceafbe-6bbf-488b-b76a-aeffe7a8c783 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Gemma 2: Improving Open Language Models at a Practical Size

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.248109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.248109Z digest=sha256:25f6a22e6c9d85c6d93c66ee0fcc7c5823a50c6b85e60f070b3b5b276a1a7a7b

Observation 83c22784-02cf-4af9-8e02-6683f78e080e · outbound

This paper cites Measuring general intelligence with generated games, 2025.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Measuring general intelligence with generated games, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.350204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.350204Z digest=sha256:0a92447bba770aa01dd3158152413a6d87b32fbd5fadb0b1705c5e4a566a8072

Observation 287aab16-f8da-4da7-aaa2-a1a53c8a67fe · outbound

This paper cites Cuts: Customizable tabular synthetic data generation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Cuts: Customizable tabular synthetic data generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.994917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:40.455814Z digest=sha256:10c6214b5ff0f3d75ea4a55d13d24f4279219e29e7608398b4c3ba157dee31f3

Observation fa9c1b34-5dae-480e-8558-4dea15fe50b3 · outbound

This paper cites The Power of LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data The Power of LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:03:42.726856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:40.514127Z digest=sha256:4ac0a2e03976b70104843733cb85e58081fc9cac5838cc574d76412567327146

Observation 12ecc956-b073-4191-9415-7a6d530c350f · outbound

This paper cites Glue: A multi-task benchmark and analysis platform for natural language understanding.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Glue: A multi-task benchmark and analysis platform for natural language understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.805852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:40.582255Z digest=sha256:5d3fce4763d4ab7b1ff12fac1172c8094044103dc468cba7f238b2cc3284bfde

Observation e5220f4a-7129-497a-b745-4f8170da5201 · outbound

This paper cites Can language models solve graph problems in natural language?Advances in Neural Information Processing Systems, 36:30840–30861, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Can language models solve graph problems in natural language?Advances in Neural Information Processing Systems, 36:30840–30861, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.621047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:40.660003Z digest=sha256:dc5d54f1211c47f5fb7a61f837164969089d3add4a878bb506f2f71d04150e23

Observation b7cc40a9-d296-4b25-ac97-2c099f57c4c4 · outbound

This paper cites Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.358576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:40.729147Z digest=sha256:7e90a755f0a920689af341ebf4a3595bb8364bd3932ded8a5b66b59daf24e767

Observation 1064b364-68fa-45f1-9753-a3f9176832e9 · outbound

This paper cites Self-instruct: Aligning language models with self-generated in- structions.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Self-instruct: Aligning language models with self-generated in- structions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.827284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.827284Z digest=sha256:5ad4250b13d72960147e7e8110946589140972aeb2a9bfe8a3e37e7b746eb4b4

Observation f672e44b-eda7-4f67-8007-da3b6ac579dc · outbound

This paper cites Pre-training with synthetic data helps offline reinforcement learning.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Pre-training with synthetic data helps offline reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.214965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:40.925782Z digest=sha256:0a0724b0ab080e7227d93021152d9b93eb17fa5112f7bcb24c8513b5eef9fb97

Observation 2ae4d87d-0c52-43f4-8558-51801c4ca646 · outbound

This paper cites Rocketeval: Efficient automated llm evaluation via grading checklist.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Rocketeval: Efficient automated llm evaluation via grading checklist

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.017392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:40.994839Z digest=sha256:618c963488fe445137d278bc33847c9cc7c7de6bb47945c403a1ab8364088dc8

Observation a5a81aa4-3368-492c-8936-94765473778d · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.104348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.104348Z digest=sha256:2cb543ecb2b8fb649ff242f5fb4ccae8966963d2b89d07836fb46835d0fbf26e

Observation 5a54daff-d92d-42da-95e8-c21d034ee6d0 · outbound

This paper cites Differentially private synthetic data via foundation model apis 2: Text.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Differentially private synthetic data via foundation model apis 2: Text

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.837302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:41.220080Z digest=sha256:3c1a2b07e18a389e4e27c6971d9230f929f5caa19bf3ebddeb03cafae277a86d

Observation d570d6b4-12d0-49a0-b1d7-6ca0dab74574 · outbound

This paper cites Automating dataset updates towards reliable and timely evaluation of large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Automating dataset updates towards reliable and timely evaluation of large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.651043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:41.331758Z digest=sha256:c7b2358d5fa817fc6fed3fb25abe9a01b5dfa3fdaf8e48ab132b6b8a647c96da

Observation 22db04ed-bee3-4fd8-a2eb-5c6de61b3b6e · outbound

This paper cites xfinder: Large language models as automated evaluators for reliable evaluation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data xfinder: Large language models as automated evaluators for reliable evaluation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.481634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:41.402978Z digest=sha256:c3cf63312aa089776c8c13a7ca0509fc45fe4c3c80ef15784da4c712cd6cfb42

Observation 36e675ce-81da-4eaf-a022-33e188051eeb · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.541623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.541623Z digest=sha256:253d90bbf6d8d96772064efbaf2b54051553536a304e60d78fed06deceaf0ed4

Observation e5df9015-eeba-493d-a863-5c1d497aa6a3 · outbound

This paper cites EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.647357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.647357Z digest=sha256:ba4d1d8707f2410093270e45923529b1fbcb7b051f8063ea41e117867423f0a7

Observation 421897cc-4d5d-4cfe-90cd-ce34ec6505a1 · outbound

This paper cites Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge Injection.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge Injection

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.749831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.749831Z digest=sha256:6ccf380cbc19f78d6b350aafdef193320339dcfe7b94bced16950381b791b249

Observation 7fff0ab4-4543-462f-9808-82a749400c7f · outbound

This paper cites Wildchat: 1m chatgpt interaction logs in the wild.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Wildchat: 1m chatgpt interaction logs in the wild

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.304019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:41.852941Z digest=sha256:970f962642363b613e9f6df09c8543b7623ec01ebca55c22b757b025353e121e

Observation 523b8033-85f6-4e45-b746-d0638c7ad679 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.967706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.967706Z digest=sha256:a39175ed4f02572020e701fe0d021e171e64fb5943ab449d94ba3f742df187e8

Observation c2b40608-3786-4bce-bc64-3ffb03f02354 · outbound

This paper cites Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.147466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:42.068582Z digest=sha256:77c1c71f23d29c5638405b9c0473a491478d034e841954d790e5f6df0b44b322

Observation 3f59b8ae-51c8-4f18-bccd-539b6abfc1b2 · outbound

This paper cites Sotopia: Interactive evaluation for social intelligence in language agents.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Sotopia: Interactive evaluation for social intelligence in language agents

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.981981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:42.161531Z digest=sha256:4f52fde403c1820c71c7504a502075c9ac8c60a1896dd92c4dcd8345904e146e

Observation 49d8ee22-76eb-4750-a158-8d9a5de1bc00 · outbound

This paper cites Dyval: Dynamic evaluation of large language models for reasoning tasks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Dyval: Dynamic evaluation of large language models for reasoning tasks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.812704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:42.262166Z digest=sha256:3fb28097f50612262d67b1d67cdcc81d605eae2c00879799cd53160058cc4371

Observation 170741c2-449c-4e06-b467-adda07e6e7b4 · outbound

This paper cites Dynamic evaluation of large language models by meta probing agents.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Dynamic evaluation of large language models by meta probing agents

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.645456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:42.372390Z digest=sha256:509ad0ed4bec0597ece3e8c00fbe0aacadcae282027d078917d6575eee4c3d3f

Observation 09208314-afdc-414f-8d4a-dc9ae8757df8 · outbound

This paper cites Top Secret.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Top Secret

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.487172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T12:03:42.484850Z digest=sha256:1cabc86bd7afae665adb2c69ea8b3cd6229ecf9f854b78e50b68a45df3698bde

Pith citing papers

No inbound Pith citation observations are available.