Pith. sign in

Paper Citation Record · LEDGER

Data Swarms: Optimizable Generation of Synthetic Evaluation Data

As of 10 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2506.00741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00741 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:03:42.484850Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact2
  • verified fuzzy33
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d016e96c-6595-4339-b0cc-cb96a723844b · outbound

This paper cites Kgquiz: Evaluating the generalization of encoded knowledge in large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Kgquiz: Evaluating the generalization of encoded knowledge in large language models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:34.567215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:34.567215Z digest=sha256:a778d042f829c0f08aa4e0baf79d459a873c3ea33f8a6d5f756c14423efb0c58

Observation 3d0e456c-a8c9-4635-8c4b-0b332bff53bd · outbound

This paper cites AutoEval Done Right: Using Synthetic Data for Model Evaluation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data AutoEval Done Right: Using Synthetic Data for Model Evaluation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:34.669899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:34.669899Z digest=sha256:ae085fe9415e8db13aff137fb948a30a1b32ca0300d51d39b20a28ad946c27af

Observation e5d3074a-1dc2-4a2c-8ffe-8c5179553f00 · outbound

This paper cites Adaptively evaluating models with task elicitation.arXiv preprint arXiv:2503.01986, 2025.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Adaptively evaluating models with task elicitation.arXiv preprint arXiv:2503.01986, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:34.895664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:34.895664Z digest=sha256:b83b5f321c7b2d633eaf7f6381212fecb0d213d2de02cf95a5e9cbb684caa7d5

Observation b921ed93-8e24-4bca-9e25-0bce68d43405 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.075359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.075359Z digest=sha256:2ff297791aa965a55f68d4f41b248a3662fc9fcd2c1807f22d5a7a7d0402616a

Observation 55493653-a93b-4192-b056-e954f47d10a2 · outbound

This paper cites Knowledge crosswords: Geometric knowledge reasoning with large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Knowledge crosswords: Geometric knowledge reasoning with large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:50.656173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:35.188332Z digest=sha256:46c7a1d17b3277ae1f53c2a3ee97c24ba64213dcd195a67e79748b63388dbb16

Observation d38a2290-f18f-47e1-a916-e0d9b553976b · outbound

This paper cites Self-Boosting Large Language Models with Synthetic Preference Data.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Self-Boosting Large Language Models with Synthetic Preference Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.341828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.341828Z digest=sha256:df0656fbe5b22b1d6b3cd5904ef2d7eb4d2743b64678b72cc49ee21c42fcce9e

Observation 013d73e7-0a88-4b2d-b3f3-9d4d03f5362b · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.Advances in Neural Information Processing Systems, 36, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Alpacafarm: A simulation framework for methods that learn from human feedback.Advances in Neural Information Processing Systems, 36, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.502201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.502201Z digest=sha256:4d7faa49256d73fd1af246a2c1571c72b4bc8261b7a368d25bf7c3f6970a5981

Observation d41b1c59-fd0e-4f48-afef-853faba12af0 · outbound

This paper cites Clas- sifying the classifier: dissecting the weight space of neural networks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Clas- sifying the classifier: dissecting the weight space of neural networks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:50.381847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:35.540623Z digest=sha256:c85e6739ffef21b53d600e55e213bfec3eec6d0a16021e3be60a0ede83c090f0

Observation e987ecb4-b7f4-4c52-b6c9-3e905dfb908e · outbound

This paper cites Heterogeneous swarms: Jointly optimizing model roles and weights for multi-llm systems.arXiv preprint arXiv:2502.04510, 2025.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Heterogeneous swarms: Jointly optimizing model roles and weights for multi-llm systems.arXiv preprint arXiv:2502.04510, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:35.613142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:35.613142Z digest=sha256:aa32dbaf33062df7aa4379812e9b1b35f4980227e8add54bfe722b09a49c59f7

Observation fd4fb5fa-c7bd-4b36-85ce-d6907ad443ec · outbound

This paper cites Model swarms: Collaborative search to adapt LLM experts via swarm intelligence.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Model swarms: Collaborative search to adapt LLM experts via swarm intelligence

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:50.181618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:35.762723Z digest=sha256:db80644ced2a33be07efde7b99b62d2275927d155ea1c32efb4200049e2b1553

Observation 6301967c-1eb0-41d1-8804-d4b16fd9704a · outbound

This paper cites Promptbreeder: Self-referential self-improvement via prompt evolution.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Promptbreeder: Self-referential self-improvement via prompt evolution

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.915732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:35.916500Z digest=sha256:a4bcd704aeef208de0c06b6788b4789c1fc6b321f6f842eb3a26d095b3b37062

Observation 07f41ea0-74e2-4f30-ae48-8840954b843f · outbound

This paper cites Open llm leaderboard v2.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Open llm leaderboard v2

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:36.045712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:36.045712Z digest=sha256:ebc324d59f71cdfed3fcc3b86522dde2b04015001af0d1763bfd7278cefc1055

Observation e1da0c0b-27ac-47ea-bb3c-1a6b8142652a · outbound

This paper cites Time travel in llms: Tracing data contamination in large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Time travel in llms: Tracing data contamination in large language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.649629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:36.184289Z digest=sha256:c91083e2650d510b78c5d8d5c8fad1266d7d705d72dcda291da7392e892df6a6

Observation 8178c412-ca1b-45f4-8c1f-4987319c36c6 · outbound

This paper cites Automated evaluation of retrieval-augmented language models with task-specific exam generation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Automated evaluation of retrieval-augmented language models with task-specific exam generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.382389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:36.332679Z digest=sha256:eb828ec05fca91ad1d607b4c1dbdfca4ba85517decc13e65d4194179775f37f6

Observation 81453eb1-4329-4599-b071-66bf8072cbe2 · outbound

This paper cites Measuring massive multitask language understanding.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Measuring massive multitask language understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:36.464445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:36.464445Z digest=sha256:c5f7c09371412787a2e0c4d5eb69bb46676360e01b8615ed87c74bd1400d234b

Observation 5b49f598-0ae1-451f-a5fa-5464a234093a · outbound

This paper cites Datagen: Unified synthetic dataset generation via large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Datagen: Unified synthetic dataset generation via large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:49.134596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:36.603676Z digest=sha256:414436a9d6acf80d14cfc98d50b7f9eeb99ddae76908c84af84aa59a94d9f819

Observation 4f27019b-e3df-4b9c-8f00-91f772bb388f · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:36.765030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:36.765030Z digest=sha256:d129122c793dd8466eb9011a1610f0a1d704f941b26046a92b3f85c6bf2908ea

Observation e55dc512-29eb-492a-8c5c-652e67ef1c96 · outbound

This paper cites Wildteaming at scale: From in- the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Wildteaming at scale: From in- the-wild jailbreaks to (adversarially) safer language models.Advances in Neural Information Processing Systems, 37:47094–47165, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.937768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:36.978535Z digest=sha256:9a4b3ca89d78f0f19bd2d266cdafed7cc5f77af583738bbd07935c5c877629f3

Observation 7002d98f-da27-43be-829b-d8edf4490dcd · outbound

This paper cites Teaching language models to hallucinate less with synthetic tasks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Teaching language models to hallucinate less with synthetic tasks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.645883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:37.116286Z digest=sha256:69e02a67e3f2efe78c0ad09f4111e63cd258e251cba95141a5b023bf9a928a9f

Observation 34ed43f9-c088-4ec6-9a3f-ff7c1b7f7e9b · outbound

This paper cites Au- tonomous evaluation of llms for truth maintenance and reasoning tasks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Au- tonomous evaluation of llms for truth maintenance and reasoning tasks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.458711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:37.251988Z digest=sha256:54e96972eb40313d9eedb0cee2cc5f13d30ae329119fddf94b61468b744a6b55

Observation 736059c6-e70f-43bd-9a8a-6dc29d8e1b03 · outbound

This paper cites Realtime qa: What’s the answer right now? Advances in neural information processing systems, 36:49025–49043, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Realtime qa: What’s the answer right now? Advances in neural information processing systems, 36:49025–49043, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:37.403279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:37.403279Z digest=sha256:9772c516fe1b9acd16dc4e14fd55beb11a4f55bf335833144a0b992dc896640c

Observation e30f5e7f-8934-4ad5-8a4c-84496bd00fd0 · outbound

This paper cites Particle swarm optimization.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Particle swarm optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:37.581123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:37.581123Z digest=sha256:06cdae6468b037d97cc813d509c17a8deebb559458497d95f1de00e9f27bbff1

Observation e7b23351-9ef6-4c55-8752-f02f950bcf67 · outbound

This paper cites Openassistant conversations-democratizing large language model alignment.Advances in Neural Information Processing Systems, 36:47669–47681, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Openassistant conversations-democratizing large language model alignment.Advances in Neural Information Processing Systems, 36:47669–47681, 2023

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:48.194529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:37.767108Z digest=sha256:c7ba3d96ff23e38a181d98ede161d9507b5755873388d443244633c9f9250c96

Observation e40cc949-e6b2-43c0-b393-e6d8927bf100 · outbound

This paper cites Eliciting Language Model Behaviors with Investigator Agents.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Eliciting Language Model Behaviors with Investigator Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:37.921485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:37.921485Z digest=sha256:d1e77dcc39b7b90192cbc87042665c2783450acb8b164c6416507b47aa7fb722

Observation f3b21320-0d42-4975-ba23-a63db57a5890 · outbound

This paper cites Autobencher: Towards declarative benchmark construction.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Autobencher: Towards declarative benchmark construction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:38.098878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:38.098878Z digest=sha256:9a3187348583270e6fe4b45440a705e4b9bc67f725efb972dafa2d4dace53c62

Observation d244d9dd-3773-4826-ad02-5471a91be604 · outbound

This paper cites Gen- dataagent: On-the-fly dataset augmentation with synthetic data.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Gen- dataagent: On-the-fly dataset augmentation with synthetic data

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.979881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:38.227102Z digest=sha256:d5d7379155e26f4e32f01ef4bbad9c58dca283e9bbcf01771a6cbde4ae205fc9

Observation 8da932b7-31ef-4c50-8c38-e62277f8c458 · outbound

This paper cites Hemm: Holistic evaluation of multimodal foundation models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Hemm: Holistic evaluation of multimodal foundation models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.710055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:38.373411Z digest=sha256:2c6e5aca08fca7a3a471005c18b2390cb3b8e412eed9bb00c25a13de103a2dd1

Observation a89be644-4aab-44c0-92d4-34e813ff1c33 · outbound

This paper cites Holistic evaluation of language models.Transactions on Machine Learning Research, 2022.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Holistic evaluation of language models.Transactions on Machine Learning Research, 2022

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.397511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:38.517906Z digest=sha256:87a18f60bc52ed09a41fd5354e6c02c91a83560d87a4aee3dd3af12e58972993

Observation 50e6751b-e4e9-4fd8-b1ca-b844c8542f33 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Truthfulqa: Measuring how models mimic human falsehoods

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:38.666791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:38.666791Z digest=sha256:1b08da274baa42c4ec7b824dd49bd5c8704a428efe2c645f45a726f3ef353c0a

Observation 34406dc8-cbb5-4a5b-a9e9-46bc8580c585 · outbound

This paper cites Best practices and lessons learned on synthetic data.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Best practices and lessons learned on synthetic data

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:47.205529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:38.854417Z digest=sha256:ebb1fe86657f6dd227a072e347facba262adf1ea72390e7890d010d50ac0a964

Observation 080387ab-ccae-423b-adc7-7451384d0e5d · outbound

This paper cites Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Fictitious Synthetic Data Can Improve LLM Factuality via Prerequisite Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.014964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.014964Z digest=sha256:d41888250ec3473fbf5adab6612839ddbaf7add3a17a68b278167ce36109fe7e

Observation ec6e65cc-1c8c-4b78-b252-8c1ea0f7441f · outbound

This paper cites Adaptive labeling for efficient out-of-distribution model evaluation.Advances in Neural Information Processing Systems, 37:70981–71003, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Adaptive labeling for efficient out-of-distribution model evaluation.Advances in Neural Information Processing Systems, 37:70981–71003, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:46.853533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:39.157235Z digest=sha256:257fd8b9dd55956bf5c7216e459337661f1535a1767fa54201d14d37ce58b365

Observation 93bb72f4-c315-4cbb-9533-827ac636cf70 · outbound

This paper cites Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.313256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.313256Z digest=sha256:80fa6ba0c1726244a0d0d4b81f6ca205994ce26606b2dbf6e0b38eda1056c46c

Observation 4c594c94-5c1a-42e8-8e81-dcd7a06e0923 · outbound

This paper cites Enhancing reason- ing capabilities of llms via principled synthetic logic corpus.Advances in Neural Information Processing Systems, 37:73572–73604, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Enhancing reason- ing capabilities of llms via principled synthetic logic corpus.Advances in Neural Information Processing Systems, 37:73572–73604, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:46.425652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:39.451503Z digest=sha256:b1d234e2b5d465d137e5fb3cce1302c9a233ef221ee6d23632fd79f58b23d9dd

Observation 68588ff6-2b66-43e7-8058-4dc7072c5046 · outbound

This paper cites Caps: Collaborative and private synthetic data generation from distributed sources.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Caps: Collaborative and private synthetic data generation from distributed sources

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:46.206363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:39.570494Z digest=sha256:77baad2398c9783ccdf558009f372e2461d4465d197d8f8c016310e1b220008f

Observation c185db6c-5de2-4728-b983-20224c9e5697 · outbound

This paper cites Humanity's Last Exam.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Humanity's Last Exam

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.704349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.704349Z digest=sha256:4a6c2b5e8b07b9e8ed4f210449cccf961a6d785a3ec8479182ef2c6347e1f080

Observation 1440dcac-f8d7-4683-9148-2dba3f03f29a · outbound

This paper cites Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.843361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.843361Z digest=sha256:975425c3807d92f4bfc70ff7267419fb38262deb7e72dfbcaa3c7f73b4f02ff2

Observation 4b304ac3-1d89-486f-a811-dd6974a27604 · outbound

This paper cites Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Explore Theory of Mind: Program-guided adversarial data generation for theory of mind reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:39.949056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:39.949056Z digest=sha256:1e78f0222948a1c1e1956cce388f1daa7f07505f2005e41a7f159d97413e857c

Observation a82e3d99-1f46-439c-93e6-25414521ab33 · outbound

This paper cites Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold.Advances in Neural Information Processing Systems, 37:43000–43031, 2024.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold.Advances in Neural Information Processing Systems, 37:43000–43031, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.030301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.030301Z digest=sha256:b23cf65f3c0cabec0abe72aea30f308450125b836baa600ca5fbc45ec9e2e17a

Observation 2d635b6f-aa87-4f42-b0eb-f251f2f356b5 · outbound

This paper cites Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Is your benchmark truly adversarial? AdvScore: Evaluating Human-Grounded Adversarialness

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:03:42.954865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:40.130964Z digest=sha256:22eaafb6aaf25bfbc053a643b28c804c93efa24bb45cd209c3e8292e70c3642e

Observation a4ceafbe-6bbf-488b-b76a-aeffe7a8c783 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Gemma 2: Improving Open Language Models at a Practical Size

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.248109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.248109Z digest=sha256:25f6a22e6c9d85c6d93c66ee0fcc7c5823a50c6b85e60f070b3b5b276a1a7a7b

Observation 83c22784-02cf-4af9-8e02-6683f78e080e · outbound

This paper cites Measuring general intelligence with generated games, 2025.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Measuring general intelligence with generated games, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.350204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.350204Z digest=sha256:0a92447bba770aa01dd3158152413a6d87b32fbd5fadb0b1705c5e4a566a8072

Observation 287aab16-f8da-4da7-aaa2-a1a53c8a67fe · outbound

This paper cites Cuts: Customizable tabular synthetic data generation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Cuts: Customizable tabular synthetic data generation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.994917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:40.455814Z digest=sha256:ec6ea62e6869bb14a2b9fa40ceea435f84ecae2f7f034d8fdadd2b6c690c9a16

Observation fa9c1b34-5dae-480e-8558-4dea15fe50b3 · outbound

This paper cites The Power of LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data The Power of LLM-Generated Synthetic Data for Stance Detection in Online Political Discussions

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:03:42.726856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:40.514127Z digest=sha256:7547fb5a4ad5ee503b5d7616e681084351c52a0f7fdb4e3c8b8b5cc9b0dccd62

Observation 12ecc956-b073-4191-9415-7a6d530c350f · outbound

This paper cites Glue: A multi-task benchmark and analysis platform for natural language understanding.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Glue: A multi-task benchmark and analysis platform for natural language understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.805852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:40.582255Z digest=sha256:467d2ef0954a6dd5221cf8269f6e99a163735732de4e312994f51bee71caf872

Observation e5220f4a-7129-497a-b745-4f8170da5201 · outbound

This paper cites Can language models solve graph problems in natural language?Advances in Neural Information Processing Systems, 36:30840–30861, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Can language models solve graph problems in natural language?Advances in Neural Information Processing Systems, 36:30840–30861, 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.621047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:40.660003Z digest=sha256:32f9fd69439224ddea9b1fcdd09c322a27c47e4fe3c02dbb87e37b1c3cafe21b

Observation b7cc40a9-d296-4b25-ac97-2c099f57c4c4 · outbound

This paper cites Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.358576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:40.729147Z digest=sha256:d0fa6f3b5728842b03095fa800f11a8673bbc9962af1686b4a83d6a995c9992d

Observation 1064b364-68fa-45f1-9753-a3f9176832e9 · outbound

This paper cites Self-instruct: Aligning language models with self-generated in- structions.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Self-instruct: Aligning language models with self-generated in- structions

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:40.827284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:40.827284Z digest=sha256:5ad4250b13d72960147e7e8110946589140972aeb2a9bfe8a3e37e7b746eb4b4

Observation f672e44b-eda7-4f67-8007-da3b6ac579dc · outbound

This paper cites Pre-training with synthetic data helps offline reinforcement learning.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Pre-training with synthetic data helps offline reinforcement learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.214965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:40.925782Z digest=sha256:806551e962503632415046bdd227c400f0f7388b9d40ce03a9fa71baf66c09d4

Observation 2ae4d87d-0c52-43f4-8558-51801c4ca646 · outbound

This paper cites Rocketeval: Efficient automated llm evaluation via grading checklist.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Rocketeval: Efficient automated llm evaluation via grading checklist

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:45.017392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:40.994839Z digest=sha256:c79ad72d70ead036b9356d0ee01177b59761064ef9abf9ce08dfbb4d5aed8dbf

Observation a5a81aa4-3368-492c-8936-94765473778d · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.104348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.104348Z digest=sha256:2cb543ecb2b8fb649ff242f5fb4ccae8966963d2b89d07836fb46835d0fbf26e

Observation 5a54daff-d92d-42da-95e8-c21d034ee6d0 · outbound

This paper cites Differentially private synthetic data via foundation model apis 2: Text.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Differentially private synthetic data via foundation model apis 2: Text

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.837302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:41.220080Z digest=sha256:f0dad02a6ffdef58925c0eb7c1429cdfc2d74cc0acf0c555695ff78b72dd226e

Observation d570d6b4-12d0-49a0-b1d7-6ca0dab74574 · outbound

This paper cites Automating dataset updates towards reliable and timely evaluation of large language models.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Automating dataset updates towards reliable and timely evaluation of large language models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.651043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:41.331758Z digest=sha256:129d4fc4056bf993928a048659849ff2fc2e26dc2e3ed394a04bcdef7277051d

Observation 22db04ed-bee3-4fd8-a2eb-5c6de61b3b6e · outbound

This paper cites xfinder: Large language models as automated evaluators for reliable evaluation.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data xfinder: Large language models as automated evaluators for reliable evaluation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.481634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:41.402978Z digest=sha256:7e451019a08e08cb6ed43900ed593509f3448743beea2051b201d706b0e12c9a

Observation 36e675ce-81da-4eaf-a022-33e188051eeb · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Hellaswag: Can a machine really finish your sentence? InProceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, 2019

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.541623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.541623Z digest=sha256:253d90bbf6d8d96772064efbaf2b54051553536a304e60d78fed06deceaf0ed4

Observation e5df9015-eeba-493d-a863-5c1d497aa6a3 · outbound

This paper cites EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data EvalTree: Profiling Language Model Weaknesses via Hierarchical Capability Trees

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.647357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.647357Z digest=sha256:ba4d1d8707f2410093270e45923529b1fbcb7b051f8063ea41e117867423f0a7

Observation 421897cc-4d5d-4cfe-90cd-ce34ec6505a1 · outbound

This paper cites Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge Injection.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge Injection

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.749831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.749831Z digest=sha256:53503e01d1f2a9ee335ba0dba2fa71b9db7a0d2b1f261a6dac2fce0b5c32bfec

Observation 7fff0ab4-4543-462f-9808-82a749400c7f · outbound

This paper cites Wildchat: 1m chatgpt interaction logs in the wild.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Wildchat: 1m chatgpt interaction logs in the wild

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.304019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:41.852941Z digest=sha256:819e655eebbab88d4a9477b952e9b5e153b8056e3aeda647aff6ca50a28b87b9

Observation 523b8033-85f6-4e45-b746-d0638c7ad679 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems, 36:46595–46623, 2023

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:41.967706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:03:41.967706Z digest=sha256:a39175ed4f02572020e701fe0d021e171e64fb5943ab449d94ba3f742df187e8

Observation c2b40608-3786-4bce-bc64-3ffb03f02354 · outbound

This paper cites Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021, 2023.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Lima: Less is more for alignment.Advances in Neural Information Processing Systems, 36:55006–55021, 2023

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:44.147466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:42.068582Z digest=sha256:ea150e3f893d001fe85e578a9e34461922d000616787453f04962a2f9b1ae506

Observation 3f59b8ae-51c8-4f18-bccd-539b6abfc1b2 · outbound

This paper cites Sotopia: Interactive evaluation for social intelligence in language agents.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Sotopia: Interactive evaluation for social intelligence in language agents

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.981981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:42.161531Z digest=sha256:b50f63e05662803e9def1ecb8e86bc64a09375e643728bd90c5283e3649b6a4e

Observation 49d8ee22-76eb-4750-a158-8d9a5de1bc00 · outbound

This paper cites Dyval: Dynamic evaluation of large language models for reasoning tasks.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Dyval: Dynamic evaluation of large language models for reasoning tasks

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.812704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:42.262166Z digest=sha256:24ed8844c57135298d6f92dbb8cfe4f5ecb8d21455ca4fefe682c2148865836c

Observation 170741c2-449c-4e06-b467-adda07e6e7b4 · outbound

This paper cites Dynamic evaluation of large language models by meta probing agents.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Dynamic evaluation of large language models by meta probing agents

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.645456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:42.372390Z digest=sha256:622da6314b0ec40cde3173dd428de33f4fd1956ded5be200047acc3d9235fc5a

Observation 09208314-afdc-414f-8d4a-dc9ae8757df8 · outbound

This paper cites Top Secret.

Data Swarms: Optimizable Generation of Synthetic Evaluation Data Top Secret

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:03:43.487172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:03:42.484850Z digest=sha256:e54384954a289ee166e17188266311e702708220522b61f8ddfd8fed8b4be374

Pith citing papers

No inbound Pith citation observations are available.