Pith. sign in

Paper Citation Record · LEDGER

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation

As of 19 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2505.12058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12058 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:44:49.143695Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ec16231-ed40-478d-a0a2-f7240c370cc4 · outbound

This paper cites A Comprehensive Survey of Text Classification Techniques and Their Research Applications: Observational and Experimental Insights.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation A Comprehensive Survey of Text Classification Techniques and Their Research Applications: Observational and Experimental Insights

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T20:44:49.320389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:49.090513Z digest=sha256:ff8b0e319429de5d70b6b158d70eba21096e0ef8a7fa762be541c53a87c117fd

Observation c70b4215-fad7-4e7f-bb3f-6c719db63b71 · outbound

This paper cites Jaikanth J.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Jaikanth J

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.103034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.103034Z digest=sha256:03c236b8f11302ce09e9c23f594a19a039804b77198883f959f1dc9fca73ce2c

Observation 6c8a1e29-fa53-45aa-abb2-22f6455cfa9e · outbound

This paper cites an unresolved cited work.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:44:49.352402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:49.106085Z digest=sha256:aaaf694c8f07210caa8b2d951fec3856f47f2caaacbd48521f04f6be8855b42d

Observation de949d48-5437-433d-81bc-bc3a912d0dae · outbound

This paper cites Generative AI and Large Language Models in Language Preservation: Opportunities and Challenges.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Generative AI and Large Language Models in Language Preservation: Opportunities and Challenges

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:44:49.244691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:49.109466Z digest=sha256:73704dc3078101c1eada2ba0bf8494125bc11defc1ec91ed4dc3f30d6b1abb39

Observation d2d73466-879d-4ae6-908b-cae8e3966e9f · outbound

This paper cites Harnessing large-language models to generate private synthetic text.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Harnessing large-language models to generate private synthetic text

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.113008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.113008Z digest=sha256:b581cfc88dab442607c0d49719f99721c5c1f2a19268aab945a34789224fe193

Observation cbc7993b-b2e9-4cc6-a779-eb973ccdb27a · outbound

This paper cites Vladimir I.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Vladimir I

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:49.341057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:49.117448Z digest=sha256:fe6106f321e1fb5c8993d15f6c6d4794fc4105b94ee07fd057d0f8320e313f1a

Observation 7befa1f8-c309-40ba-9a43-c5188c9435d8 · outbound

This paper cites On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.128978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.128978Z digest=sha256:535ee9f693df4437800ab47f5c6601f013238061ddca0e86bd3d5e70cdf27f83

Observation d40aa916-bdc1-4071-838d-1464500e4dee · outbound

This paper cites Dataset Inference: Ownership Resolution in Machine Learning.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Dataset Inference: Ownership Resolution in Machine Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.132230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.132230Z digest=sha256:1eea18ade18a7f75d4a48cc70eb7eeac4e4df0e8da6fdf6c88e330e623eed216

Observation 9f93db07-3bcc-4c17-a883-30816218d60c · outbound

This paper cites Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:49.330745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:49.135686Z digest=sha256:7c24e1963e1e7ed479c609462f7b66af19c0b094339b3b816dd9d99ce2f868b7

Observation 77c98e17-d408-4441-84e7-13f7fcd1f49a · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation tinyBenchmarks: evaluating LLMs with fewer examples

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.139785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.139785Z digest=sha256:061dc9910cef8b377ca573c55fdbda9ca52bc599206bc0de71a4f0021c42c57e

Observation 235e33c2-bda0-4b2d-a0b9-c767fbf4d179 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.143695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.143695Z digest=sha256:7a3d9704b2ab753a598d1045b2a78f45125f83ae5928e8ce0d8f635ce30a7549

Observation 166db7f0-7f4d-4bf6-bdc2-f6e8ebf26969 · outbound

This paper cites Holistic Evaluation of Language Models.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Holistic Evaluation of Language Models

Reference 1966

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.120283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.120283Z digest=sha256:66f9b4ef708093187f0b032b485bd84aaf59fd2692f8988b48ef650c522f40b1

Observation e6485176-c99e-4289-9f32-f0cd467d1647 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.100163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.100163Z digest=sha256:b82873c5fc8ded0b57ffe8575a7c24101192a82d16bb2a2215a3485eba6f5c59

Observation 97ef3422-0d34-4338-86fd-75cbb5ef5f1f · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation AgentBench: Evaluating LLMs as Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.123971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.123971Z digest=sha256:5843d41bd192eb968ee769be7d3b5c55db14c9b93f27e23933c4d11608cdf12e

Observation 01174b6c-276e-48be-a948-70db6bfb2556 · outbound

This paper cites an unresolved cited work.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:44:49.363461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:44:49.094201Z digest=sha256:2cbbe803d9bd20910da3b686e32be3d5024628b8e16b94a83bcee359ef756a42

Pith citing papers

No inbound Pith citation observations are available.