Pith. sign in

Paper Citation Record · LEDGER

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation

As of 18 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2505.12058.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12058 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:44:49.143695Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ec16231-ed40-478d-a0a2-f7240c370cc4 · outbound

This paper cites A Comprehensive Survey of Text Classification Techniques and Their Research Applications: Observational and Experimental Insights.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation A Comprehensive Survey of Text Classification Techniques and Their Research Applications: Observational and Experimental Insights

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T20:44:49.320389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:49.090513Z digest=sha256:0e874f809d094a40aa7db69f9d4b5886b2e46212052013298a2fc8b7207bd4c8

Observation c70b4215-fad7-4e7f-bb3f-6c719db63b71 · outbound

This paper cites Jaikanth J.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Jaikanth J

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.103034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.103034Z digest=sha256:7df5eafc8a4793b2d0ffc93f9f729a7199860d7b54a4dccb5c0723fae3c4f396

Observation 6c8a1e29-fa53-45aa-abb2-22f6455cfa9e · outbound

This paper cites an unresolved cited work.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:44:49.352402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:49.106085Z digest=sha256:34a3acd1b8288d11d54ab9ac261ea42e2bf2e4a6b0f3ce3aa3596cb9b4c95eb7

Observation de949d48-5437-433d-81bc-bc3a912d0dae · outbound

This paper cites Generative AI and Large Language Models in Language Preservation: Opportunities and Challenges.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Generative AI and Large Language Models in Language Preservation: Opportunities and Challenges

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:44:49.244691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:49.109466Z digest=sha256:3e4b8528a19251fed8576c3fd90013daef7ee3380472c3808d9aba1ec7f2fa8e

Observation d2d73466-879d-4ae6-908b-cae8e3966e9f · outbound

This paper cites Harnessing large-language models to generate private synthetic text.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Harnessing large-language models to generate private synthetic text

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.113008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.113008Z digest=sha256:34a62476a4ad98e56786a4da0b4af9f423818565c865c05647098adfd81e416d

Observation cbc7993b-b2e9-4cc6-a779-eb973ccdb27a · outbound

This paper cites Vladimir I.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Vladimir I

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:49.341057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:49.117448Z digest=sha256:1bb8fe8197200ec546a16456421abf41bc88c0279ace54108dea7d3ef2799145

Observation 7befa1f8-c309-40ba-9a43-c5188c9435d8 · outbound

This paper cites On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation On LLMs-Driven Synthetic Data Generation, Curation, and Evaluation: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.128978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.128978Z digest=sha256:2f57a28bdeac94667329c6e29e937b319a4c1b9920f8cead8ce21b5c7d52a0af

Observation d40aa916-bdc1-4071-838d-1464500e4dee · outbound

This paper cites Dataset Inference: Ownership Resolution in Machine Learning.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Dataset Inference: Ownership Resolution in Machine Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.132230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.132230Z digest=sha256:34d386887363ce0848a60f2f215a7d8b877405c80282d3e32099da6331d6806c

Observation 9f93db07-3bcc-4c17-a883-30816218d60c · outbound

This paper cites Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:44:49.330745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:49.135686Z digest=sha256:8094d0d9922fe084ceda06a52dfbc8851effd6f82440b8e71cad01e8b324b1c8

Observation 77c98e17-d408-4441-84e7-13f7fcd1f49a · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation tinyBenchmarks: evaluating LLMs with fewer examples

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.139785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.139785Z digest=sha256:dafc1736064e1bd63c8a7b911d20508f61cf4d63065ca97a380852d5e490b370

Observation 235e33c2-bda0-4b2d-a0b9-c767fbf4d179 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.143695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.143695Z digest=sha256:965ac93160e4dce095f849ebc8cf88342a8bc31b8ce84ed4a6c67ea147c0e187

Observation 166db7f0-7f4d-4bf6-bdc2-f6e8ebf26969 · outbound

This paper cites Holistic Evaluation of Language Models.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Holistic Evaluation of Language Models

Reference 1966

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.120283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.120283Z digest=sha256:47b5047ba0e457fd1cdae88528c72c5a363581e8b4221b7b74956a27faff23ea

Observation e6485176-c99e-4289-9f32-f0cd467d1647 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Measuring Massive Multitask Language Understanding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.100163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.100163Z digest=sha256:d9ef4ffcbc5db94ac4a7a452fbc76e55e212afc9f8f822fb42d989e0455e2b6d

Observation 97ef3422-0d34-4338-86fd-75cbb5ef5f1f · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation AgentBench: Evaluating LLMs as Agents

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T20:44:49.123971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:44:49.123971Z digest=sha256:7aa3ade3f51cfa6b71505e2fcfed2b4a203f613d1e99e07eab6eb0b8c0783a07

Observation 01174b6c-276e-48be-a948-70db6bfb2556 · outbound

This paper cites an unresolved cited work.

Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:44:49.363461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:44:49.094201Z digest=sha256:0db5968ca211d310b8a2d9af68f6baa4950bca0aff53892090588e0d1c288708

Pith citing papers

No inbound Pith citation observations are available.