Pith. sign in

Paper Citation Record · LEDGER

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute

As of 12 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2608.09351.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09351 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:01:29.631887Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact7
  • verified fuzzy1
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7faa10d-4099-41b5-a650-21c53b820358 · outbound

This paper cites Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Enhancing LLM Robustness to Perturbed Instructions: An Empirical Study

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.532259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.532259Z digest=sha256:5307ec9f926ac5d76609932dcd9afc991fb36cde51323aed4d4735d34d597892

Observation 267c094b-32aa-4078-97f7-149f5ed65537 · outbound

This paper cites The Surprising Effectiveness of Test-Time Training for Few-Shot Learning.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute The Surprising Effectiveness of Test-Time Training for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.539084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.539084Z digest=sha256:17e57752eb59a76767e136c55bab1b7047e2905907e58fa47f9ce5008bc89668

Observation 274cfb02-528f-42a0-9697-9dcd0ab32c86 · outbound

This paper cites Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Text Data Augmentation for Large Language Models: A Comprehensive Survey of Methods, Challenges, and Opportunities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.545090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.545090Z digest=sha256:e6e418a6ad65f15bcae8b43192771d26efa3347a05f1d8ef2bc8a79f958df6c9

Observation 9e6fa001-a216-4e73-9165-179998869704 · outbound

This paper cites Exploring LLM Reasoning Through Controlled Prompt Variations.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Exploring LLM Reasoning Through Controlled Prompt Variations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.548200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.548200Z digest=sha256:9dab5c07a70444474541b0bf2466b5254ba5de76caaaa698fcf7efbd08ab08f4

Observation e28a2820-d83a-4870-a3e4-ad5d68cb7093 · outbound

This paper cites Teaching Large Language Models to Self-Debug.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Teaching Large Language Models to Self-Debug

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.551448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.551448Z digest=sha256:6fc72976d65992daaf47aa24b05b989921df9aa4ae2fcc2cfc0ca850666b8bc2

Observation dba996a4-ac77-43ba-b281-90b9f444d1e3 · outbound

This paper cites No Language Left Behind: Scaling Human-Centered Machine Translation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute No Language Left Behind: Scaling Human-Centered Machine Translation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.556706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.556706Z digest=sha256:03d3b133dd4c69983c31486714ffb75568ae7e1c5d5dfb90f41488493d800f27

Observation 138c76ef-b0c1-414d-9ba9-ae9aaba98965 · outbound

This paper cites Frustratingly Easy Test-Time Adaptation of Vision-Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Frustratingly Easy Test-Time Adaptation of Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.562315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.562315Z digest=sha256:09b517d7648d31f11cc9d0d0ca8c6b7316dfb02083edd9043a1f631bd1ae1551

Observation 1d6734f8-74a2-4378-b1f6-cc53c7332200 · outbound

This paper cites M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute M-Ped: Multi-Prompt Ensemble Decoding for Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:02:25.243674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.565203Z digest=sha256:57c0b4bae418e8c869ad441dc75cbd60ded0f7ecd36f51b0387c25b30b2b9e12

Observation a6e6da89-ec2d-4ea2-8465-ff9cf098849e · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Measuring Massive Multitask Language Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.568253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.568253Z digest=sha256:11f699a7ff868037d9b06ee351a999d962b10f470e252d26ab64173165b70b9c

Observation 3b7c528a-9b4b-4149-94d9-e81195fc1496 · outbound

This paper cites Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Efficiently Learning at Test-Time: Active Fine-Tuning of LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.573473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.573473Z digest=sha256:542602de7c654de23e00c6888aa27a56da78d58bfabbfadff336ad7d00baf0be

Observation b3c73951-a3d0-45ea-9c3a-2043d58f2330 · outbound

This paper cites Calibrating language models via augmented prompt ensembles.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Calibrating language models via augmented prompt ensembles

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:02:40.437350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.576659Z digest=sha256:2d4d5dcbded9184e3697a00f91d1825dded80ecae9728e9f548e7e2b9d94817c

Observation cff92cfd-bdf2-47bf-a7a1-fc5c5e406a62 · outbound

This paper cites Test-time Augmentation for Factual Probing.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Test-time Augmentation for Factual Probing

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:02:10.183656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.579337Z digest=sha256:5cc4463caeb5a95b4a37bc09a86ccb00f0529fb98cae7b2ec602b4495cb18dd5

Observation 4bd67d26-faff-47b6-bb9f-2fcad9165c9d · outbound

This paper cites Improved Text Classification via Test-Time Augmentation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Improved Text Classification via Test-Time Augmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.582126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.582126Z digest=sha256:e087df922c7b3a9afb45be9b52cf12d7108ec5e201582bce4b7a1e5274bee532

Observation c6f71437-e619-4f1d-afcf-889046fc80fb · outbound

This paper cites ReFT: Reasoning with Reinforced Fine-Tuning.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute ReFT: Reasoning with Reinforced Fine-Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.584800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.584800Z digest=sha256:62280cb511c04330ed8ca871ec071fee4eba2bfb5d92500e7b63181bdf252890

Observation 1bbf7138-9f82-4521-ba43-69ffb2f84fcf · outbound

This paper cites Humanity's Last Exam.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Humanity's Last Exam

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.590449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.590449Z digest=sha256:65f17f1f24dd07d892d11a8dc690c19fed3613e22adbce44de2d67bbd804cfbd

Observation 71adee13-e605-4423-b1c8-3a5ba9bb01dc · outbound

This paper cites Boosted Prompt Ensembles for Large Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Boosted Prompt Ensembles for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.593127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.593127Z digest=sha256:db5f5145c6f0add73d20f5fab73f49a7d050ad63d82953b3bf870caa793be24f

Observation 925c9d0a-1c63-4e28-9703-13a0e88c3936 · outbound

This paper cites When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute When Punctuation Matters: A Large-Scale Comparison of Prompt Robustness Methods for LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.595863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.595863Z digest=sha256:f3c239d54a12f0e461f1b77d99fde8411b19fd62c0843f0d5a7147fb653d1282

Observation 7e0924ec-8f7b-4c4e-99a4-539415e9c089 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.598569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.598569Z digest=sha256:1899ac7aa940f9e80fffbe3aa92c6b98b853609a4383cdc57c83deda788b6869

Observation 9ccfa0ef-9c32-4dcb-bb91-eb9c8f7731fa · outbound

This paper cites Better Aggregation in Test-Time Augmentation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Better Aggregation in Test-Time Augmentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.601164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.601164Z digest=sha256:99964c0edf1e8856d4e06dc3111fa8a173972b4ec4ef2ff3c085050dd365217d

Observation f4a00aed-f634-41ce-9bf7-0cc24d88e6c7 · outbound

This paper cites On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.607177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.607177Z digest=sha256:b8d2a2007b85b2e9d6a0c9f91894c52ae5e32fa1d5095944d97a7dc1a6c3658d

Observation f8e096ca-cdad-4c11-a066-6d3d0b38de34 · outbound

This paper cites Confidence improves self-consistency in llms.arXiv preprint arXiv:2502.06233,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Confidence improves self-consistency in llms.arXiv preprint arXiv:2502.06233,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.609878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.609878Z digest=sha256:dbbdc3a97f8a80943bba455037352d166ef4c9ae0a5d14a92c2e64a553a64794

Observation aa0e2dae-a3e0-4897-ad3b-cf91d50542d9 · outbound

This paper cites Can Large Language Models Really Improve by Self-critiquing Their Own Plans?.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Can Large Language Models Really Improve by Self-critiquing Their Own Plans?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.612699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.612699Z digest=sha256:a0f0bc34cea4d9cbbdd3ed93a54475116c662350c4bff3a9832b3632b02742ef

Observation 5391658d-cc28-48e7-8402-61fa8f74db9d · outbound

This paper cites Paraphrase types elicit prompt engineering capabilities.arXiv preprint arXiv:2406.19898,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase types elicit prompt engineering capabilities.arXiv preprint arXiv:2406.19898,

Reference 30

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:01:55.066549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.615330Z digest=sha256:82d102b4998e688a093353718ec23c7a49f05cd68e8e872112ec9847d91f56a2

Observation 999d48c4-4613-4f7c-8c55-3b4b207a152b · outbound

This paper cites Paraphrase and Aggregate with Large Language Models for Minimizing Intent Classification Errors.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase and Aggregate with Large Language Models for Minimizing Intent Classification Errors

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-11T19:01:44.750369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.617854Z digest=sha256:48cc43fa1370a7aafb0ffcf4da3fffb9fda725a05711af523eb6a4132fa9e111

Observation 74699a4a-1fd0-4095-98af-9a2eed0326d4 · outbound

This paper cites Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Training.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Your Language Model May Think Too Rigidly: Achieving Reasoning Consistency with Symmetry-Enhanced Training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.620769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.620769Z digest=sha256:431a4a4b708b1cfd0b888143ae645611daa35c51c34ebd8f069004d50c10a647

Observation 99260707-88f6-4fe3-86fc-3d63e485bc8a · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.623654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.623654Z digest=sha256:3f82d6f5bc497ffe0359d8eb17695cbae8e669b124cfd05297cfdf3e8c31c51d

Observation 21afcc6d-3ff9-4e5e-a30c-b57482d666d9 · outbound

This paper cites PREFER: Prompt Ensemble Learning via Feedback-Reflect-Refine.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute PREFER: Prompt Ensemble Learning via Feedback-Reflect-Refine

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.626588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.626588Z digest=sha256:741df5ed4e42661c1b4cca1b42d098c5e104ac9f09d6d8b36cda7d1318cbb5b8

Observation 40454ea8-570d-4c8b-97ec-ca34c0314d78 · outbound

This paper cites Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Paraphrase and Solve: Exploring and Exploiting the Impact of Surface Form on Mathematical Reasoning in Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.629219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.629219Z digest=sha256:8a3b48116182db72560732ebef279a46acb89376c685026577fb52ad199a0654

Observation 24f0135a-0377-40f5-bc64-063830293537 · outbound

This paper cites output as array.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute output as array

Reference 36

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:01:44.714610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.631887Z digest=sha256:2e714a29a6070d4e98ec64353b81660dec4ddaefba98254a9a32e291277b8f04

Observation 0139a27c-89ff-4c33-9b54-781123c46005 · outbound

This paper cites Dynamic Sentiment Analysis with Local Large Language Models using Majority Voting: A Study on Factors Affecting Restaurant Evaluation.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Dynamic Sentiment Analysis with Local Large Language Models using Majority Voting: A Study on Factors Affecting Restaurant Evaluation

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.587591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.587591Z digest=sha256:aa80da2b6f8953cea6c0a724053a6308b3e7f6d5899517ac9d9aef7b21be5b1b

Observation 9acfdb9f-c477-417a-b29c-08e840dc63b5 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.604113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.604113Z digest=sha256:06863b2282109944697d7c67f32ab95d0ef286766ca1cf9a495058bf97918a21

Observation 7c8d72bf-8bd1-451c-bc23-a2fae51aa21b · outbound

This paper cites Mirror-consistency: Harnessing inconsistency in majority voting.arXiv preprint arXiv:2410.10857,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Mirror-consistency: Harnessing inconsistency in majority voting.arXiv preprint arXiv:2410.10857,

Reference 2021

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:02:25.226343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.570968Z digest=sha256:328d2e864c304fbb9989ca0ffb0dea339dda23c7f885f047c732ec6ed59f53cf

Observation 966d0eb5-9ada-45fe-b15f-ed69797d3ebc · outbound

This paper cites Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Rephrase and Respond: Let Large Language Models Ask Better Questions for Themselves

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.559468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.559468Z digest=sha256:3b0239774f41c892ccaf2b9c280d8603566cc142af03142cdf92971747f6a429

Observation af04e517-c00e-417e-9224-a77163cee234 · outbound

This paper cites RoParQ: Paraphrase-aware alignment of large language models towards robustness to paraphrased questions.arXiv preprint arXiv:2511.21568,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute RoParQ: Paraphrase-aware alignment of large language models towards robustness to paraphrased questions.arXiv preprint arXiv:2511.21568,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.554277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.554277Z digest=sha256:f355cca4bb02e5e931545873ee38661672e59f43e25fdb467c7f56bf9d111d1b

Observation c64b5ec6-e754-426b-a1af-6a6b80e2f5e4 · outbound

This paper cites Finding the sweet spot: Trading quality, cost, and speed during inference-time llm reflection.arXiv preprint arXiv:2510.20653,.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Finding the sweet spot: Trading quality, cost, and speed during inference-time llm reflection.arXiv preprint arXiv:2510.20653,

Reference 2024

Resolution
verified exact
raw_fallback, observed 2026-08-11T19:02:40.409119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T19:01:29.542183Z digest=sha256:f1a3e6152a6b59d6e9f27739233b898ee946fd6b645388b7a0a92f4ad2f100b5

Observation 84b3b5fa-2ca9-4e13-b78f-4a9887b3e700 · outbound

This paper cites Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information.

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute Beyond Majority Voting: LLM Aggregation by Leveraging Higher-Order Information

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T19:01:29.536143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:01:29.536143Z digest=sha256:679d8cb2f8eaeaa67757ffd4b557e6b46693b1f92616301ee27550c67fd58856

Pith citing papers

No inbound Pith citation observations are available.