Pith. sign in

Paper Citation Record · LEDGER

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation

As of 8 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2505.21941.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21941 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:22:43.743614Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:51:32.303810Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:51:34.729631Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 41a1d5bd-d8e3-4943-95d5-8355e9cfffa8 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.332471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.332471Z digest=sha256:376d7a2d6434350bef79581c65965f9aa8b0cfa2cc612c96127dccc0dbeed59a

Observation 626451ef-7109-41b0-9c61-96cb3dc9d052 · outbound

This paper cites Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.440527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.440527Z digest=sha256:a9b3ccf85573d7ab8caa02e886e3faee26dba41cf92028ba478b18beddaaac21

Observation 8ef982de-4648-49a6-ab6c-3a7bbb57e0f6 · outbound

This paper cites The Llama 3 Herd of Models.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.484105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.484105Z digest=sha256:4c3396dfca5e85b4ec5add954959e0e18d1ab1c5faafc1321eda47a55e0d01aa

Observation 95536eae-6b54-4d11-bdc8-9f8ea417005d · outbound

This paper cites M-RewardBench: Evaluating Reward Models in Multilingual Settings.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation M-RewardBench: Evaluating Reward Models in Multilingual Settings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.658486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.658486Z digest=sha256:2fde4a3ffe62e152e1661d3a0956d380943a647dc76bef6c038da4e2a7eb70bf

Observation 1c266c9c-4760-4735-802f-c143fd4fd2f5 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation RewardBench: Evaluating Reward Models for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.818840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.818840Z digest=sha256:eefbde13620a0b289e7b87f56e28ad854a9e6d98eb7c56adcff03b58989969b7

Observation 34f63ecf-aaac-429c-8775-beb645de9a9d · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.911718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.911718Z digest=sha256:4109d98c7a07e590c8991769e7b802a0a1629b78fbd211b5b2b48b160f49713c

Observation 21bd6289-0180-41c4-8cd2-23152b44ea54 · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.998884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.998884Z digest=sha256:486e788ad92f2c7d77fe38db08793cfa0906e755a51de7d329ccda0a54f8a024

Observation 726a627a-496b-4340-bccb-86ca399fcdfc · outbound

This paper cites s1: Simple test-time scaling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation s1: Simple test-time scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.121015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.121015Z digest=sha256:0bf32f7d6f458ac7f94b544258f91358a07b4a488e22d95a1bce81d2828debcd

Observation 863fd582-d7e9-48c3-ac1c-95901518b98c · outbound

This paper cites Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.281232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.281232Z digest=sha256:5612e429a832632975e2195a1301642bd30d8cdcfb9f25a3d0dc0fc2f1f5a957

Observation 5a4ddb01-3c61-4e32-a477-e4e04a915e8d · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.393276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.393276Z digest=sha256:f788bf19370d9aecc49099f01550217c04dc87f2fa254d6b81a726549ae21dc6

Observation 822cb2f8-5903-45d1-8226-9d669659ca0f · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.499614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.499614Z digest=sha256:a717f316c673135106daf9ebd2f9c074109e3aa9f08af45cbfb8e87c40a7afad

Observation 5115f042-f7de-44e9-8033-783d0859fe69 · outbound

This paper cites We abbreviate these names in figures to make them more readable.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation We abbreviate these names in figures to make them more readable

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:22:44.453195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:43.743614Z digest=sha256:548b62f97cd4a65a160412ab20d6ecddd79719947208db001249aca957cde57e

Observation ebcbb6ad-83cf-4c9f-8494-2c995f9ac964 · outbound

This paper cites A Other Experimental Details We provide some additional experimental details regarding the models and the hyperparameters used in our evaluation.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation A Other Experimental Details We provide some additional experimental details regarding the models and the hyperparameters used in our evaluation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:44.624366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:22:43.620849Z digest=sha256:00025ae9f13a7083c4c1082d8a83e3fcb1aee55c10746c2886a5ea66013001b2

Observation 1b23aeb3-3861-484d-8f0f-23ae9a909a79 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.270421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.270421Z digest=sha256:b836e22ce2dbd9ec864bc1a5c30ee00c2af1b98fd56c5097a8bbb43eefd6422f

Observation d4a2c687-b1d7-4733-8c38-14a52b70dc62 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.566750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.566750Z digest=sha256:03a95b7c7407c06a94fddf3049f28a5b69691f229f1a48f100a5a8b9e422e500

Pith citing papers

Observation efef82c8-fc47-4fb1-b1ee-f585b8daaa35 · inbound

When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs cites this paper.

When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:51:34.734802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:51:32.303810Z digest=sha256:fa3369b8084c02734e6ad72b05a813d5ddcd355e02ec0a0092a6563fdeb6c684