Pith. sign in

Paper Citation Record · LEDGER

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation

As of 15 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 1 inbound Pith citation observation for arXiv:2505.21941.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21941 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:22:43.743614Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:51:32.303810Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T22:51:34.729631Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved13
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 41a1d5bd-d8e3-4943-95d5-8355e9cfffa8 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.332471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.332471Z digest=sha256:0c35bdf9ab1d0fdcf418758f6d66c67d7e83cae2af24c26c913dd79fe675bd00

Observation 626451ef-7109-41b0-9c61-96cb3dc9d052 · outbound

This paper cites Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.440527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.440527Z digest=sha256:00cbeff5168379d1dd9ed85e5175a84dd668823bef295be05a274533b9b90e21

Observation 8ef982de-4648-49a6-ab6c-3a7bbb57e0f6 · outbound

This paper cites The Llama 3 Herd of Models.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.484105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.484105Z digest=sha256:70e0812cc70a77c72d95437121051fb5ebf0f0176c2b598810539cc3049a9144

Observation 95536eae-6b54-4d11-bdc8-9f8ea417005d · outbound

This paper cites M-RewardBench: Evaluating Reward Models in Multilingual Settings.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation M-RewardBench: Evaluating Reward Models in Multilingual Settings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.658486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.658486Z digest=sha256:f0d3581ac90e3730785751469f02e3971039bb9b7b3879d5e1bb5482d094b499

Observation 1c266c9c-4760-4735-802f-c143fd4fd2f5 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation RewardBench: Evaluating Reward Models for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.818840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.818840Z digest=sha256:5f241155b6e97e59f1b8e1fc2547a6b128bb273197f42d1a2beea905af6fff2e

Observation 34f63ecf-aaac-429c-8775-beb645de9a9d · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.911718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.911718Z digest=sha256:0abd1ed1743a6f1a9a5660d554bd444759be53235ce28972edb78378972434d6

Observation 21bd6289-0180-41c4-8cd2-23152b44ea54 · outbound

This paper cites Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Uncertainty-aware Reward Model: Teaching Reward Models to Know What is Unknown

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.998884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.998884Z digest=sha256:73311fa50bc75c567e1d5e4861944449c712494d83fee574f44f66a580673282

Observation 726a627a-496b-4340-bccb-86ca399fcdfc · outbound

This paper cites s1: Simple test-time scaling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation s1: Simple test-time scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.121015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.121015Z digest=sha256:8beacdee6b9f8a1368f59f9103cde7279c0122c930a1880ef72cfb2d44149686

Observation 863fd582-d7e9-48c3-ac1c-95901518b98c · outbound

This paper cites Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.281232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.281232Z digest=sha256:faf25e84d115a507e9de958fdbb25f2c2f7af42e3c15a5e8bb541d8213549f4e

Observation 5a4ddb01-3c61-4e32-a477-e4e04a915e8d · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.393276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.393276Z digest=sha256:d9946805dd4759db07208451a7f06447c9e242016107a40d05fdea9c32b04195

Observation 822cb2f8-5903-45d1-8226-9d669659ca0f · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:43.499614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:43.499614Z digest=sha256:29382ff0035f8fc2b92d2bbde9e65e53137f792f344cbf98bdafdcf20066e709

Observation 5115f042-f7de-44e9-8033-783d0859fe69 · outbound

This paper cites We abbreviate these names in figures to make them more readable.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation We abbreviate these names in figures to make them more readable

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T13:22:44.453195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:43.743614Z digest=sha256:6c763fa58a660ecf2c138f887662879a31f3a855665263ef3e793b879b635643

Observation ebcbb6ad-83cf-4c9f-8494-2c995f9ac964 · outbound

This paper cites A Other Experimental Details We provide some additional experimental details regarding the models and the hyperparameters used in our evaluation.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation A Other Experimental Details We provide some additional experimental details regarding the models and the hyperparameters used in our evaluation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:22:44.624366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:22:43.620849Z digest=sha256:1d7e610a8a2a71a39e84578efe7d203f742c52e5bf801c23615f22cc5d03dbb6

Observation 1b23aeb3-3861-484d-8f0f-23ae9a909a79 · outbound

This paper cites Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation Smaller, Weaker, Yet Better: Training LLM Reasoners via Compute-Optimal Sampling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.270421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.270421Z digest=sha256:7d68c888be15b4b2703b805094345deaf4791ae51e183515e4cd58f51d36f313

Observation d4a2c687-b1d7-4733-8c38-14a52b70dc62 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:42.566750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:42.566750Z digest=sha256:56e79eeadf5cbd22614290f25406c1501d31b603b1ee14f3183599af9bef7c0c

Pith citing papers

Observation efef82c8-fc47-4fb1-b1ee-f585b8daaa35 · inbound

When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs cites this paper.

When Life Gives You Samples: The Benefits of Scaling up Inference Compute for Multilingual LLMs Test-Time Scaling with Repeated Sampling Improves Multilingual Text Generation

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:51:34.734802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T22:51:32.303810Z digest=sha256:6e3f65e1fa136ac61d31facf2d444570043ece04161ef414b428c7a3d7b9dffc