Pith. sign in

Paper Citation Record · LEDGER

Does More Inference-Time Compute Really Help Robustness?

As of 11 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2507.15974.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15974 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:26:22.907922Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T02:32:49.034790Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9982772f-ef9e-4a0b-a257-3c77f1f215de · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Does More Inference-Time Compute Really Help Robustness? Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:19.829032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:19.829032Z digest=sha256:7cce148caa3857aa054c666fd2786b4f69ada4196a23f863ffcd235fda1242d1

Observation 9cb1b559-ab18-4b7b-9095-be7da16f85f2 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

Does More Inference-Time Compute Really Help Robustness? Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.051022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.051022Z digest=sha256:ea1e9104f4e607ef3613256ff7829292b1a2964e7c34bd8097e7227a99d0c6a1

Observation 8841390c-91ff-4d0e-a924-f0ef47f607be · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Does More Inference-Time Compute Really Help Robustness? ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.117161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.117161Z digest=sha256:cc48cf462b7e2172e940f7f937e0a1d1de5da2b01d0174b380b1d5d590b2de54

Observation 8417ebd9-59b6-4aa6-b8fa-646522081d01 · outbound

This paper cites org/abs/2506.15674.

Does More Inference-Time Compute Really Help Robustness? org/abs/2506.15674

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.188675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.188675Z digest=sha256:72ee3890f6f7b6a7f44a2be8f49cf02e829f2242c7eadfe62706ebb4608696c5

Observation ce9ddaa7-5fae-4fa2-824c-2542d8430729 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Does More Inference-Time Compute Really Help Robustness? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.346907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.346907Z digest=sha256:c0bc40cfbe969a4f4dfccb6a4dcab6d2b93541ae01922bfae14312e36adb90fc

Observation a75f6f51-0fb7-481f-8422-f5fe79d0a145 · outbound

This paper cites OpenAI o1 System Card.

Does More Inference-Time Compute Really Help Robustness? OpenAI o1 System Card

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.417484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.417484Z digest=sha256:44df8421a72073445c6ca64042e4dfa82ee0fd60088b471e3aecf2ce84e4fce9

Observation 785f0dd7-84af-428d-90f6-45692600b60a · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

Does More Inference-Time Compute Really Help Robustness? SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.476585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.476585Z digest=sha256:d6fef61d355339d6f605774c3366a5c88eab82497d1a34ac573f7fc934ac6993

Observation 5d37befe-25be-47b2-9984-2d4825deed40 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Does More Inference-Time Compute Really Help Robustness? Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.542266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.542266Z digest=sha256:a9be7590f812c9b2628ab74e826195208a06ddc15a0f2fbb263b8013a7c0be5a

Observation 190ac5e0-6c50-4b7e-b6b9-841f8408844a · outbound

This paper cites H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking.

Does More Inference-Time Compute Really Help Robustness? H-CoT: Hijacking the Chain-of-Thought Safety Reasoning Mechanism to Jailbreak Large Reasoning Models, Including OpenAI o1/o3, DeepSeek-R1, and Gemini 2.0 Flash Thinking

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.640185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.640185Z digest=sha256:8b9fb4487b69c64d277d98ea83e7b859915cd2d780c13aa443feb6df44f028fe

Observation 6bb401ba-6a7d-4b8e-9227-e3c5f6ffbc8d · outbound

This paper cites START: Self-taught Reasoner with Tools.

Does More Inference-Time Compute Really Help Robustness? START: Self-taught Reasoner with Tools

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.726516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.726516Z digest=sha256:c930b4bfb122f11b459de9fff929ec0b88a16276234fcd7d350d070b8abb9d2c

Observation 95143e5e-7b84-4c41-a28c-302303f4f8ca · outbound

This paper cites Deepseek-r1 thoughtology: Let’s think about llm reasoning.

Does More Inference-Time Compute Really Help Robustness? Deepseek-r1 thoughtology: Let’s think about llm reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.812119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.812119Z digest=sha256:ae966c27ad70c1dcf4843d4e1f5bb9c2f2a184e13049b55d3c7bb20556ffb3a9

Observation dabc4faf-01cc-42a2-a2a3-0011ce77004e · outbound

This paper cites SaRO: Enhancing LLM Safety through Reasoning-based Alignment.

Does More Inference-Time Compute Really Help Robustness? SaRO: Enhancing LLM Safety through Reasoning-based Alignment

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.968631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.968631Z digest=sha256:76da88426777f4647b0ad7d71fcd1a27b87cbc53cbc2bd2dd3606061a5be5c12

Observation 47269b30-f004-490a-a72c-d33a882633a6 · outbound

This paper cites s1: Simple test-time scaling.

Does More Inference-Time Compute Really Help Robustness? s1: Simple test-time scaling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:21.053444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:21.053444Z digest=sha256:127af715f46148936afd8b5dc3a84442f4965031c7e7d783dd9457e74d882bbd

Observation ed13eb12-23ff-46ad-8414-828b4f8a9d4e · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Does More Inference-Time Compute Really Help Robustness? Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:21.122575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:21.122575Z digest=sha256:ccf8a466d83fcceeb1bfdf98a22149da6156a6aa62002758e5ef6d1f429736f3

Observation c93cf7cc-4420-49af-a43e-3dc3d2457b82 · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

Does More Inference-Time Compute Really Help Robustness? R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:21.218421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:21.218421Z digest=sha256:42e8ec413f1e523e3dd9f78b4c9c806eff6082c8c823597a5347cd0aa024923d

Observation 41bb8317-881a-4f74-93a2-718796aefabe · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Does More Inference-Time Compute Really Help Robustness? The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:21.270522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:21.270522Z digest=sha256:5d6f498be83101c29ea382106e028a6b73ed28bda7d25e2a7b6df5e42f43040c

Observation c2fb6f89-a64b-4de4-8f05-950781929a9e · outbound

This paper cites Safety in Large Reasoning Models: A Survey.

Does More Inference-Time Compute Really Help Robustness? Safety in Large Reasoning Models: A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:21.358800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:21.358800Z digest=sha256:8c1ff5ece6f4b07e211defe9391ed9fa20df5026f31fbab7fd842dad5ec62936

Observation 157d35bb-6d66-4e12-92df-7ea7156ebf16 · outbound

This paper cites Star-1: Safer alignment of reasoning llms with 1k data.

Does More Inference-Time Compute Really Help Robustness? Star-1: Safer alignment of reasoning llms with 1k data

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:21.479404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:21.479404Z digest=sha256:4ae0945921b5ce8a34087165510ab8186a8bc0756d074e7b9f7870c99d0de752

Observation c7898956-3875-476d-9020-8fd8c11592cb · outbound

This paper cites Effectively Controlling Reasoning Models through Thinking Intervention.

Does More Inference-Time Compute Really Help Robustness? Effectively Controlling Reasoning Models through Thinking Intervention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:21.597162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:21.597162Z digest=sha256:3f0c759d6816586c9b16b8406ab89015e79a9b5fa181b4fe7e92b09b717acbf4

Observation 402c6651-5ef4-42f9-9eaa-dfbd928f89b0 · outbound

This paper cites Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models.

Does More Inference-Time Compute Really Help Robustness? Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:21.750058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:21.750058Z digest=sha256:07afc20901792ee1719c3c15746ae3589128766248b45de46a87e5a14d752dfd

Observation 333ca84c-6509-48c8-a1da-bbe269f53930 · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

Does More Inference-Time Compute Really Help Robustness? SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:21.912851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:21.912851Z digest=sha256:fdc8b8499329a41c02625e352a66147f9b3e5e050be9655c29a791c08a59f3fe

Observation 16badb67-6dc0-4516-99f2-d6926619ca6a · outbound

This paper cites Qwen3 Technical Report.

Does More Inference-Time Compute Really Help Robustness? Qwen3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:22.109737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:22.109737Z digest=sha256:65773bdf2e81663e32ca6306b3a134ce666208f0ae613d8507b7a995a6f2edf2

Observation a9d9edf9-7a4c-42c1-a5b5-0a423b16bca3 · outbound

This paper cites A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos.

Does More Inference-Time Compute Really Help Robustness? A Mousetrap: Fooling Large Reasoning Models for Jailbreak with Chain of Iterative Chaos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:22.339685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:22.339685Z digest=sha256:3fc2a6ad4b50373a40911f1293152ebd5fe747321642900c3226a597c66a8e0a

Observation 8404eecd-d58e-467f-ad4d-29ee522e8a54 · outbound

This paper cites Trading Inference-Time Compute for Adversarial Robustness.

Does More Inference-Time Compute Really Help Robustness? Trading Inference-Time Compute for Adversarial Robustness

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:22.432003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:22.432003Z digest=sha256:2667d47889a3d5a5a2b40f64c154cd1b667c8381133e843681bc20868caf4368

Observation 14fa33a6-25c9-4749-b8f2-378c165ee02b · outbound

This paper cites RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability.

Does More Inference-Time Compute Really Help Robustness? RealSafe-R1: Safety-Aligned DeepSeek-R1 without Compromising Reasoning Capability

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:22.599759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:22.599759Z digest=sha256:675826de9a5860c95b1673254e2c26e401465180c87bb48d988153e9ea05f657

Observation 4a9fc1eb-2807-4103-96d4-9595684e0300 · outbound

This paper cites Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models.

Does More Inference-Time Compute Really Help Robustness? Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:22.673177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:22.673177Z digest=sha256:191f9446d61f519c3bd84282116ab09f45fab246ed9be83f2b080f2f83ad679f

Observation 1d5034c9-425f-4991-a90b-d3e295c21f02 · outbound

This paper cites The hidden risks of large reasoning models: A safety assessment of r1.

Does More Inference-Time Compute Really Help Robustness? The hidden risks of large reasoning models: A safety assessment of r1

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:22.727553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:22.727553Z digest=sha256:f035752c00da8ee7de322331276835699319bf76f6aa2a2f136a0b974785e9db

Observation 5436148c-4927-40f6-891f-611ab62ed35c · outbound

This paper cites 13 Preprint.

Does More Inference-Time Compute Really Help Robustness? 13 Preprint

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:24.220754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:26:22.800518Z digest=sha256:ee0e786b28fb19b3e493c799b7b8967e81e50f6ab7eba2fb4f71af674dd05fe5

Observation 819db80c-6f7a-41c8-95cc-43838aab726f · outbound

This paper cites The main instruction, associated data, low-priority query, and witness are shown.

Does More Inference-Time Compute Really Help Robustness? The main instruction, associated data, low-priority query, and witness are shown

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:24.205877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:26:22.856307Z digest=sha256:8fcd39288eccec862dca67216a369349f9bfd444e44f49346417f9e59749b60f

Observation af8b48a9-fb0f-4028-879f-2730257b0277 · outbound

This paper cites The system instruction and malicious user prompt are shown.

Does More Inference-Time Compute Really Help Robustness? The system instruction and malicious user prompt are shown

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:26:24.190719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:26:22.907922Z digest=sha256:3d4289fa54b7af088776a8e8fa4fea8b1245b1762bab4743e05d4f22f110a056

Observation 9d527b4d-6f21-4db1-ba27-14c2408157a1 · outbound

This paper cites Theoretical guarantees on the best-of-n alignment policy.

Does More Inference-Time Compute Really Help Robustness? Theoretical guarantees on the best-of-n alignment policy

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:19.894110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:19.894110Z digest=sha256:d1b9c2bc1c4cc54066e7f59afc59970ba07361fe2eb5ea774e5aa2a345d8a7ce

Observation 30e9e55d-4c40-43f6-9cfd-cc8d56e62fdf · outbound

This paper cites ISBN 9798400702600.

Does More Inference-Time Compute Really Help Robustness? ISBN 9798400702600

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:20.262437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:20.262437Z digest=sha256:1a7bc7726f9622f7e83859cf17c4cadaacfcbe1c82cece8a214ca968be44f1d7

Observation 0b2d3b75-6674-4780-b91a-a1ebd25c0e05 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Does More Inference-Time Compute Really Help Robustness? Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:19.981630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:19.981630Z digest=sha256:e77114da99ef45e4cd6a97130790a6fe921f4c1cc9180432233ee369ed9e0dc7

Observation 8868e2aa-1fe6-41e6-8068-78143ef4ebae · outbound

This paper cites Phi-4-reasoning Technical Report.

Does More Inference-Time Compute Really Help Robustness? Phi-4-reasoning Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:19.738283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:19.738283Z digest=sha256:415f9c34164cf09c95b8c73fdcd60e8a017ad0db9d863fa6d6bb54b885b59331

Pith citing papers

Observation 026b3643-7260-481c-8717-988f3de28d2c · inbound

Adaptive Probe-based Steering for Robust LLM Jailbreaking cites this paper.

Adaptive Probe-based Steering for Robust LLM Jailbreaking Does More Inference-Time Compute Really Help Robustness?

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:33:54.907156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T02:32:49.034790Z digest=sha256:2d8fb14442d31169abe65603fc860f5fbf4b4a935894b73dcec2f6a9065ad5f9