Pith. sign in

Paper Citation Record · LEDGER

MetaSC: Test-Time Safety Specification Optimization for Language Models

As of 10 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 2 inbound Pith citation observations for arXiv:2502.07985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07985 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T11:15:45.016767Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T08:17:10.481202Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 2022d310-a28d-4ce0-8134-f1ba8f434780 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

MetaSC: Test-Time Safety Specification Optimization for Language Models Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:44.972604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:44.972604Z digest=sha256:e8cbb36b4749a3b8f7afabd23c366d941367a004974703081661d4a76bea2f10

Observation 599b5898-dc2b-4146-843a-b5fcd295d9aa · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

MetaSC: Test-Time Safety Specification Optimization for Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:44.977046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:44.977046Z digest=sha256:91bd5fd593e8f2c447b78ed7a5ef3aed6bf61631241ad5da80c3966074de210f

Observation d282f1f9-1687-410f-9a03-37dda913f7c1 · outbound

This paper cites The Llama 3 Herd of Models.

MetaSC: Test-Time Safety Specification Optimization for Language Models The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:44.984746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:44.984746Z digest=sha256:99da58fc12cb836557951e91ebabb0f2eb0e454647a8edf5dacd9371eb32c5cc

Observation 2e7d7c88-ac9b-42c2-9fa1-c29ab2ee356b · outbound

This paper cites Fight Back Against Jailbreaking via Prompt Adversarial Tuning.

MetaSC: Test-Time Safety Specification Optimization for Language Models Fight Back Against Jailbreaking via Prompt Adversarial Tuning

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-08T11:15:45.087589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T11:15:44.988029Z digest=sha256:e2203ec8cf9c2107f5674287d17f46ccb0159d5ef6e8502b045cec276635da99

Observation 4b5db248-9976-4374-b751-22a3d133c73a · outbound

This paper cites SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks.

MetaSC: Test-Time Safety Specification Optimization for Language Models SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:44.990935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:44.990935Z digest=sha256:4a7dae45f5a7e6b8fd4886a741b56808138610f90d06344e38de6821e074a52f

Observation ec7ad2e0-4ac6-4a6c-9c60-59326b626e3b · outbound

This paper cites ” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models.

MetaSC: Test-Time Safety Specification Optimization for Language Models ” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:15:45.184439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T11:15:44.994573Z digest=sha256:cbe669e92b0292266a37628a99aa4777def665f506bae84b5cd2e01f6ec8b621

Observation 169c8b7a-e174-4a73-a2ab-c6fe5b378ac2 · outbound

This paper cites Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao.

MetaSC: Test-Time Safety Specification Optimization for Language Models Noah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan, and Shunyu Yao

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:15:45.172974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T11:15:44.998969Z digest=sha256:0a91422574be4c9365c5a752eb45c256ab67ee93131770fef6e7689e99a8bbd8

Observation d99bd8a6-60a0-4f3d-ae45-01b02d561c2a · outbound

This paper cites The ART of LLM Refinement: Ask, Refine, and Trust.

MetaSC: Test-Time Safety Specification Optimization for Language Models The ART of LLM Refinement: Ask, Refine, and Trust

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:45.002434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:45.002434Z digest=sha256:03d378e9dc0312c8feee78be59bf27c3dfc201798129fb7bfd6fdae18b37f700

Observation 2335f18c-ae00-4d3a-aa4c-0bb93ef66df7 · outbound

This paper cites Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate.

MetaSC: Test-Time Safety Specification Optimization for Language Models Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:45.006636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:45.006636Z digest=sha256:caedc11ea2fdfe509ba299ed029806d47f9e7d980f89e2ea5134e5abf5862fa5

Observation e2a488f1-08c4-43a6-a2a4-94b88f78a080 · outbound

This paper cites Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations.

MetaSC: Test-Time Safety Specification Optimization for Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:45.010325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:45.010325Z digest=sha256:c369eb801a243c42f1783a9061f629f3196ceda4865c195143ed1e428b7f6064

Observation 8f65bae9-4b97-4834-8339-ca1575789c22 · outbound

This paper cites From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks.

MetaSC: Test-Time Safety Specification Optimization for Language Models From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:45.013448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:45.013448Z digest=sha256:bfba466eb706035462faf2ac9ff904bad50db50f25f614d86c8378b1f1e9b555

Observation 61923c79-eda0-4f66-ace6-2420611a4b4f · outbound

This paper cites t spect 0 Safety and harmless.

MetaSC: Test-Time Safety Specification Optimization for Language Models t spect 0 Safety and harmless

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T11:15:45.163848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T11:15:45.016767Z digest=sha256:eb469a7dba972bf13b2bb57202a9250d4a3fa1f5b2881dba8e1f2c37982f766d

Observation 8d0a213c-c91b-4670-93f4-8d7e9cdff5f3 · outbound

This paper cites Configurable Safety Tuning of Language Models with Synthetic Preference Data.

MetaSC: Test-Time Safety Specification Optimization for Language Models Configurable Safety Tuning of Language Models with Synthetic Preference Data

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:44.963870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:44.963870Z digest=sha256:887487182e93bb849570cff6b097f29ed028e1e69bf7c30697e23f8884f0b3b9

Observation 35a6b2f9-8d39-4f06-adff-8953b346837d · outbound

This paper cites Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding.

MetaSC: Test-Time Safety Specification Optimization for Language Models Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:44.960566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:44.960566Z digest=sha256:73534856a73348cdd19f9a1505101417c253a415c06b109ded2af028d9da344f

Observation 87375eb2-47a2-4606-a742-5d48d49c85f9 · outbound

This paper cites A Survey on LLM-as-a-Judge.

MetaSC: Test-Time Safety Specification Optimization for Language Models A Survey on LLM-as-a-Judge

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:44.968673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:44.968673Z digest=sha256:e6b9013f31040639d8606be6c61870eed40fc3315fd91faf7ecd23c6a8119705

Observation fbc5c325-868a-42c1-87dc-236572a5aafb · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

MetaSC: Test-Time Safety Specification Optimization for Language Models Constitutional AI: Harmlessness from AI Feedback

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:44.956043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:44.956043Z digest=sha256:3f31d18f0e4724bd21aab4e8827bf39a9157fd5fce1af4a8eeb9d4df5eb9128b

Observation edf4a044-44f1-432b-a217-65ca0d19f1f8 · outbound

This paper cites The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models.

MetaSC: Test-Time Safety Specification Optimization for Language Models The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T11:15:44.981245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:15:44.981245Z digest=sha256:3d6ad0a49f3e821ea0a6d9e6f551e3098e566f8cfa98b350b3a9bf960611691d

Pith citing papers

Observation 3fd0c1e7-a6d4-41c3-87cf-042ecdf831cb · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models MetaSC: Test-Time Safety Specification Optimization for Language Models

Reference 194

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:41:23.482650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:92921cb7040a516cdcbe2bf2d3ffe8f6c37f1119d9800d327c6068150607c947

Observation ba2fea89-06d9-46f3-aa23-ece80465df79 · inbound

Prompt Governance? On Governing Technologies Governed by Natural Language cites this paper.

Prompt Governance? On Governing Technologies Governed by Natural Language MetaSC: Test-Time Safety Specification Optimization for Language Models

Reference 107

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:25:32.875448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T08:17:10.481202Z digest=sha256:5b19ee9f9c9446d5b50c8a7ee34d4e9b98dd4d7e6f6ce3c54c7cb1d847962c09