Pith. sign in

Paper Citation Record · LEDGER

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming

As of 18 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2506.04302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.04302 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:20.444420Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f175e702-865c-41fb-9b1a-e41681f97f73 · outbound

This paper cites Training language models to follow instructions with human feedback.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Training language models to follow instructions with human feedback

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.772227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:19.959731Z digest=sha256:fd753250cb35c40b95cac986c45850a9ca79ce9645719b9259782bcb55ff4e90

Observation dc690450-026d-45d9-979c-2b506f6cd357 · outbound

This paper cites Unsolved Problems in ML Safety.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Unsolved Problems in ML Safety

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:19.968574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:19.968574Z digest=sha256:6b05cab6f44cfebee4d40603e53de3947ebd7bf7d9878ccd5d7ba8ebfe1803bb

Observation 17cc6a00-b7c4-4ae6-8e15-cffb15d9a430 · outbound

This paper cites Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:19.978526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:19.978526Z digest=sha256:6d3616bd6c6ae5ee4a2a7a4ba9280d986af33b2b9bdf5eb94de1375029e0c1ca

Observation 7a20bf65-7fe3-483d-939c-4c47a640d3ac · outbound

This paper cites Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:19.989451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:19.989451Z digest=sha256:22de720a9d48ec21a7d979c97b38d1a4f679b15fcfad7ea9054fba92a8ae0e62

Observation 381be7f1-ef81-4ae9-b907-d175e55a1aa0 · outbound

This paper cites AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.000342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.000342Z digest=sha256:206d8f6c2b6cdc39ddad1b1e4505fefe332f0afbed6f290ae013d26e6ebe0b49

Observation 0c503a99-6875-4ae6-9011-00c5c4b72415 · outbound

This paper cites Query-Efficient Black-Box Red Teaming via Bayesian Optimization.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Query-Efficient Black-Box Red Teaming via Bayesian Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.011756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.011756Z digest=sha256:2add4f82298e88e0e64c58e2b02cc873ce11f43376a7103f53c86b9788ab3e7b

Observation 853fdc58-f2a8-47a3-aa54-12b83aee6b31 · outbound

This paper cites Large Language Models as Optimizers.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Large Language Models as Optimizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.027907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.027907Z digest=sha256:82e23cf415db4f42d0f09a7c9d791f48dcaf6b3705aebbb8b8f2862a6e965385

Observation cac5389b-1754-4149-881a-773667e75f80 · outbound

This paper cites Red Teaming Language Models with Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Red Teaming Language Models with Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.037233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.037233Z digest=sha256:1d92000396ce5f070ef00d254b0429bdddd248f50ace0b48b1ffbd643ed17b0b

Observation 382dc7af-da47-42fe-b859-be941488cdc2 · outbound

This paper cites Discovering language model behaviors with model-written evaluations.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Discovering language model behaviors with model-written evaluations

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.752229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.043431Z digest=sha256:a02fb87211f1b5cda4206a097fd02edc3356883ce31a214f13f6f534668684ad

Observation b06c2a86-7fd2-4537-92e4-511cf268330d · outbound

This paper cites Attack Prompt Generation for Red Teaming and Defending Large Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Attack Prompt Generation for Red Teaming and Defending Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.050296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.050296Z digest=sha256:36a61533800f345e91aa2af48b6465036e2a315e1d92882efcb6d4d3333b0f3f

Observation ae34b8d3-fbac-43ec-a207-7dba2ef5b36d · outbound

This paper cites Explore, Establish, Exploit: Red Teaming Language Models from Scratch.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Explore, Establish, Exploit: Red Teaming Language Models from Scratch

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.057855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.057855Z digest=sha256:33826e4ec5eeed98e27aa041ee613608a3b38ece4f95cb94e6e7cd105d883f5c

Observation 3e93d054-5718-4965-a335-c1904c8f7421 · outbound

This paper cites Curiosity-driven Red-teaming for Large Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Curiosity-driven Red-teaming for Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.071156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.071156Z digest=sha256:3bfecbc223a30bdeab3ef5f8ff5a6717e06fc6b00debde4fe2f11c59eae71fc2

Observation 87957fe8-5de1-449b-9104-f13c3a4e5268 · outbound

This paper cites DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming DiveR-CT: Diversity-enhanced Red Teaming Large Language Model Assistants with Relaxing Constraints

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.732369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.078058Z digest=sha256:90f4979e0baeb7dcacb553c41572d0e5c3a9b3f2d9cabc949ac35281ac8aefe9

Observation 38a38d83-3fff-443d-a6f5-8bef94cf5d1d · outbound

This paper cites CALM: Curiosity-Driven Auditing for Large Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming CALM: Curiosity-Driven Auditing for Large Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:55:20.946358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.085750Z digest=sha256:0afb6c7da7cf122768c60ef8a31917b476a7971e82526167f89a5eedb98f524c

Observation 9b2881ed-b30b-4c61-9fb6-045bdf8030f9 · outbound

This paper cites CIM: Constrained Intrinsic Motivation for Reinforcement Learning.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming CIM: Constrained Intrinsic Motivation for Reinforcement Learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.716179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.096049Z digest=sha256:a5fbf49c5e1493d51d88ef8e2fc99482f4d7a576633d9181bfda4dd8f9390284

Observation a9f03fec-4556-4d63-a544-00dd68645c7e · outbound

This paper cites Proximal Policy Optimization Algorithms.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Proximal Policy Optimization Algorithms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.103060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.103060Z digest=sha256:72c2967a76f2597d6cffb825454943f3d05f1b37f9e75c853f32d9d86480c7ca

Observation d7c316e8-2e99-4ee5-8f7d-4ed2de119e3a · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.110906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.110906Z digest=sha256:370b96bd631e5000e4853a189fba50fc4442c9d85baeb7e4204b46101f762d6a

Observation 33da88f3-d4b9-4424-ba63-b3cec21e6766 · outbound

This paper cites Tree of attacks: Jailbreaking black-box llms automatically.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Tree of attacks: Jailbreaking black-box llms automatically

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.699718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.118783Z digest=sha256:329a673c3e661cc28a27c698a15c826ccc24ef14fd354341dfcedb6cbdaba7d2

Observation 9a313817-9950-4f83-a6e9-e015f0e9b66c · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.683610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.124594Z digest=sha256:955603c406b1aae9b4a1f5f4bfccee04de10a4b452fa554961c3f73e771402ba

Observation ad3fbefa-428f-42c0-8b06-d0a1cd315538 · outbound

This paper cites GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.132349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.132349Z digest=sha256:de5915b129b83dd8e5051de1503cd88a565ff3f7504a01f469c069e8000ede70

Observation 8ea00912-2743-44fc-9d6c-47e46ac6d4a0 · outbound

This paper cites https://github.com/ huggingface/trl.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming https://github.com/ huggingface/trl

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.667522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.138450Z digest=sha256:7033ca502857230a2833caf8446318b2d37fc523dac5a02da7f04f2c6ff9252f

Observation 381dc6ef-ac5c-49ae-b65b-9b98b331c786 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Direct preference optimization: Your language model is secretly a reward model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.650520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.147018Z digest=sha256:a88edfcd6016e5b3614f17d2990109614718adefac5a60c6967d4c6930ef9d7c

Observation 908c75a0-f94a-4865-a9f4-8d9bbe7e8e9f · outbound

This paper cites OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.161901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.161901Z digest=sha256:4ad217bf6b053b0e6f54f006cc36a79776f70e0745c03386e64e2a4af508a09e

Observation 08285b22-1570-4643-add2-818f40fc8436 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.167213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.167213Z digest=sha256:f85950845e8c6dad29d449cd059fb08613b201cf9aecf66fe5e4dda2a49f5375

Observation 4ed6cbbe-99a6-41cd-ae6a-9d427b1592b3 · outbound

This paper cites Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Is Reinforcement Learning (Not) for Natural Language Processing: Benchmarks, Baselines, and Building Blocks for Natural Language Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.175303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.175303Z digest=sha256:91e24104355dcf1a84a435da6348bd380894cc702af05cc20d875a9639e4a11a

Observation 87e49c1d-7f06-4080-852f-d4786f7fc0d8 · outbound

This paper cites GPT-4o System Card.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.188583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.188583Z digest=sha256:59f46b23ee4aeee1a5cb696836bdd800aa8ee6fbab38c77fa9da5f0945ae31cd

Observation 8aeef9cb-e71d-4f57-998c-6696040a3799 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.197684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.197684Z digest=sha256:60529593c7dccb3523d508d1243d9b57fc7f96ca83e85aaaf679d26cc04bd50d

Observation b4b230da-84bc-4768-b4c3-268d49f2c8c0 · outbound

This paper cites What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.205546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.205546Z digest=sha256:305f48c8525602921280f1d250e817c7a96068128d77b602ff68b2029b7e66a2

Observation 0ab1fa87-310f-4571-ad83-a3e2d187c08c · outbound

This paper cites URLB: Unsupervised Reinforcement Learning Benchmark.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming URLB: Unsupervised Reinforcement Learning Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.212381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.212381Z digest=sha256:cab843c407ba1db3827d6ee69fa3c9ed9a006ed0d1a618e8a40aa98a4a46f4fd

Observation 40351ced-22cb-4d3f-a5c6-0c5bc8cea678 · outbound

This paper cites Exploration by Random Network Distillation.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Exploration by Random Network Distillation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.217369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.217369Z digest=sha256:29088974f07565de6b71c03a5c51e4307a46722b4bbf41aa50ccceb835b10211

Observation a180c0b8-31ba-47a6-8997-c72253260825 · outbound

This paper cites Large-Scale Study of Curiosity-Driven Learning.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Large-Scale Study of Curiosity-Driven Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.225697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.225697Z digest=sha256:cd6fb397115194759740337ad90294753821dd38bd66884d527abf4559390a6d

Observation 2a95946f-6344-4dda-9356-a73851590d2b · outbound

This paper cites Made: Exploration via maximizing deviation from explored regions.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Made: Exploration via maximizing deviation from explored regions

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.632220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.231970Z digest=sha256:aeb4041b593339ca36dff9582717846e9c6f027d50db4e8285013ab7e57d378f

Observation 87a2b6a5-e034-47bd-b063-0a4a9d478ccf · outbound

This paper cites Aps: Active pretraining with successor features.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Aps: Active pretraining with successor features

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.615141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.242868Z digest=sha256:7f49ecbb21cd608b37433354be1b21c71354a44a62112f060bbfd196db5cb5b0

Observation 9902af5a-2f2c-4bf3-9db7-8616ceb4b872 · outbound

This paper cites Provably efficient maximum entropy exploration.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Provably efficient maximum entropy exploration

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.597103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.248052Z digest=sha256:7a54d9ad24c47c4535bf9a1cbd9aa7faaa5ddbaf87e828ddc288040d8c8063f2

Observation 46edf65b-7e79-4371-8a30-383151690748 · outbound

This paper cites Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Task-agnostic exploration via policy gradient of a non-parametric state entropy estimate

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.579769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.254621Z digest=sha256:70996b7d38a51c1b1bd1cd2814cac447eba6ad78685738eadc060a90719627d3

Observation 63dc068b-a992-475e-836b-fa81c7bae7ed · outbound

This paper cites Behavior from the void: Unsupervised active pre-training.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Behavior from the void: Unsupervised active pre-training

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.563908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.260652Z digest=sha256:e07bb1c6eb02f3198218f05de72d1e533dd41899897b4736b31927e62c477d72

Observation 9623f91d-a38e-4822-b268-deb1e410a141 · outbound

This paper cites METRA: Scalable Unsupervised RL with Metric-Aware Abstraction.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming METRA: Scalable Unsupervised RL with Metric-Aware Abstraction

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.265825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.265825Z digest=sha256:f76d422ccf4f134ce7d142d0fcce631157347420c390f1301b8d2ba2dbd48b9e

Observation 5debbec7-ef51-442b-821b-4bede4fed053 · outbound

This paper cites Variational Intrinsic Control.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Variational Intrinsic Control

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.272855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.272855Z digest=sha256:572521e7dc4df36a40cf7914fe1350ce626dbf1cbfa4153c3184dbcc660083f4

Observation e03d3241-ec77-42e9-b744-9191bd314146 · outbound

This paper cites Dynamics-Aware Unsupervised Discovery of Skills.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Dynamics-Aware Unsupervised Discovery of Skills

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.279140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.279140Z digest=sha256:bb75786aea4a0f25af01b5fa64f8dc45caeb3904fc43c6c09b0e72ddeaabd948

Observation 3b9a831c-c371-4a1a-8e9f-7b704e083dd4 · outbound

This paper cites CIC: Contrastive Intrinsic Control for Unsupervised Skill Discovery.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming CIC: Contrastive Intrinsic Control for Unsupervised Skill Discovery

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:20.286572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:20.286572Z digest=sha256:336001d775b124f48b0066e36613ecc52da2c5e7a97abc27ab68419e6b56f6b8

Observation 7edfc9e9-e9a3-4498-861d-cc2b94dbd0e2 · outbound

This paper cites Lipschitz-constrained unsupervised skill discovery.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Lipschitz-constrained unsupervised skill discovery

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.546382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.294060Z digest=sha256:5d00a72a178017216d9c7accccfb0cfdb770c590c159afdc80806ea247ff3d6c

Observation 400927c6-059a-448b-99bc-51f51c43c6dc · outbound

This paper cites Contrastive learning as goal-conditioned reinforcement learning.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Contrastive learning as goal-conditioned reinforcement learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.529985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.300397Z digest=sha256:1f7d81d0fdba8e0472b40b49de1f6deef53ee7f82ce3c1f36e8e1c80efaab737

Observation 663a31b7-a700-41ba-ad81-9d45e84ce291 · outbound

This paper cites Behavior contrastive learning for unsupervised skill discovery.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Behavior contrastive learning for unsupervised skill discovery

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.514043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.307494Z digest=sha256:ad250f873c6e4372dd2b005c4056c7de1197221b8771b0198437e79826f0e09e

Observation 58364df4-b6e3-41b3-b97f-1b27c1a821fb · outbound

This paper cites Constrained Intrinsic Motivation for Reinforcement Learning.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Constrained Intrinsic Motivation for Reinforcement Learning

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:55:20.543202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.314427Z digest=sha256:2be327000b88baa70496a0fa65b71d139c97cb3e789c416f20bc1d71bbc6f4d8

Observation 6827f958-515b-4993-9e21-ce34312fc62d · outbound

This paper cites Texygen: A benchmarking platform for text generation models.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Texygen: A benchmarking platform for text generation models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.495004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.322001Z digest=sha256:12979e878b09bd91341d553cd1a2f2f2d36135ce4ea514c22d5c95678e940194

Observation 54a87242-300d-4a7e-ac4e-925cfcd01723 · outbound

This paper cites Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the abstract and introduction do not include the claims made in the paper

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.471942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.328339Z digest=sha256:82be77fa25a33941bbf8e6de45f2fc5c166cd3a153061eac80d40a490fec365b

Observation 3b7f25ba-2680-4287-a94a-9c13dfc3475a · outbound

This paper cites Limitations.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Limitations

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.445445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.333557Z digest=sha256:32216feea3c75c59b860598b242907c9e11f79a60d3fe73e9651ce956fdb3732

Observation 43689bd1-c7ff-4d64-983c-b29be3ea708a · outbound

This paper cites Thus, we do not provide original theoretical results of the each baseline method.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Thus, we do not provide original theoretical results of the each baseline method

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.421981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.338073Z digest=sha256:1266b5f357bfb84ef9e2b08a1d7447cbbcc289071c397da1d1bb41429c8d1fee

Observation 2e769d34-be7d-4293-9717-f63132a532ba · outbound

This paper cites Our experiment results can be easily reproduced with a simple environment setup, the default config files, and the prepared shell scripts.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Our experiment results can be easily reproduced with a simple environment setup, the default config files, and the prepared shell scripts

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.401126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.344055Z digest=sha256:01556023e72b9957eef99252a4971e64e448dc8d69fffe59612192f888e73460

Observation 729bd66c-a64a-44e7-8592-905e34f9ad0e · outbound

This paper cites Guidelines: • The answer NA means that paper does not include experiments requiring code.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that paper does not include experiments requiring code

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.375072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.350750Z digest=sha256:4e7b23cafaee3eaf23783e808dfe9b430d6357b33558b89f72b01755770b3c86

Observation 541639fb-fc2c-4f21-b843-2c2018df7c94 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not include experiments

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.353669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.359237Z digest=sha256:951d01c319f70ab4e4f01e6049ff000f621844a7512cbe3bb112ca50c6955eb3

Observation b11c78d3-4c70-461a-bdad-6b0d36c732ab · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not include experiments

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.331413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.364627Z digest=sha256:ac173b1b3bc79148f65a66be600bc9c84b9c1115fab1db5048ee76728477ba97

Observation ac8777d8-1fcb-4c87-9979-fd85f64a6bfc · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include experiments.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not include experiments

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.308529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.385226Z digest=sha256:9093fc1426ac1b44a711d5722156ab2bc019424ad86d7deb954193ac9c6c7e2a

Observation 5f88aa89-5bf3-4ead-bab4-b454b1aa1367 · outbound

This paper cites Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the authors have not reviewed the NeurIPS Code of Ethics

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.289513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.392099Z digest=sha256:4e3b16bdfa2c8bf4c580dc9b675ea91886e5a1be0a133d348e7bbf77f0ee510d

Observation bc5483a2-b170-4143-a810-6eb120b8953f · outbound

This paper cites Guidelines: • The answer NA means that there is no societal impact of the work performed.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that there is no societal impact of the work performed

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.269397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.398177Z digest=sha256:c4aaa0f7ed534794075ec36443681bb9a635c437015778cc5c77de4cae1799a1

Observation 87174eda-9204-4f40-ad21-68fde8abdc1c · outbound

This paper cites Guidelines: • The answer NA means that the paper poses no such risks.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper poses no such risks

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.251148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.402830Z digest=sha256:3d7c790f56514e18d8f6c7a3880c6eaa8cd4622c48b125c884f2c3cfd8dc681c

Observation bdc3f549-f526-4d11-8653-dc3203cafff8 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not use existing assets.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not use existing assets

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.234174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.407677Z digest=sha256:9fb56d0846425752db22e00fbe9e527b22bbcb7cc4b0ab5682bfcea75a05492d

Observation 3a24e7df-fd5c-44df-ba83-91600cc58698 · outbound

This paper cites Please refer to https: //github.com/x-zheng16/RedRFT.git for the document.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Please refer to https: //github.com/x-zheng16/RedRFT.git for the document

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.217962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.414775Z digest=sha256:c6790b9f6a9b9b87ddf87ab201401a1407c584bb8787dcf7c579816b78e13b67

Observation d7fb2682-a2e7-4ba4-af7f-21bf833ea638 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.200211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.419930Z digest=sha256:0e8dad24f194042076b58ec2dc716c1a1f90f90f7ebeec45a1f9230898f2c16a

Observation 431ab6db-b213-4dd0-aa74-f9bc67e192cd · outbound

This paper cites Guidelines: 25 • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Guidelines: 25 • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:55:21.180137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.425912Z digest=sha256:6d04752e31ef70f23fdfc842073619ba2c31950ab1f1095e4612f7eafeff4502

Observation 8f459fa7-9125-4e1c-9e78-53da5a87cf8b · outbound

This paper cites an unresolved cited work.

RedRFT: A Light-Weight Benchmark for Reinforcement Fine-Tuning-Based Red Teaming Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:55:21.162162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T10:55:20.444420Z digest=sha256:210eb86e0e937a8d24319da32b34bfc39667e63b1ff8ba03374812278dd31b66

Pith citing papers

No inbound Pith citation observations are available.