Pith. sign in

Paper Citation Record · LEDGER

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

As of 22 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 9 inbound Pith citation observations for arXiv:2506.15651.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.15651 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:55:19.365804Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:52:06.705135Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.705912Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5d924e4a-6e1c-4a66-a136-cfe0b695ed8a · outbound

This paper cites The claude 3 model family: Opus, sonnet, haiku, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning The claude 3 model family: Opus, sonnet, haiku, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.878456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:16.078806Z digest=sha256:0393c9aecc00a39193b8e5f151d0419a41377adfdacef945564ea8ea9eb734d2

Observation 6c02a401-4af8-4245-ac7f-1578f504b921 · outbound

This paper cites Training a helpful and harmless assistant with reinforcement learning from human feedback,.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Training a helpful and harmless assistant with reinforcement learning from human feedback,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.627779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:16.146217Z digest=sha256:d7f044086bf1e70e3d52af758f284b5dd7d272ed94d50d6bb11d25bd39e74954

Observation d409543e-9789-4a80-a823-91960e13630e · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:16.515092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:16.515092Z digest=sha256:cb1b718f69d45f06057ceb695a932a4746363a257df7678ecb64852ef5bcf400

Observation bf3ebc87-f99e-47ec-bca2-45f2ea52e711 · outbound

This paper cites Odin: disentangled reward mitigates hacking in rlhf.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Odin: disentangled reward mitigates hacking in rlhf

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.516900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:16.653707Z digest=sha256:700c62e4f71cc7aee9c4efb41e0ceceafeae338bd8825ee9ee5662e45daea6f7

Observation 8e933ef1-de09-41c8-b638-2b2aa1613da6 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Ultrafeedback: Boosting language models with high-quality feedback, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:23.220705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:16.808290Z digest=sha256:36b72942b315778cf6f9e8b6586ec0c5ec41e9327e47978493eaf4b07b86cbb4

Observation dc5c5bfc-8105-49c2-a982-d9c7d62efdea · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:16.919221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:16.919221Z digest=sha256:dd92e34b436db153e1edac6520c7bee20d59b30d9a23d428e9f9d0d54a4bc6de

Observation 6c38fd67-dc42-46cd-b62f-8d1ad1879e81 · outbound

This paper cites Length-controlled alpacaeval: A simple debiasing of automatic evaluators.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Length-controlled alpacaeval: A simple debiasing of automatic evaluators

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.952432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:17.098160Z digest=sha256:1e88cfa01682cdb1fc75ff74c745c31c748cfc8eb6997622b0c07357b0221e66

Observation 1632e14a-4c0e-4f8b-a419-19fc043d74dc · outbound

This paper cites Reward shaping to mitigate reward hacking in rlhf, 2025.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Reward shaping to mitigate reward hacking in rlhf, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:17.220976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:17.220976Z digest=sha256:834d3a202f1596293b13ef98bcfdb80c284500cf2d6889a7cae4fccbf95678bf

Observation a17db659-2111-4ca2-a188-08f38718066a · outbound

This paper cites Scaling laws for reward model overoptimization.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Scaling laws for reward model overoptimization

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.687833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:17.370411Z digest=sha256:911f32536f1233f68fe9062008ad34d24b544d3466819a38d3f484b61e573ce0

Observation 3b3496ff-6536-4ee3-8b78-27e842cb7a96 · outbound

This paper cites S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning S., Green, R., Mokrá, S., Fernando, N., Wu, B., Foley, R., Young, S., Gabriel, I., Isaac, W., Mellor, J., Hassabis, D., Kavukcuoglu, K., Hendricks, L

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.395484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:17.475196Z digest=sha256:bad88b0e80221146d2ad3216c08c6d975010d3f188724284339f00138f584951

Observation 33486952-d28f-4202-89ff-01085059ccad · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:17.559872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:17.559872Z digest=sha256:22f87522b3653d95bbcc3791fd33aaafb2f50da6bfb5152f07054887e95bc334

Observation e46063c1-7856-4808-89d6-2be2185331d0 · outbound

This paper cites Openrlhf: An easy-to-use, scalable and high-performance rlhf framework, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Openrlhf: An easy-to-use, scalable and high-performance rlhf framework, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:22.089056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:17.694338Z digest=sha256:5b0d3415aca2615d01040d05cd1436ffd90824559433cc621d37bda4c3dc78cc

Observation 2c941af4-2b1b-46cc-81cb-7852c8327d7c · outbound

This paper cites The Llama 3 Herd of Models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:17.770085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:17.770085Z digest=sha256:6d796e21986b5b9f3af6abf08aea860a41150cec5497cb6064b804b2455d3b05

Observation 168800b7-45b3-482e-883a-7fb8975cd15d · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.776544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:17.841397Z digest=sha256:cb8874e1bf085b413316980db2521ec045b7bc3e1bd64ef4eda5cc6f33a25415

Observation 96037b73-3393-4f8f-9b99-617ffb2e20f2 · outbound

This paper cites Rule based rewards for language model safety.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Rule based rewards for language model safety

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.471082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:17.907897Z digest=sha256:08adfab401550f5a2c2a0fda094661eb360b18213f8d958d11abd93b46f14f4b

Observation 3be88102-f3ab-457f-a736-06037b856f0e · outbound

This paper cites GPT-4 Technical Report.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning GPT-4 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.001414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.001414Z digest=sha256:5b4f8252d9ed81c972fd9340e9882e3296c2efcfb3e97258b17f75d71c84d72b

Observation 6baabdaf-39d8-49c0-9a0a-198d02f2d057 · outbound

This paper cites F., Leike, J., and Lowe, R.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning F., Leike, J., and Lowe, R

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.197182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:18.056217Z digest=sha256:8f5845658206f9403f94726dc490f4e5d5639afc9e9b7da360f4f816d4e04ed8

Observation f4682ebe-c9bd-4415-aa2b-018b7cde02f4 · outbound

This paper cites D., Ermon, S., and Finn, C.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning D., Ermon, S., and Finn, C

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:21.034443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:18.139566Z digest=sha256:22690131620444d1f9bff497485f69f5f24780ff1dd91951e2a46aef29adf26f

Observation aef2970b-9987-4889-a699-3eeb389b82eb · outbound

This paper cites Warm: on the benefits of weight averaged reward models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Warm: on the benefits of weight averaged reward models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.818159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:18.263261Z digest=sha256:18d22636b8df93e899cc2cd05e1fa77d5f6c678923d3295d50fc75df544e6634

Observation 91244e0a-5e96-480d-92cc-fd7a69542366 · outbound

This paper cites Proximal Policy Optimization Algorithms.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.379988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.379988Z digest=sha256:6f6829533910ef1edfb1afc39ff5cd7dfdccbf405404a7f03eaf17a37f4cf76f

Observation 1a06f961-8d08-4459-8a29-ff3f0708672a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.568219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.568219Z digest=sha256:c6b196f941ac8a1e25fc5cb1cc0952315edfe425ad5fe54259fb78a431109088

Observation 38a3a96d-3812-48de-9d11-7143f4d24650 · outbound

This paper cites A long way to go: Investigating length correlations in RLHF, 2024.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning A long way to go: Investigating length correlations in RLHF, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.621509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:18.684210Z digest=sha256:07e8dce0c5cada69966f4eb9ce1e5d36c14800e8928af767ded8a763dd72d8ad

Observation f1bce08e-2765-4650-ad99-e26236de101d · outbound

This paper cites Learning to summarize from human feedback.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Learning to summarize from human feedback

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.885733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.885733Z digest=sha256:1930b119e9ed6c21ab2db9f134f32742d5eec6a54033f63f974dc549acd59df9

Observation fe93116b-d111-474a-a295-271f3d6f21d5 · outbound

This paper cites Interpretable preferences via multi- objective reward modeling and mixture-of-experts.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Interpretable preferences via multi- objective reward modeling and mixture-of-experts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:18.993349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:18.993349Z digest=sha256:db6a93605be4b69d1ccfa1c8dffbc68e03c41c4b84b6886c2a681242701e6879

Observation 5f3038dd-c30c-4500-b988-cc7392c35913 · outbound

This paper cites Transforming and combining rewards for aligning large language models.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Transforming and combining rewards for aligning large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.466949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:19.143851Z digest=sha256:b40b90b92108a1301381a00b4e7dd2b34ddae017ad157317e3b9182976972e39

Observation 2bfa2f02-2fad-454c-b3c1-0db4f0351f45 · outbound

This paper cites E., and Stoica, I.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning E., and Stoica, I

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:20.274208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:19.259824Z digest=sha256:5d27e363a3b104067925535995fb9bdae51b595e7161895df3ff4689cb4e6131

Observation 3e40a4df-025b-4e51-bfef-1be957b8cf7f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T23:55:16.329410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:55:16.329410Z digest=sha256:1dba7c9c43f88cb228f855edc77a5f1cbb160efb177c65716c29e1cac64a1c24

Observation 29deb25a-cbbc-4ee4-ba28-98eb74ad27e1 · outbound

This paper cites confidence.

AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning confidence

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:55:19.995163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:55:19.365804Z digest=sha256:0626a6be6c612d4d63cc89f15901839425c80ab40546f32c4f048808e114ea67

Pith citing papers

Observation ce1cbc83-1f7a-418a-9f3a-c78c75413aaf · inbound

A Survey on Progress in LLM Alignment from the Perspective of Reward Design cites this paper.

A Survey on Progress in LLM Alignment from the Perspective of Reward Design AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:06.705135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:52:06.705135Z digest=sha256:ebad3a352ee99c9b9ac81dc279d9d322afa5ff4c80b1c7612539b13b6c62c29e

Observation eb9dc29b-af21-4a33-837f-297378dc94bc · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 262

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.294691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:467b5a13f6e52732dd29f6d14b2849e896e5b461e0f41e9ee2d48a1e3ba7cc9b

Observation ada27720-584f-40ab-a0a2-a0c9f39baa49 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 170

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:42.225595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:42.225595Z digest=sha256:361464a6df2c6c03c9a881dc691f33d932ff8d0420ab744ead49ea63a567aace

Observation 0ca00ec7-4d2a-4e55-9fcd-e12388adae42 · inbound

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration cites this paper.

Training LLM Agents for Spontaneous, Reward-Free Self-Evolution via World Knowledge Exploration AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T12:10:23.451741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T04:36:27.381942Z digest=sha256:4e239aa017de024d502bff28d9c1349a7b5364c373d9f3f863e04e0405d1b140

Observation bfb0290e-40fb-4bfc-a48a-58a088745556 · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 151

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:26:05.330212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:ed88d1616c655903384fbca1b88086f74e921fb289c50722349639ffa0cde58d

Observation abdb9b9b-5c93-48c9-94f5-4a5ab4fc5248 · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.112770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T12:11:23.775843Z digest=sha256:1fb36f22781717c97975930b0443ba260a14295a01ddaae8adbf53c29d8202f2

Observation cee7903f-9cbe-4bc3-b9c4-dbedb0d01671 · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:21.439793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T09:19:39.848194Z digest=sha256:013b74ef0b4888e8604b57f5b915dc28452b25e79dec200d98929330d16f9cf5

Observation e2dc734c-b2b6-479d-9a4c-e9bd28f1bad1 · inbound

Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge cites this paper.

Generating and Refining Dynamic Evaluation Rubrics for LLM-as-a-Judge AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:23:12.618105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T07:21:37.937934Z digest=sha256:ab70a973ee998b5cbc4a83b7adee1a41c5c5046daa56a261507740e3953bddd7

Observation 2107dac5-c1d6-48f2-8801-6c2dd0e50e20 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning AutoRule: Reasoning Chain-of-thought Extracted Rule-based Rewards Improve Preference Learning

Reference 219

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.707284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:a5786773f2c443a4e7f2302e6ae4cce99f0ecf129deae2a0fb9cb77a51c1bc98