Pith. sign in

Paper Citation Record · LEDGER

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment

As of 18 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2506.12446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12446 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:56.794234Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:15:37.083838Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:17:30.651801Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4912a7b3-e65e-41ef-8303-80c8726670f7 · outbound

This paper cites GPT-4 Technical Report.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.681605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.681605Z digest=sha256:5400c7b6ecaa69cec63e98d02417efb46cd45ee7984c4a1995ffba99b3a0ac9a

Observation 36b19547-6eef-4627-ae27-4dcc131ed85a · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.685536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.685536Z digest=sha256:58c2d082b5ff936c3bbfad0d370d19a69285843f74d32d4bbe06ec840ef90516

Observation 5a301271-6912-400a-b58d-7c8d8420db2b · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.689287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.689287Z digest=sha256:406e6c073c67e6053aa1bc08f76a7c81429c71560e5aad6ee7c7c9a953090fcd

Observation 64b8a358-7d5d-46f4-b1bd-763f83fc57b1 · outbound

This paper cites Transfer Q Star: Principled Decoding for LLM Alignment.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Transfer Q Star: Principled Decoding for LLM Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.692630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.692630Z digest=sha256:212acf832bec4e5295194e184e606e3dff47f8aa1b6c77e1977e18c0b5bbb160

Observation bd759edf-e7f4-41a2-b8c4-a9a8ed339231 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.695734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.695734Z digest=sha256:fac78f51a042d23df59e83ccf590187e1450cedfd4df7df764d05cf7149ced6a

Observation be50a20e-9127-4da4-90a6-8a96f6782934 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.228881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.698732Z digest=sha256:b96add813947f987eab5f6de97cc3b293e001ddda138cb474692f6ffcc237321

Observation 5ba7954f-67b5-4ade-a03c-f8e5177afea9 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.218508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.701902Z digest=sha256:9e81e76ab84efec4f9eb65acb097f701fbc2b079897b055aca847fe1b9df123a

Observation f4b0db48-f16d-4119-9208-e0a95dcea152 · outbound

This paper cites The Llama 3 Herd of Models.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.704410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.704410Z digest=sha256:bc71faf4ad2952f6b743ea7c56eb17c6d49109c409b7a14b60390699733dc893

Observation 7c3de4f8-d9f8-4c1e-8feb-4f2346d6cdee · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment KTO: Model Alignment as Prospect Theoretic Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.707328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.707328Z digest=sha256:2c124d5d9c66fa214ed97243f48afa0a48af8d6b1f0485737d93f9ee4d55db76

Observation f374665c-7c97-4133-b1d5-1da0c4d3329d · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.207916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.710065Z digest=sha256:3c31274f5ea757cb13ccd04476c1c1f66527ac960c6415f6b55e8f40d4dfe7b3

Observation 750e2506-feb8-4427-b82d-5e28a9a52517 · outbound

This paper cites Value Augmented Sampling for Language Model Alignment and Personalization.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Value Augmented Sampling for Language Model Alignment and Personalization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.712904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.712904Z digest=sha256:8bbc52f5c84191876b2b2c6b5845d2301db6c9d82e6a10f6ae4ab66488bd3bad

Observation ff29bdb1-8785-49a8-a405-a1467e673158 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.716070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.716070Z digest=sha256:19ebb205415e74c0e6d9792cdac8dcd4dfa2fce85003a3b9a1266784e9380b23

Observation 30413e44-3eea-4f9f-99c3-37fa6e7a001b · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.198063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.718781Z digest=sha256:7d7f153657126723b260691914d776adf4bdb18265e38cb75a5941c123d5f18a

Observation 2d75fb37-7379-41ef-9b21-9d795adc87fe · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.188134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.721421Z digest=sha256:b8e2f6c266e55968f9c7900046bea0885d2d362e182341413afdcffb80daf0bd

Observation e2c5cddc-97b0-4acb-9b40-bb6750515811 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.178228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.724342Z digest=sha256:af35eef6c4a72da2f127b71cf2f8fd185f1f1767dda5965d29d235d075174aa7

Observation d2fea9c8-b210-42d0-8d78-732189a2b6d3 · outbound

This paper cites DeepSeek-V3 Technical Report.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment DeepSeek-V3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.727160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.727160Z digest=sha256:03983408d3cd670fbd232b8fd0fc4947b08114dcb7b2161bafc641d8dfd8b108

Observation 63beea3b-8356-4f0c-96e8-868abdd8ef81 · outbound

This paper cites Chain of Hindsight Aligns Language Models with Feedback.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Chain of Hindsight Aligns Language Models with Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.730028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.730028Z digest=sha256:e9c252c4767e0c50b08e72ec7cff46a23e068de6aad62b968cc096f9e69df21b

Observation 90b1afa9-5f82-4ca0-a092-fe93568fe61a · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.168071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.733069Z digest=sha256:6a97b23ca2df0e45534e0c1070519ffc0b95908f487e32bb6ebc7acb6c871b39

Observation b034acf2-58e1-40be-b0af-8d389838d0e4 · outbound

This paper cites Controlled Decoding from Language Models.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Controlled Decoding from Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.735807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.735807Z digest=sha256:50654650d5c10efef331e367e21bca0a357e9b9a39116aecd5edaa6eddbf5781

Observation 1367ce34-8eba-4b7f-a077-1e9959661349 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment WebGPT: Browser-assisted question-answering with human feedback

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.738980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.738980Z digest=sha256:b3757e169f6043e42e93a4e5e04cc98978d44abeb6a3e57eead4a18dc14eeb23

Observation cd673f1c-19fe-4958-a75d-b6523f875956 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.742029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.742029Z digest=sha256:cc3384bf290aab8a40bcf67c87aac214656c8466161ae8140335869f8ff5c4ce

Observation 922294fd-5611-495f-9001-32f2266d9f50 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.744661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.744661Z digest=sha256:cd0351bb09709742dcb2e98982dbb30ede43f4b4ed68e6e25eff5b09672f96c7

Observation 143a9fd4-240f-44aa-bc2f-f615c4e8db85 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.747399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.747399Z digest=sha256:4bc60e39e3b6248635e5fce80c45694044eaf154780a29e20e69d70ae0a5b70b

Observation 3199fd39-8f9e-45ef-ae6a-b68988faee9d · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.750234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.750234Z digest=sha256:5e0079785616ca6da3fee709d39dcec3ad484475d879f5e85a5872fe22c21234

Observation a7787224-2806-4057-ac3b-82b765031810 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.131045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.753004Z digest=sha256:5ca94a7e1ce2e1ba161ad18d6ff37de9cfdeadd2658204ae09f5b93c0db0673b

Observation 78ed9d1c-1e6c-47a6-8d80-cb1291d8bbef · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment LLaMA: Open and Efficient Foundation Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.755879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.755879Z digest=sha256:42d9080d3496e426dfb53f1b974ec89324ae8f9b9be9db6deb8fcddf9a4e61f2

Observation 62b97dd5-3282-43c5-8392-2a43316d5d4c · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.119669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.758899Z digest=sha256:d2c29493a703d3418bf43b6328a0253c8528428c7de38ac5c6a3b7dbaec12f61

Observation b835bba0-5dc3-49a4-a69b-fdd4e7ec0ca3 · outbound

This paper cites Aligning Large Language Models with Human: A Survey.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Aligning Large Language Models with Human: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.761990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.761990Z digest=sha256:84ffc4c53049ea6c54967cc254e58885faf9695a1aac5ce020c358acde2b2989

Observation 4e89d054-047c-48db-be8a-bd8991fc20d9 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.109086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.764821Z digest=sha256:0f04c74ba894c50b1eba5b2d9273b48aff351d95b37ead96c53e1d5926a36607

Observation ac75c50f-f1d4-4a42-9063-586384f48b1e · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.097961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.767542Z digest=sha256:45144bf0aac94fe8f71133baad744e70632b7ace8dea911e9170cbd83e2cd2d5

Observation affda21e-2a61-4a2e-a931-ae8f8241e175 · outbound

This paper cites GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.770058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.770058Z digest=sha256:27450919312799841af66f7aff63acf60425016f491a3a93f383c94e4c4ed9c9

Observation aca8fcc6-d181-4273-9fc6-653d0d315f8c · outbound

This paper cites Inference-time alignment in continuous space.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Inference-time alignment in continuous space

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:57:57.087942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.772720Z digest=sha256:1071727ec6a79b3c4dbc53c33907b613b2c6cc1837227878a0146e86eaf13ee8

Observation 71993794-c053-464e-aee0-5b8a05fd2004 · outbound

This paper cites Secrets of RLHF in Large Language Models Part I: PPO.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Secrets of RLHF in Large Language Models Part I: PPO

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.775426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.775426Z digest=sha256:52a50c6b4ead6348136b28f27d0144dddbf2c1df2b76c5622a4f54dbb3b9de0c

Observation 1b4a77da-8cf0-4db5-ad51-48d485e998c6 · outbound

This paper cites Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Beyond One-Preference-Fits-All Alignment: Multi-Objective Direct Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.778437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.778437Z digest=sha256:0dc7b5f4f8b34315c9eec3897971572c43941a44d714fc46cd7f00e49ccb2b5e

Observation f423252d-98fe-4eb1-b8ba-ddcb0c443197 · outbound

This paper cites an unresolved cited work.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T00:57:57.075895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-07T00:57:56.781278Z digest=sha256:cad189be18d3931c9c7a259ea40f19d40337530e2299cbee954023c929c955ee

Observation 5cdc6326-be2b-460e-bb3f-bacb141ca3a5 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Fine-Tuning Language Models from Human Preferences

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.785108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.785108Z digest=sha256:11864f0621b082d6f1bc357f5e96770f762803243c44123ecb03a1bfb25b8fa7

Observation a3988930-adf1-455e-8a44-bdacb97b37f9 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.788097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.788097Z digest=sha256:f93148ce56090e310ce394f61e5877869a28a4694dd7fa02853e5f74a94fb2bd

Observation 89243b78-9e17-42f3-8ef2-3cb6b72e0e5b · outbound

This paper cites online" 'onlinestring :=.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment online" 'onlinestring :=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.790928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.790928Z digest=sha256:b3b0d846910797c2258638d5d14c3e87b52b53e78d3c5695582f861e0ae6db15

Observation abab279f-b4a8-4765-8e2c-21f35eb5758f · outbound

This paper cites write newline.

From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T00:57:56.794234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:57:56.794234Z digest=sha256:46e16016268ed91adaff20117c90bd30828599764c78b6bfaaccaeabc92e88fd

Pith citing papers

Observation 3b792169-b2d2-4d3e-9265-9f0db8bf8c72 · inbound

Gradient-Guided Reward Optimization for Inference-time Alignment cites this paper.

Gradient-Guided Reward Optimization for Inference-time Alignment From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:17:30.653219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T16:43:04.984866Z digest=sha256:a2805f3e24a828b7c3d3479e9d9be89e303797cdc2ae0a14d44e1f63ce627646

Observation 033c006d-117b-49b8-afbe-e84dce3a2697 · inbound

Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation cites this paper.

Crayotter: Learning Long-Horizon Video Editing Agents via Group-Relative Preference Backpropagation From Outcomes to Processes: Guiding PRM Learning from ORM for Inference-Time Alignment

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:37.083838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:37.083838Z digest=sha256:757b082d5c15b08616ecad313189b6f87035cff4da48f062eaf9f8f13965df79