Pith. sign in

Paper Citation Record · LEDGER

Activation Reward Models for Few-Shot Model Alignment

As of 14 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 3 inbound Pith citation observations for arXiv:2507.01368.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01368 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:02:42.537188Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T13:46:04.570352Z

Reference resolution

75 of 75 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a309e831-ad2e-4be8-83ef-60ed9188627e · outbound

This paper cites an unresolved cited work.

Activation Reward Models for Few-Shot Model Alignment Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.321263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.321263Z digest=sha256:b6f5e5bed7752bce6cf49d3fd1e0020cabc60c4801c854d359d1c7e7986ee6f0

Observation 3392bfb3-0a1b-4fa3-b2c9-0efd1d22b87f · outbound

This paper cites Qwen2.5-VL Technical Report.

Activation Reward Models for Few-Shot Model Alignment Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.325145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.325145Z digest=sha256:4801331aafda4bbd61a53f445ea5efe18909c1d955a033a90da580032c53edc2

Observation e39bdd43-c782-4440-85ce-e37b73592ede · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Activation Reward Models for Few-Shot Model Alignment Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.331317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.331317Z digest=sha256:1f57c003b639e147f3a521902951762ab2b837b72ca0bd1eef721d6a29883557

Observation ca3027ac-b74b-4c0e-b83f-5a15eface521 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Activation Reward Models for Few-Shot Model Alignment Constitutional AI: Harmlessness from AI Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.337121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.337121Z digest=sha256:87842925cb8a0af5a9bd6f2b49c74c1ea807c50474d9f6c2efb069fbe4eda6a6

Observation 71de3fff-e5fe-4a2f-8a1e-0c330a8d5a01 · outbound

This paper cites Capturing individual human preferences with reward features.

Activation Reward Models for Few-Shot Model Alignment Capturing individual human preferences with reward features

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.339704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.339704Z digest=sha256:f41d918c56e40ce8ca48de7c23489b34f24dac52a9c9d4c07f0d81c4b9166b8c

Observation 0e72bd98-4df8-4b46-ab47-2f018686a0d8 · outbound

This paper cites Network dissection: Quantify- ing interpretability of deep visual representations.

Activation Reward Models for Few-Shot Model Alignment Network dissection: Quantify- ing interpretability of deep visual representations

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.378656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.342712Z digest=sha256:c562b0ed8e67a8ae3f7639de568692fe2388f64e68d61c14962b997be6cc6cee

Observation 61cdfbcf-9caf-459b-81d2-7543ce242d59 · outbound

This paper cites Understanding the role of individual units in a deep neural network.

Activation Reward Models for Few-Shot Model Alignment Understanding the role of individual units in a deep neural network

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.370334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.345424Z digest=sha256:237520e70b58dc1b12c9e0199d5f39c172a017d6e4dbbd61d28fc9ed48d63734

Observation 53587fdb-b7bb-48fe-99fd-bbac1ee0e9d9 · outbound

This paper cites Language Models are Few-Shot Learners.

Activation Reward Models for Few-Shot Model Alignment Language Models are Few-Shot Learners

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.348325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.348325Z digest=sha256:3208cf8659bd1e9e5504091896edc6d31b91f6dfd930a84364432d316e8a2011

Observation 73ebb749-b745-4825-9bd6-38f9b129794b · outbound

This paper cites RRHF-V: Ranking responses to mitigate hallucinations in multimodal large language models with human feedback.

Activation Reward Models for Few-Shot Model Alignment RRHF-V: Ranking responses to mitigate hallucinations in multimodal large language models with human feedback

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.361377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.351096Z digest=sha256:b7d3507c04c0488c2b43ab498f547ca679c93bbe1101568c3062f1738fcb3419

Observation a7db32e6-3734-400e-ae30-4dc63aedd773 · outbound

This paper cites Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Sam Bowman, Jan Leike, Jared Kaplan, and Ethan Perez.

Activation Reward Models for Few-Shot Model Alignment Denison, John Schulman, Arushi Somani, Peter Hase, Misha Wagner, Fabien Roger, Vlad Mikulik, Sam Bowman, Jan Leike, Jared Kaplan, and Ethan Perez

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.344698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.357004Z digest=sha256:b19c9ead2b9eef0024dcffe091089e73926d77ab0b635f4acf2dd2f82723c848

Observation bcab6f08-6cb1-457a-992b-c32bdb4344f6 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Activation Reward Models for Few-Shot Model Alignment Christiano, Jan Leike, Tom B

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.336451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.359782Z digest=sha256:428a9b57e7129d44bed0b76dc1c8d811e165d10eaa13857288fa9a5055198202

Observation 333f3104-8309-43ff-a71a-fa32a444418a · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Activation Reward Models for Few-Shot Model Alignment Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.362383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.362383Z digest=sha256:caf2f08851ff44af06230870b820bad793d28e333b7b82a13dd0928fbf13c432

Observation 0b8325f5-f3fe-49db-925f-63231a3714a4 · outbound

This paper cites Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking.

Activation Reward Models for Few-Shot Model Alignment Helping or Herding? Reward Model Ensembles Mitigate but do not Eliminate Reward Hacking

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.365297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.365297Z digest=sha256:583ff142dc53ba6ca0d0187fc684af3c602d1a4665218e738d870825771c357c

Observation dd0d2183-911c-4a10-84e4-5364a54d0916 · outbound

This paper cites Paint by Word.

Activation Reward Models for Few-Shot Model Alignment Paint by Word

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.368595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.368595Z digest=sha256:011f6871a909c4ae15c76a8d60aeaa219d8e95c89f42beb7e8a2306b05a678f9

Observation 565ec317-f552-4704-8525-3be5fb04fa51 · outbound

This paper cites A Survey on LLM-as-a-Judge.

Activation Reward Models for Few-Shot Model Alignment A Survey on LLM-as-a-Judge

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.371326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.371326Z digest=sha256:33abc78ba1f0b2ae207a4ad8b2e6d97fa4d170a1c1aef28e5a3ba73039ae9587

Observation 7b27f7cc-25e1-485c-9580-29ab4868c157 · outbound

This paper cites M-RewardBench: Evaluating Reward Models in Multilingual Settings.

Activation Reward Models for Few-Shot Model Alignment M-RewardBench: Evaluating Reward Models in Multilingual Settings

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.374334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.374334Z digest=sha256:bd13edd44c8c78f4c19da5ec1dd267df7aed93525b75e28505cd276016be76cd

Observation 701e53e9-6d25-4896-916e-b67d26311382 · outbound

This paper cites In-context learning creates task vectors.

Activation Reward Models for Few-Shot Model Alignment In-context learning creates task vectors

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.327878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.377311Z digest=sha256:f40bfb0a2b3db6e3b1e32c05f2dd306edaf1fff579227e636b450e3713e711ec

Observation 2b2b4cff-a11b-420f-ad15-f45d39767669 · outbound

This paper cites In-Context Learning Creates Task Vectors.

Activation Reward Models for Few-Shot Model Alignment In-Context Learning Creates Task Vectors

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.379822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.379822Z digest=sha256:9b5a931fd8feb81348265377ab871bc3c7215e7c52e624374a0116c2d03db8b5

Observation c8ad2182-a253-43d3-9764-a0c2b1c06f15 · outbound

This paper cites Inspecting and Editing Knowledge Representations in Language Models.

Activation Reward Models for Few-Shot Model Alignment Inspecting and Editing Knowledge Representations in Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.382687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.382687Z digest=sha256:8a2ee3d69ab70e99975bae0f76b63ad2a0745dc7b59407543f5fffd2426ab703

Observation bff11584-e388-4db1-b95c-74ee3e411e88 · outbound

This paper cites Finding visual task vectors.

Activation Reward Models for Few-Shot Model Alignment Finding visual task vectors

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.318883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.385383Z digest=sha256:504f9b7d71aa955d4f884f656356c2c506c69aec303158e07d398ef322f8e390

Observation 26c76699-badc-4c7e-861a-fd81a78130c5 · outbound

This paper cites SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality.

Activation Reward Models for Few-Shot Model Alignment SugarCrepe: Fixing Hackable Benchmarks for Vision-Language Compositionality

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.387975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.387975Z digest=sha256:40e6eefda8b2190899468240a1d271788ec62a4800bb9c8567419602366cdbcb

Observation d8fe77af-5b37-4b9d-ac1d-607631c37add · outbound

This paper cites Multimodal task vectors enable many-shot multimodal in-context learning.

Activation Reward Models for Few-Shot Model Alignment Multimodal task vectors enable many-shot multimodal in-context learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.309709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.390963Z digest=sha256:804e63c222f456f10877e5248c32c2d5d6cfd839a6a69f48eec3fb8d10ce92cb

Observation 739cb215-0073-4d93-91f7-2a70c9221e3d · outbound

This paper cites Multimodal task vectors enable many-shot multimodal in-context learning.

Activation Reward Models for Few-Shot Model Alignment Multimodal task vectors enable many-shot multimodal in-context learning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.300705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.393812Z digest=sha256:663a613f2b8b91b9c3b23026a4916a8809e762f708b1193b966d18cf5936fcb3

Observation ca3a6e68-0545-4c7a-aadb-9fb082c0ee22 · outbound

This paper cites RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment.

Activation Reward Models for Few-Shot Model Alignment RAG-RewardBench: Benchmarking Reward Models in Retrieval Augmented Generation for Preference Alignment

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.396463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.396463Z digest=sha256:d7c735186e175ff95c2c826d9352d4d9dba8b71f5a593e4fa10be41ca0b7d83c

Observation 8acdb000-0686-4932-9d12-8726a2b6d43c · outbound

This paper cites Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes.

Activation Reward Models for Few-Shot Model Alignment Few-shot Steerable Alignment: Adapting Rewards and LLM Policies with Neural Processes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.399277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.399277Z digest=sha256:7221bc9a1796349ab47d4823b86cecf3fb9cbd74db1fb414221a405b39ec0ae3

Observation 5529c393-9ff4-4b69-8aad-fd628fabf51a · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Activation Reward Models for Few-Shot Model Alignment RewardBench: Evaluating Reward Models for Language Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.402165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.402165Z digest=sha256:46a8c5bd7ea4876736d0e1f83981039150ca6dc7b8b1d7cbf8a6a667be556846

Observation 1a064690-906b-4d33-b571-b637800c65ef · outbound

This paper cites RLAIF vs.

Activation Reward Models for Few-Shot Model Alignment RLAIF vs

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.291393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.405063Z digest=sha256:44198ade63e9fe2553a1009d260e7e47b3f3b258669da5c3b79c46474801d170

Observation 1bd11883-f8cd-44a9-b664-ddd46ee6c5ff · outbound

This paper cites RLAIF vs.

Activation Reward Models for Few-Shot Model Alignment RLAIF vs

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.281451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.407759Z digest=sha256:840ba9c86d141cfc2d66bb92aeef1bd748b042d61f92d642c528fabc50f12daf

Observation 301eebe0-a8d2-4d91-9e22-d502c04e8352 · outbound

This paper cites The power of scale for parameter-efficient prompt tuning.

Activation Reward Models for Few-Shot Model Alignment The power of scale for parameter-efficient prompt tuning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.272260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.410251Z digest=sha256:045179bce5331d0fb1b3ae90fdba5b0f3c915f1ff62c81d2f94ad6beb4b983ff

Observation d5696a21-e77d-48dd-87c1-d6158f0f5657 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Activation Reward Models for Few-Shot Model Alignment LLaVA-OneVision: Easy Visual Task Transfer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.413245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.413245Z digest=sha256:b6b7699560ef4d6ef805afc656463b942304c744c62a784bd85290f54efc4766

Observation 0ac81866-ea23-4e7a-89f0-c1a62e6a1509 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

Activation Reward Models for Few-Shot Model Alignment BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.416301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.416301Z digest=sha256:1b361298d56dd20d7478d4891b1f36f13b9425e118d6cebcfec461b42eeb69f4

Observation 0495770b-eb30-4685-bb38-1008d01ebe57 · outbound

This paper cites Evaluating text-to-visual generation with image-to-text generation.

Activation Reward Models for Few-Shot Model Alignment Evaluating text-to-visual generation with image-to-text generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.256118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.418758Z digest=sha256:754a320a2b9ea7a21bd54288aff3849fdd418f531ee6254882bf8718ff0453c4

Observation 82779433-6465-4df6-9e46-967589565a75 · outbound

This paper cites Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling.

Activation Reward Models for Few-Shot Model Alignment Inform: Mitigating reward hacking in rlhf via information-theoretic reward modeling

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.247216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.421311Z digest=sha256:977e6e08ee4e15091017f080a77e7c40d6549cf3e73fc3e53bc4bfe43ca8b0b4

Observation e87e8d74-2e2e-4461-b28c-1e519c258d34 · outbound

This paper cites Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features.

Activation Reward Models for Few-Shot Model Alignment Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.424201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.424201Z digest=sha256:11a4af0fe49f3f7bd36d275e239fab21fd2f101370b27c4f7767d927fab32238

Observation 2295d038-75a9-4a7b-92da-5174726ff034 · outbound

This paper cites Rule Based Rewards for Language Model Safety.

Activation Reward Models for Few-Shot Model Alignment Rule Based Rewards for Language Model Safety

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.427164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.427164Z digest=sha256:43bc8e016c53c51453c7d6c1de36ef280582e8019d1735f8e44d6d78735664c2

Observation f932738d-68f2-4a43-9acd-b74640b2bcbf · outbound

This paper cites In-context Learning and Induction Heads.

Activation Reward Models for Few-Shot Model Alignment In-context Learning and Induction Heads

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.429971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.429971Z digest=sha256:6d6d1b2a7c8ef00444fb4b250fe8e320b480eceb53306ce0e1dcfe0f400e6cb9

Observation 9de138df-d363-4716-a9cf-d5e70f13e740 · outbound

This paper cites GPT-4 Technical Report.

Activation Reward Models for Few-Shot Model Alignment GPT-4 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.432668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.432668Z digest=sha256:63658905441e14e47540e6d46b798197af611ec32fe02119277c303dda249f82

Observation 2e53c4cd-a86c-401b-a243-ea10fb95f939 · outbound

This paper cites an unresolved cited work.

Activation Reward Models for Few-Shot Model Alignment Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:43.238094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.435269Z digest=sha256:7550ec9caae83737523182d9521515b972b550cae8e95ac761a59e79af753aeb

Observation ed740f6a-768d-440b-ae61-43cb915c2319 · outbound

This paper cites Training language models to follow instructions with human feedback.

Activation Reward Models for Few-Shot Model Alignment Training language models to follow instructions with human feedback

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.229102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.438395Z digest=sha256:4a2464c55df11a35aaabe01a7ed8d7fbcc1faf06661530fd5cb75770274920b1

Observation 09a7c555-8945-4616-ac78-ffc005f697fa · outbound

This paper cites an unresolved cited work.

Activation Reward Models for Few-Shot Model Alignment Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:43.219895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.440917Z digest=sha256:44bea3ef1a98488bf95d3f807303c548482aa323b566fe643702b4d19180cca4

Observation 18c8f537-ed32-4557-a495-86e6bfb803a1 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Activation Reward Models for Few-Shot Model Alignment Steering Llama 2 via Contrastive Activation Addition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.443531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.443531Z digest=sha256:6644cb70aaf252f48bf796df9b25c2240a41fbd7cb01070f859ae2153d7c6b17

Observation 6ef9ecfa-3019-442a-90d9-88d37689425d · outbound

This paper cites Pytorch: An imperative style, high-performance deep learning library.

Activation Reward Models for Few-Shot Model Alignment Pytorch: An imperative style, high-performance deep learning library

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.446511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.446511Z digest=sha256:04cf5a05068b39b39bd8adccde55ff5026900b3a18f4f2ed61f256d8681e6d49

Observation bdc8313f-0b58-48f9-b918-8a89e203f6ba · outbound

This paper cites Red Teaming Language Models with Language Models.

Activation Reward Models for Few-Shot Model Alignment Red Teaming Language Models with Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.449385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.449385Z digest=sha256:4b6afded7374cd7cbb7c9b0ab0376f2a8a394429e7345b9d2aa30d83db1f7f7d

Observation 89618b88-420f-45bd-ae5d-8471c85695b0 · outbound

This paper cites Improving language understanding by generative pre-training.

Activation Reward Models for Few-Shot Model Alignment Improving language understanding by generative pre-training

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.452328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.452328Z digest=sha256:95d4578ba9e78a17b6c6b939fb34d9a53fa8a91d7f03385fdea77d9b4faa69ba

Observation f0234c10-4552-46d9-a723-ca951d5ca19e · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Activation Reward Models for Few-Shot Model Alignment Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.455001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.455001Z digest=sha256:8a5ed39b116568e4fb0cfc0f8f29589a7764bf9a0d093eb93111c449a6ed188c

Observation 9b59a616-ab3d-40e1-a961-41b0ddcf8921 · outbound

This paper cites GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation.

Activation Reward Models for Few-Shot Model Alignment GPT-4 Jailbreaks Itself with Near-Perfect Success Using Self-Explanation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.457877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.457877Z digest=sha256:9595b5112a51b607ca0556477c531f45f62e4f7e4f81b32542b5bed8973b540c

Observation 08a76f31-3f9e-4e7a-923b-1454c5381435 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Activation Reward Models for Few-Shot Model Alignment Proximal Policy Optimization Algorithms

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.461296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.461296Z digest=sha256:18db6b7905f4cdf735b6a991884d211ee72ebf21c092c507d75c1253fe0fedfd

Observation 68bbd057-2773-4fa5-a82f-ded588030d21 · outbound

This paper cites Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations.

Activation Reward Models for Few-Shot Model Alignment Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.463921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.463921Z digest=sha256:ffef1656a6ccf37daa1a35c4e0022eb42536fcdbdc1c938c0582400ccc7cd2e6

Observation a6b977d1-c958-4b0e-8354-29a257170a05 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Activation Reward Models for Few-Shot Model Alignment DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.466647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.466647Z digest=sha256:cf05d946d0e01b01b7ff25d4030bac4cac8424d26682353a0f077f03289f19ca

Observation 416fede6-7ea0-435b-b28f-6c9561e43e39 · outbound

This paper cites FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users.

Activation Reward Models for Few-Shot Model Alignment FSPO: Few-Shot Optimization of Synthetic Preferences Personalizes to Real Users

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.469412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.469412Z digest=sha256:955e0687953630d7d7bb70a1bee0de088b968234f95b3a19f0d74fe88e6b1b88

Observation 01895fb7-6d99-4260-ad20-7e6d22c9fcd4 · outbound

This paper cites Alpaca: A strong, replicable instruction- following model.

Activation Reward Models for Few-Shot Model Alignment Alpaca: A strong, replicable instruction- following model

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.199194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.472472Z digest=sha256:8f1c6e041840301e9d5384f48db248cb9222bdea36e1929092d494ae6c315261

Observation 6a231fd1-79dc-4ccd-9ef3-66cb51d24458 · outbound

This paper cites Learning to summarize with human feedback.

Activation Reward Models for Few-Shot Model Alignment Learning to summarize with human feedback

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.475288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.475288Z digest=sha256:de33ec126f4c51b1c4263c5d86d854209ed62ed6481967c278daed89b5b478cd

Observation 01348201-5f79-4333-a21a-8e69d5d86b58 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, Paul Christiano, Jan Leike, and Others.

Activation Reward Models for Few-Shot Model Alignment Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, Paul Christiano, Jan Leike, and Others

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.183491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.477978Z digest=sha256:f760f65fa842d953f1a4ff73524dd7987d9cf855607b5b9453c98e00725d5421

Observation 7597d84e-84c7-4f0c-ae61-22ea44d0e9ff · outbound

This paper cites Extracting latent steering vectors from pretrained language models.

Activation Reward Models for Few-Shot Model Alignment Extracting latent steering vectors from pretrained language models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.173972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.480561Z digest=sha256:f3c2c53912b844ab78e0e7d40da004537a23566f26837d6197438606f8b64299

Observation 6c308a2a-af65-4032-94a0-27c9dc16d7ee · outbound

This paper cites Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence.

Activation Reward Models for Few-Shot Model Alignment Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.483528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.483528Z digest=sha256:966f137e62d64cffadf708fdacba394153020a14402a0cf8b3741008f8bfc282

Observation d87fff3e-fc59-4c55-8b05-953b8182f488 · outbound

This paper cites Li, Arnab Sen Sharma, Aaron Mueller, Byron C.

Activation Reward Models for Few-Shot Model Alignment Li, Arnab Sen Sharma, Aaron Mueller, Byron C

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.163979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.486151Z digest=sha256:b46adc9ab4f3393c713ecd3e886734eb1154a4f1a27f78e84ac478b8d3d915ba

Observation 5b0d810c-d934-4b46-a0ab-e1a0b4d32077 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Activation Reward Models for Few-Shot Model Alignment LLaMA: Open and Efficient Foundation Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.488755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.488755Z digest=sha256:ae7776753471651b1ebd7de792855cb944d4b9dc8f4cc107373fa4ebe20c5974

Observation d292f8cf-a3bb-4424-bba5-1c9b73ac1628 · outbound

This paper cites Steering Language Models With Activation Engineering.

Activation Reward Models for Few-Shot Model Alignment Steering Language Models With Activation Engineering

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.491453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.491453Z digest=sha256:a0a98ffc5d4324368111925fbdb08b60e1cdcb12d33b8d45b3f02a9db8c80e7d

Observation c0166f73-c608-4948-9c7d-e922a1945039 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Activation Reward Models for Few-Shot Model Alignment Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.494467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.494467Z digest=sha256:af33696ed031bd971e7ee2275c918b01c84bd193159ca5b50cc08e09df681d41

Observation 885d840c-9f1f-4e78-bc23-e9c322835eef · outbound

This paper cites Large language models are not fair evaluators.

Activation Reward Models for Few-Shot Model Alignment Large language models are not fair evaluators

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.154796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.497393Z digest=sha256:aca0e2f7b49f66a9de8a5c70760dde30403d4a1f7357f5e00ffd9843a104fd73

Observation e8db6c18-6cf8-4e95-9d2c-fd2e7aa05d7e · outbound

This paper cites Williams.

Activation Reward Models for Few-Shot Model Alignment Williams

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.144570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.500110Z digest=sha256:e307e1ee7f8bec5144a424cde6acf9433fe74a46ccb6b8e700b2123228dac3ca

Observation 5c653327-9475-4a00-8979-ce6e5ce22713 · outbound

This paper cites rewordbench: Benchmarking and improving the robustness of reward models with transformed inputs.

Activation Reward Models for Few-Shot Model Alignment rewordbench: Benchmarking and improving the robustness of reward models with transformed inputs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.502647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.502647Z digest=sha256:7a5f60033e011f3defd02aaf6016e061f0291b453caefa6b61b4be624bd28cb2

Observation e2908560-f838-4916-a340-625aab08f338 · outbound

This paper cites Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models.

Activation Reward Models for Few-Shot Model Alignment Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.505501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.505501Z digest=sha256:d58e039bbc3d54beffad11cd7a3d1222e8f660e39143f2dc211347a20107722c

Observation b9ac0d03-a65d-4f0b-ad1a-586be8f418bb · outbound

This paper cites Zettlemoyer, and Marjan Ghazvininejad.

Activation Reward Models for Few-Shot Model Alignment Zettlemoyer, and Marjan Ghazvininejad

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.135993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.508936Z digest=sha256:ef94a88e00af53a5aef1577dd210100dd88d81f917bdd3fc5da05fd76e1e7434

Observation a898471a-c5b5-4a9c-bb5c-6c89f20c46d1 · outbound

This paper cites Which attention heads matter for in-context learning? In Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025.

Activation Reward Models for Few-Shot Model Alignment Which attention heads matter for in-context learning? In Proceedings of the 42nd International Conference on Machine Learning (ICML), 2025

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.127194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.511545Z digest=sha256:84bcc9206d508b15eb29f5c257b2190b1eae32dde9527c1e02d7b6c4c97d2fc3

Observation fdfc61a0-60bb-41b1-8192-1c5893abc6df · outbound

This paper cites ICPL: Few-shot In-context Preference Learning via LLMs.

Activation Reward Models for Few-Shot Model Alignment ICPL: Few-shot In-context Preference Learning via LLMs

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.514668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.514668Z digest=sha256:ffd20540c4a59b07bed84aa624a9afe6303e6282780cc1eed7f1fdd469e977cd

Observation 4d7b7262-f799-465f-a7e2-0578d1c3d2a1 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

Activation Reward Models for Few-Shot Model Alignment Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.517736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.517736Z digest=sha256:b9d475b5a12ff74f9a8b511c6f93c7d0eef4ae6d1b3a079dc16f5fd3c8eab08c

Observation e8f5b9bb-35ec-4f6f-84c6-d5d8950363da · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

Activation Reward Models for Few-Shot Model Alignment RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.520388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.520388Z digest=sha256:421b74fbae07e23570e1fbd4bba3a0a65e5bea3ed363f00ba2d9c871e097148c

Observation eaef4565-ec43-4aa0-8dd9-7c359a83de49 · outbound

This paper cites Rag-reward: Optimizing rag with reward modeling and rlhf.

Activation Reward Models for Few-Shot Model Alignment Rag-reward: Optimizing rag with reward modeling and rlhf

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.523148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.523148Z digest=sha256:72af9318dc9ffee2deb7b8036eafc11a66d7ae4055c66c9f19f85159fb9059b5

Observation 206120a9-5041-4a56-a862-e9e6169bdc5a · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

Activation Reward Models for Few-Shot Model Alignment Generative verifiers: Reward modeling as next-token prediction

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.110641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.526351Z digest=sha256:cf451c1bbe630c7fbb48c42a179759b53dca353a2a489594afc8e836bc5188d2

Observation 0ac44e47-5c44-4ac3-9f80-4bbcdfea8319 · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

Activation Reward Models for Few-Shot Model Alignment Generative verifiers: Reward modeling as next-token prediction

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.100892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.529059Z digest=sha256:6486a7ea913b31e1a2b7e225d6094ca4d820c4bad7cf078e7924100c349087f1

Observation 7508d138-a168-4880-8c5b-b1f4a844fd42 · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

Activation Reward Models for Few-Shot Model Alignment MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.531667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.531667Z digest=sha256:8d187ab8dc927ee4de6c6dba924b05e3eb1f1d3ccb14b58357e2f2803f952531

Observation 39407f16-bc69-4dcc-afc7-8c670e2c90e6 · outbound

This paper cites Interpreting deep visual representations via network dissection.

Activation Reward Models for Few-Shot Model Alignment Interpreting deep visual representations via network dissection

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:02:43.090900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.534698Z digest=sha256:e911102c966481ced997d2ee493dec4e707862544468cd2cfb9c06090013ca36

Observation a0b6a405-a8af-46ec-8d6b-3882c7fdc397 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Activation Reward Models for Few-Shot Model Alignment Fine-Tuning Language Models from Human Preferences

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T21:02:42.537188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:02:42.537188Z digest=sha256:91d9d82b1585530e36f263dbc854e99576a03b1387907bb98d59f2284c75b822

Observation 97ed4056-3f3b-4b9c-8b63-918eb0850aa2 · outbound

This paper cites an unresolved cited work.

Activation Reward Models for Few-Shot Model Alignment Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:02:43.352767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T21:02:42.354176Z digest=sha256:90295eebc9949a5cdcf2790672ce2ccdc3ba57007eb9f7d65ed3b55beb67108c

Pith citing papers

Observation df15135a-6ad7-4bf5-9d50-2dfb086d01af · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Activation Reward Models for Few-Shot Model Alignment

Reference 222

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.437494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:12cd515806791c500bf65a1be23af8c1aad772cc54870d57c95defa011443d86

Observation c697a69a-b83f-45aa-8222-9158dc93e9e7 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Activation Reward Models for Few-Shot Model Alignment

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.573192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:7636216e67a73d196db2baa1c550aca878438b59e758491536c0b386a01c4910

Observation e10fc225-4be8-474f-9947-e9332a7ffc41 · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning Activation Reward Models for Few-Shot Model Alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:0109d8916c6023b961eda19427ffcd067e4c45aaa27a7a55c95461c6bd8fb819