Pith. sign in

Paper Citation Record · LEDGER

Tiny Reward Models

As of 19 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2507.09973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09973 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:48:07.210826Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb79d5da-4c7a-40e6-b4d7-0ed61c769ad0 · outbound

This paper cites write newline.

Tiny Reward Models write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.089065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.089065Z digest=sha256:71c57c526a89726f51a58bce6176bf59b4b42bd4f4aea048e77396f2b7842294

Observation 5a83696c-c906-427c-90cf-c33df3bad179 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Tiny Reward Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.092692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.092692Z digest=sha256:7ec6514d9cc52a6af7bbbd3c8128779258d0fdd1277aed0847430b5c85ea4d22

Observation 9fc9e5de-1c04-4e9a-8616-6d0672d9b7d1 · outbound

This paper cites PPL-MCTS: Constrained Textual Generation Through Discriminator-Guided MCTS Decoding.

Tiny Reward Models PPL-MCTS: Constrained Textual Generation Through Discriminator-Guided MCTS Decoding

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.695107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.095711Z digest=sha256:8c0e900a5bcfc5bb8dd32f972d49161494e756d69fc3fd03daecb9305d907307

Observation 86b77f26-9ca0-4d57-8d65-8bcf86e0ee1e · outbound

This paper cites Rm-r1: Reward modeling as reasoning, 2025.

Tiny Reward Models Rm-r1: Reward modeling as reasoning, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.098573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.098573Z digest=sha256:27a99b03599b27fa458160a9ca58a0660372078a192d5c3ff966bf770ff9f298

Observation 8b4b6dfb-eff9-41c9-a1a2-99a9b1cd50b8 · outbound

This paper cites B., Martic, M., Legg, S., and Amodei, D.

Tiny Reward Models B., Martic, M., Legg, S., and Amodei, D

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:07.718376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.101195Z digest=sha256:79c444e700a1d10daad515d2640024494ad2b8e2b9587671032a3f1b827da88c

Observation 5027851b-45cc-41f4-9282-970a71d218c4 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Tiny Reward Models Scaling Instruction-Finetuned Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.103616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.103616Z digest=sha256:00fa07dde124b2139c6595daf478ae57c3c2c683247abc068dbb18be7ee4a1e6

Observation f86633c8-6b4d-4265-b4bc-f7dbf11bcbe6 · outbound

This paper cites It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers.

Tiny Reward Models It's All in The [MASK]: Simple Instruction-Tuning Enables BERT-like Masked Language Models As Generative Classifiers

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.607922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.106359Z digest=sha256:76be0e6af35c2b262118d214e6e023c13acecfeb8c2b2e53fa14b8a3370bebb5

Observation c0138e18-c753-41df-a667-1b813b9a507e · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Tiny Reward Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.109354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.109354Z digest=sha256:3979d2c9babe786827fac15cbf9c40cc826f9824029b0ebd298e7f2556f70d64

Observation 7ffd1b75-1817-4897-87f9-9b0a4fe3efaa · outbound

This paper cites TinyStories: How Small Can Language Models Be and Still Speak Coherent English?.

Tiny Reward Models TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.111658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.111658Z digest=sha256:6c79f3c4e71462c88b184cc108ead6e0f16d11475600a58de67c4a8ceda40746

Observation f98dd840-a992-4e95-b9f3-932fd9983cac · outbound

This paper cites L., and Vigliocco, G.

Tiny Reward Models L., and Vigliocco, G

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:07.710863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.114627Z digest=sha256:40dbc5863007d035671b8e359a7461402546a2dda9650cdcfb6765fab4a4d636

Observation 17de84cb-7b24-49a6-ae43-762183081e98 · outbound

This paper cites Scaling Laws for Reward Model Overoptimization.

Tiny Reward Models Scaling Laws for Reward Model Overoptimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.117097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.117097Z digest=sha256:a5b93e6d387b7aca55b961bbd440345a8d091b76929116a49c7c3effb2cdbf50

Observation 07f91a81-d6f9-4dca-9acd-b5e61a593cd5 · outbound

This paper cites Making pre-trained language models better few-shot learners.

Tiny Reward Models Making pre-trained language models better few-shot learners

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.119680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.119680Z digest=sha256:8ed48b5beece2721d404b1665320fc223f41740888c2d155f8478461efbd580c

Observation 02458464-c5a9-4475-8514-9a64609c1b63 · outbound

This paper cites Towards Data-Efficient Language Models: A Child-Inspired Approach to Language Learning.

Tiny Reward Models Towards Data-Efficient Language Models: A Child-Inspired Approach to Language Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.575213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.122038Z digest=sha256:8d1f7b96dc4c48b6ab46eb9a3547390a299d50256a3954d5ff557934cfd6a743

Observation 80c8c4af-44f5-4f4f-8a0c-f6eb2c026317 · outbound

This paper cites The Llama 3 Herd of Models.

Tiny Reward Models The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.124633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.124633Z digest=sha256:d8f56bcace21450e432772bde737e242619774bef10e7a01618535ebb922d1c7

Observation c11d3822-6db0-494b-a4f7-02e261dfed68 · outbound

This paper cites Textbooks Are All You Need.

Tiny Reward Models Textbooks Are All You Need

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.127115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.127115Z digest=sha256:dbe43da8af644430cbc16363d08675e70af706fea5cf469913ea53e8a34b6e78

Observation a15d9dde-7f19-4637-a4f5-ecb0c0c58b8e · outbound

This paper cites Does RLHF Scale? Exploring the Impacts From Data, Model, and Method.

Tiny Reward Models Does RLHF Scale? Exploring the Impacts From Data, Model, and Method

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.129543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.129543Z digest=sha256:7b5f3c23a047be6cc53afc0dd2f358b718681590336e22b14454175658eb417e

Observation 841364ec-0573-4c03-a26a-f6ec90b140b8 · outbound

This paper cites Universal Language Model Fine-tuning for Text Classification.

Tiny Reward Models Universal Language Model Fine-tuning for Text Classification

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.131920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.131920Z digest=sha256:d0e77c9d2ff3edd2fd2237b879970d27117f9d6fe111b06b5f92827ae356926e

Observation db601eee-593a-4c61-a1d8-3d38dfbc423c · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Tiny Reward Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.134619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.134619Z digest=sha256:52881289d3192d22cc530d126622c02b6bf63bf6bd45ec2d1debad381f8c06a5

Observation 42fe3188-5383-423d-b63f-299496c0b3a7 · outbound

This paper cites PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference.

Tiny Reward Models PKU-SafeRLHF: Towards Multi-Level Safety Alignment for LLMs with Human Preference

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.137146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.137146Z digest=sha256:6bb669388a7dc1f7ad3f0342cdda2e64695c569da96cd79bac12aed355df2f9b

Observation 53dde41c-fc6c-4f35-97a7-5e515f38a3fa · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

Tiny Reward Models OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.139475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.139475Z digest=sha256:73799c3c110a7894d8958418a03ff2c0937699f0496e9f68986eeebe6515733e

Observation 459455e7-a116-4b41-a7fc-2cb1cfac5cc5 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Tiny Reward Models RewardBench: Evaluating Reward Models for Language Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.141889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.141889Z digest=sha256:1262f1a3b4be4b5dc15ccfee8a73c7ee0f9bd9915e3feb715759538e5c6cd82d

Observation 8156c2eb-62bd-4796-9ebe-ca1460e7ab41 · outbound

This paper cites What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning.

Tiny Reward Models What Would Elsa Do? Freezing Layers During Transformer Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.144511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.144511Z digest=sha256:5fb1ca239e79b29588705c2ff1cbf06dfe783865330d924883b67bb8982ef61b

Observation 8e9201f9-ecad-4092-8e74-1c7c14d616c4 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

Tiny Reward Models BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.146902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.146902Z digest=sha256:ea09eeee654fd91a54baa0f557da395de39cf420d6ced6a5c003c8688e4b8174

Observation 5a54112a-3525-44c0-9a96-08ef27a9f9dc · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

Tiny Reward Models Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.149469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.149469Z digest=sha256:7c9dc20e832862cc5910d063a0a57fb64e5411f3fb48a39b5573de7d927236c3

Observation 34cc45f9-ca75-409f-9df3-578986314c80 · outbound

This paper cites Let's Verify Step by Step.

Tiny Reward Models Let's Verify Step by Step

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.152003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.152003Z digest=sha256:e263a71d5a54f344bf79c712af00b6c6d91049e74b2d9a5ba16fa2ff63243698

Observation e99e856f-6e33-4b1c-817e-67625d4bcaca · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

Tiny Reward Models Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.154579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.154579Z digest=sha256:4bb7d0b35aca2f3b6f633a6b69af2e89f928808892bbefc1e8742a1e5e422d7e

Observation 8c8ae2a7-6cf2-4a81-9f5d-2416827e6c61 · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

Tiny Reward Models DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.157051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.157051Z digest=sha256:3b849f0fda033c5a5eb5c06306e1413d7a3959119077a232034b77cf1cd65150

Observation e8578aff-60bd-4b87-a2f6-fbdf9c4f6b1c · outbound

This paper cites Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models.

Tiny Reward Models Routing to the Expert: Efficient Reward-guided Ensemble of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.159727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.159727Z digest=sha256:3475ae870f957123cc05fb69393a9e4973249cf0c35c2c4396b246b7aae73d79

Observation 7cb9e172-efc3-4fa9-932a-ea5822a66474 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Tiny Reward Models Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.162023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.162023Z digest=sha256:de2f079e4cdca57b81e1cdae0ab683d0750522b12c7d1a8a9c99ad4ccacf69ad

Observation f02cbb16-bc9d-4043-8623-29554abc7501 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Tiny Reward Models WebGPT: Browser-assisted question-answering with human feedback

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.164185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.164185Z digest=sha256:a96c508356e926ca08e279f141c079a413c822d8e323b00f6ae261396114146a

Observation 3279575f-3339-45b6-9782-4f22b9fa39ce · outbound

This paper cites Passage Re-ranking with BERT.

Tiny Reward Models Passage Re-ranking with BERT

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.166590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.166590Z digest=sha256:115a1187e3e511adc038e6f5211d3a3be6c6bbdd1029770074343bc7df8a547a

Observation 57c0428d-05be-40a2-bc31-aae00aa506bc · outbound

This paper cites Nemotron-4 340B Technical Report.

Tiny Reward Models Nemotron-4 340B Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.168990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.168990Z digest=sha256:303c75dc0c803fc4edf6650cee4cb0e36a393d5dfd860137c9acc8682c1c3fb1

Observation 6a490881-e7e0-4992-9191-c5b2b4fd2089 · outbound

This paper cites Training language models to follow instructions with human feedback.

Tiny Reward Models Training language models to follow instructions with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.171387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.171387Z digest=sha256:95b9864937ff43ddd22f67c3d2ce78a5741e12e8db5cac21410e8cbe81ab6ce7

Observation 7a8c7288-d584-4bb4-95e9-f3b0ba4fb327 · outbound

This paper cites Let's Reinforce Step by Step.

Tiny Reward Models Let's Reinforce Step by Step

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.173925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.173925Z digest=sha256:1db8d0d63ff8b9c8347985f8805e81dee07008ac8a825f61ed99b49ef0a28960

Observation af90b18e-5bee-4183-85a5-62a94e9b2624 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

Tiny Reward Models Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.176314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.176314Z digest=sha256:32d415e86d0534120badd17872499f388eed418f705ca00e5b21c8d7cc915897

Observation c96758dd-57fc-40a8-9c3f-1eaab651fd41 · outbound

This paper cites WARM: On the Benefits of Weight Averaged Reward Models.

Tiny Reward Models WARM: On the Benefits of Weight Averaged Reward Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.178654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.178654Z digest=sha256:b17f497ae3356c830dfbf83ec957af5bf07ceb28746a5aee8cf82d7cfca96771

Observation e265842c-7e61-40cc-afc5-6409dd85f99f · outbound

This paper cites BERTs are Generative In-Context Learners.

Tiny Reward Models BERTs are Generative In-Context Learners

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.181168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.181168Z digest=sha256:44b08040633b029af4c17b9c79b3539080389d87a4f758b75f7b5c334f6eeaf0

Observation 55c7ffa2-a916-4d5a-9f42-37c62c029a48 · outbound

This paper cites Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference.

Tiny Reward Models Exploiting Cloze Questions for Few Shot Text Classification and Natural Language Inference

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.183516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.183516Z digest=sha256:66c7c1a604fb2edaabace63f002012e6a9e9e0dde2196271d26f57c7521d91fa

Observation abb3c454-889b-487e-9e31-95e36072bc82 · outbound

This paper cites It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners.

Tiny Reward Models It's Not Just Size That Matters: Small Language Models Are Also Few-Shot Learners

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.185809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.185809Z digest=sha256:8ce9d87aa7504af67feb2b95e1e7efc87146d3b9d28ca7dc871730a1dd3927ac

Observation 2237a1c5-ba21-4367-b8f6-dc294cd7fbbc · outbound

This paper cites The Trickle-down Impact of Reward (In-)consistency on RLHF.

Tiny Reward Models The Trickle-down Impact of Reward (In-)consistency on RLHF

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.188338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.188338Z digest=sha256:cd19da7b00b5491a115c75f789398b77dad27d9cebe9cc03e08fa0abdf54bd4e

Observation 986bf3e7-c475-4689-8ffb-a0a7508a53d1 · outbound

This paper cites Layer by Layer: Uncovering Hidden Representations in Language Models.

Tiny Reward Models Layer by Layer: Uncovering Hidden Representations in Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.190751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.190751Z digest=sha256:0d3e92562095a8136559bdab84d4f459fdd428c397ba9cc5b39602e1ea0af66d

Observation 60319909-5b82-4069-bfa0-023280b205c8 · outbound

This paper cites D., and Su, W.

Tiny Reward Models D., and Su, W

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.193757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.193757Z digest=sha256:277d96b5f98aedeb79dafe2f1a1fb9688a2bb49f976bf172d58ee933ef5a4a7a

Observation faba13a6-cc57-4e9c-9c9b-8f8b81993a6d · outbound

This paper cites Large GPT-like Models are Bad Babies: A Closer Look at the Relationship between Linguistic Competence and Psycholinguistic Measures.

Tiny Reward Models Large GPT-like Models are Bad Babies: A Closer Look at the Relationship between Linguistic Competence and Psycholinguistic Measures

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:48:07.285710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-06T17:48:07.196004Z digest=sha256:f6a484bbf1a8acde0dd37d39a88d5adba02a75b25a12a660b84c67997de9856c

Observation d63fe4c4-efae-4003-a284-2e58a6e4d6a0 · outbound

This paper cites How to Fine-Tune BERT for Text Classification?.

Tiny Reward Models How to Fine-Tune BERT for Text Classification?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.198467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.198467Z digest=sha256:840774dc3d43402f1dacd3b6a0ea1502778ae5706b5c38e1bd5fd82d70f85cce

Observation 5be4873b-40e3-45a8-896b-91a648f1633d · outbound

This paper cites HelpSteer2: Open-source dataset for training top-performing reward models.

Tiny Reward Models HelpSteer2: Open-source dataset for training top-performing reward models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.201291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.201291Z digest=sha256:05bae08fff7e1bca2e404c4f1b80ea58d0c158be59923d7619c42d2ad04bab7b

Observation 79311ea0-1930-4be7-87c8-9a9fc00b6d25 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

Tiny Reward Models Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.203840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.203840Z digest=sha256:3540ca7e709f1bb927862bff7bf7c1c0dce623fe295493e3f16a079bc8c72951

Observation de32fa4b-0030-450f-a9c6-4897c78b760c · outbound

This paper cites Finetuned Language Models Are Zero-Shot Learners.

Tiny Reward Models Finetuned Language Models Are Zero-Shot Learners

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.206091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.206091Z digest=sha256:06bfa169513ea7c800d741f9482de8521dacfb6ad57549a7d5ca6f4372816895

Observation a71856fb-f940-4e8f-8098-d0b4794b6b94 · outbound

This paper cites MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences.

Tiny Reward Models MetaMetrics: Calibrating Metrics For Generation Tasks Using Human Preferences

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.208459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.208459Z digest=sha256:bd745c6d86000bbc642579017816f2ebbaedb4396692ea574480c0f22bbb16a9

Observation b5ece352-57bb-4918-bb9e-91cd6961c748 · outbound

This paper cites How transferable are features in deep neural networks?.

Tiny Reward Models How transferable are features in deep neural networks?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:48:07.210826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:48:07.210826Z digest=sha256:aef6dc34204849309316a170724e1daa1f5b202a83bb1a5f605f33a7f6737f8d

Pith citing papers

No inbound Pith citation observations are available.