Pith. sign in

Paper Citation Record · LEDGER

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards

As of 14 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 3 inbound Pith citation observations for arXiv:2507.09104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09104 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:12:27.792351Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:38:37.969788Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:46:46.529546Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved35
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 27e4bece-34d3-4655-b59c-ecbd55ba1e41 · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:25.776024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:25.776024Z digest=sha256:e51c4004a3ece06ec77aa461a605861d3f3d86117c7992c9d9946ddaa2454899

Observation a8229587-89f9-4973-8272-a18480bbe021 · outbound

This paper cites Internlm2 technical report, 2024.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Internlm2 technical report, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.814458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:25.812538Z digest=sha256:2827dddf2d5fcd03e60c08838c435b3f51cc40f721c262480ff60b343741d8d0

Observation fd88d8ab-459a-4851-a8ea-a196281f8614 · outbound

This paper cites CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards CompassJudger-1: All-in-one Judge Model Helps Model Evaluation and Evolution

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:25.845529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:25.845529Z digest=sha256:70bf30bb71231963ab548bc659c4d08644d5856a1abf0a89c2f5562a6ae1b8cb

Observation 5a244f4a-55b8-4433-9bca-bcbeb8a535f6 · outbound

This paper cites xverify: Efficient answer verifier for reasoning model evaluations.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards xverify: Efficient answer verifier for reasoning model evaluations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:25.879190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:25.879190Z digest=sha256:1b9757e25915e1719419d03a956ae49839f62b0412940b4badefbad8abf4b4e7

Observation 5eb6f3f5-f2ea-40a4-83e4-8bca12bbe728 · outbound

This paper cites Rm-r1: Reward modeling as reasoning.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Rm-r1: Reward modeling as reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:25.913376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:25.913376Z digest=sha256:3c95422576c09fc50257472939c909e79886c3a9937827c86d707a7079cbbb19

Observation 48c74d69-9918-4ca3-b307-9e401294ac89 · outbound

This paper cites BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:25.951994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:25.951994Z digest=sha256:45b67e09904a5915c36d873fa6789f38e2328932b86f84bab4c25cc26c34085c

Observation 7930b945-138f-406a-a80f-7e223673d29a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:25.991199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:25.991199Z digest=sha256:f304e060a7335cfaa241aa6d2806354224fbb5a026fe689c5dbd2ee1ee010155

Observation dc3bf25f-644f-44e9-b4d7-b2e32884a422 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.https://github.com/open-compass/opencompass, 2023.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Opencompass: A universal evaluation platform for foundation models.https://github.com/open-compass/opencompass, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.020066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.020066Z digest=sha256:c2e4d791d2380c3e33656f74725f43e3d0d50ae8eb8068f215c96011a2a8d323

Observation d7109e07-544e-4427-875a-300a852d6170 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.036730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.036730Z digest=sha256:fdea8640d399562ccff1a7ccf60d8ce29152d25b4670a3054e6528841c2cfe36

Observation 91a8f682-4fdc-4488-aa2b-87c2bdbc8ea5 · outbound

This paper cites Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs,.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.769707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:26.074429Z digest=sha256:c62a0329ed5063b0f0a71116155cb49fd71c2329d6c6209d473d857da9c5ba5e

Observation 5f67fc17-b412-4718-9e5d-bdf7b4b5a5a1 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.136339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.136339Z digest=sha256:0a5da688f8d2c471f6b8460ddafb953e753736da38f3d3573f703c8965efc733

Observation e6d3a2aa-d945-423f-b32c-cef8de7c4e35 · outbound

This paper cites The Llama 3 Herd of Models.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.161345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.161345Z digest=sha256:dfef529f5152ec6b7dcc7b93a401a4f0bdee350e23d9b0dee4ac170463a94c5b

Observation 9134ff68-0f5c-428b-a629-e43243b37694 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.192280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.192280Z digest=sha256:8f4c98fa80da7699d5d95413da8be2647a43550b0723d8cb894a27880f75aef1

Observation 89da10b7-2134-4928-8d0a-08808659ec0e · outbound

This paper cites OpenAI o1 System Card.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards OpenAI o1 System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.225242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.225242Z digest=sha256:7d6c2dc28a64a8f08279c7a87c69b0fcb025a71122f96080789d14b546e72b44

Observation 5545b802-e87f-4476-8211-b6640beb4990 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.258243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.258243Z digest=sha256:3e6aaa86118d4fdbf3837a80e542458d5e467dd84ddd3e66231b18685842791b

Observation bf909548-303c-4f02-b969-475d1922f1e8 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards RewardBench: Evaluating Reward Models for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.293859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.293859Z digest=sha256:cb62f6f08891441d1bcc5c5307206ef4e9197d3711237b0b68702052cd24bc54

Observation 45344ba6-acfd-4334-9793-2a18459719d4 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards CMMLU: Measuring massive multitask language understanding in Chinese

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.322948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.322948Z digest=sha256:8d02c4010a517fb372c3a1579f12713cdc567bf98f64b9e0226330887b1da77e

Observation d00481c9-9872-4460-ba2a-3481dbf6bb5f · outbound

This paper cites Generative Judge for Evaluating Alignment.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Generative Judge for Evaluating Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.348410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.348410Z digest=sha256:a2169c778b9d4b54d63e140123476ed6791b68e33d8577ac45896c63f230f678

Observation d6c40686-fd6e-4553-ba9a-5cb72a9223e6 · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.373443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.373443Z digest=sha256:04dbc5aa704c72915a8c3e08ee8b67b9b79b995da0ab8dcf6ae39b475fb51a1d

Observation f42bd137-6ecf-49af-9c2d-538a32a95b5d · outbound

This paper cites WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.407047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.407047Z digest=sha256:f24cb6dfd26eac6c546fbca16c8c573fcda7421b2d73f6fc45aeb67b505241d2

Observation a5a72ff1-95cc-4ccc-b44a-811e87db24a3 · outbound

This paper cites DeepSeek-V3 Technical Report.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards DeepSeek-V3 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.432070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.432070Z digest=sha256:15f223c7613b39f8f0cc74da584ba6dc7172bf1757aa3d31d0b0d6510956a8d5

Observation 2285295d-a65f-46bd-b053-39e90300167b · outbound

This paper cites AlignBench: Benchmarking Chinese Alignment of Large Language Models.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.444272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.444272Z digest=sha256:d5716c0daabbf202d835a00f8fde8023a71e849351e41e4b52a37fd3b050f4ac

Observation 37b0be6b-e5c5-4f9f-b230-041943804558 · outbound

This paper cites Inference-time scaling for generalist reward modeling.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Inference-time scaling for generalist reward modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.459035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.459035Z digest=sha256:d8d9727007d52772bd4357a6c89e36ec4dc15bedb0f8a2f70641a84209fce858

Observation ff6b0cc0-9d3a-4009-94ad-90d0bacab3d3 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Gpqa: A graduate-level google-proof q&a benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.466847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.466847Z digest=sha256:d8c1e90b773a3d2c7bf5812f55b3e698056e7e7b9ef9f783bfa50bd6655c514d

Observation b0341f35-be9f-494e-9a5d-7d227634c76f · outbound

This paper cites Skywork critic model se- ries.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Skywork critic model se- ries

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.725029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:26.498372Z digest=sha256:ac4e0215778d12d41472ea932cea0e3d8806aa6d7edaad92cbed5a0409924ab2

Observation 68938448-2a4b-4ea9-98d3-3cad8e1b3909 · outbound

This paper cites JudgeBench: A Benchmark for Evaluating LLM-based Judges.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards JudgeBench: A Benchmark for Evaluating LLM-based Judges

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.531419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.531419Z digest=sha256:92ccb875fa668528abb76fb351e4e97c4c5ae9613a15b910bca96f679b20dbda

Observation d5e015e9-c44e-49ef-a69e-1cb7dff45951 · outbound

This paper cites Qwen3: Think deeper, act faster.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Qwen3: Think deeper, act faster

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.690604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:26.554703Z digest=sha256:abf442d34a71dd86e0f735d6094855311e08fbaa306029823a7437de103cff71

Observation 795da8ab-7e65-43f8-a394-5160207ec1b2 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.567396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.567396Z digest=sha256:b03f98355660b15edf724ebdc2988c9e6ed85e50c00e34562e02db0a71e088c4

Observation d56107c2-04cf-435c-b5e0-09f1b4efae95 · outbound

This paper cites Qwen2.5 Technical Report.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Qwen2.5 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.579312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.579312Z digest=sha256:275d1adc5f37225266d5eae05bc926bfaceb3fed876e620d94813f46d2f9b0cf

Observation 18917754-a9e0-4cd7-9bb8-634b69e74c0a · outbound

This paper cites Beyond Scalar Reward Model: Learning Generative Judge from Preference Data.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Beyond Scalar Reward Model: Learning Generative Judge from Preference Data

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.607519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.607519Z digest=sha256:442d6dacbc651e062c7c2b661e2382960f0033d3499b6684ab225ca92a79ddc7

Observation d6997f2a-8ef3-4a86-bacb-be09c4ac33fe · outbound

This paper cites Learning llm-as-a-judge for preference alignment.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Learning llm-as-a-judge for preference alignment

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.645710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:26.642755Z digest=sha256:72391256c3a80523b7c5309264ce60990966787242bfc99e2f3badb3b8ad8c7f

Observation 9bca6223-97c1-4117-b5e5-ea6569d54fe0 · outbound

This paper cites Improve LLM-as-a-Judge Ability as a General Ability.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Improve LLM-as-a-Judge Ability as a General Ability

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.670313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.670313Z digest=sha256:41098f99e0ad113b86bcc68f359a8ae4e5ae271cefc00f475658f46587f9640f

Observation 87686c92-4e7f-4a09-a47c-0fcb6213e753 · outbound

This paper cites RMB: Comprehensively Benchmarking Reward Models in LLM Alignment.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards RMB: Comprehensively Benchmarking Reward Models in LLM Alignment

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.702068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.702068Z digest=sha256:cff2930e8d70f93068a4676dd9231d7e1a0c080bc4fb2200c395bbe9e9af5f15

Observation 5f1839a8-3f0f-422d-b8dd-b5106ee3dee2 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Instruction-Following Evaluation for Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.737680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.737680Z digest=sha256:61d677894643724d1a7d04b3e815af0a286d7a4c79f470b2e2c656c23220d1f0

Observation f79f2ab3-661c-4cde-afaf-718d10fe32d5 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 35

Resolution
malformed identifier
no resolver link, observed 2026-08-06T18:12:26.763993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.763993Z digest=sha256:6bf695b4ac8055132f25cefd9d00dc117c986dee5d99d2f7e7adee1c8f10d0a9

Observation 39290a96-bdbd-414f-8930-1f51ff5d31fe · outbound

This paper cites an unresolved cited work.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:12:29.612393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:26.799679Z digest=sha256:fcd1a5cc38fdef4b01a6e2b0286651fe4ac50fc5a8bc868deadfeb7ba3498198

Observation 7bd4b6f4-784d-46d7-a72e-30a8c34b6ebd · outbound

This paper cites Consider how well it addresses the user’s demand, meets the user’s constraints, and how well it serves the intended purpose.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Consider how well it addresses the user’s demand, meets the user’s constraints, and how well it serves the intended purpose

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.578949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:26.834630Z digest=sha256:f8b0e551c9d9939055d689bfb2e0d8d6a6bb07c2ce33c7c82aa7b0d41a6176e7

Observation 69757312-1fbb-4263-b5dc-45a1a4838573 · outbound

This paper cites What aspects of the response fail to meet the user’s request or constraints? What could have been improved?.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards What aspects of the response fail to meet the user’s request or constraints? What could have been improved?

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.548703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:26.868124Z digest=sha256:d396b4b30bf3c7ebd59d9e85c09bd26b72f5cf18e1f6b8fb929ef78860904f30

Observation 4fc1365d-f39e-4d45-b641-6d7a7e2ab138 · outbound

This paper cites Consider how well it addresses the user’s demand, meets the user’s constraints, and how well it serves the intended purpose.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Consider how well it addresses the user’s demand, meets the user’s constraints, and how well it serves the intended purpose

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.510120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:26.888800Z digest=sha256:8596cb26e760e47f3c4c3a470d524892dce2f222042420c615ad8d7e4603beae

Observation 19933869-5ee9-4471-9f7f-896d60fd0d03 · outbound

This paper cites What aspects of the response fail to meet the user’s request or constraints? What could have been improved?.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards What aspects of the response fail to meet the user’s request or constraints? What could have been improved?

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.457387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:26.921244Z digest=sha256:8b2d15a8492d02e9363f2d01a785c49cf2f468c13b00ddb82d1df5084610786e

Observation cafbf833-f115-4365-afba-5190523b1938 · outbound

This paper cites Discuss which model’s response is more suitable given the user’s request and con- straints.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Discuss which model’s response is more suitable given the user’s request and con- straints

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.428329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:26.952087Z digest=sha256:25c5c1bf05dd32c1baf410237bfc45b3fd81f96ebcf274898b5e0a82bd0597f9

Observation 5e250681-443e-4915-9fe0-d0e4e3906f9d · outbound

This paper cites User’s Demand.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards User’s Demand

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.391320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:26.988135Z digest=sha256:0285e33228c638bcb0ccfb233554cdaa14c77951347e1b15e0a5a962e89de653

Observation 97afeb97-ec0c-4f5a-9392-f99d8ad9989d · outbound

This paper cites hushed, waiting world.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards hushed, waiting world

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.165270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.190901Z digest=sha256:6a4fdf752eaca2770d4c3d4acea423b3c23d018c0fea1be522ca6676378dd512

Observation 9b1d7471-1463-420d-ad97-a81dc080490b · outbound

This paper cites an unresolved cited work.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:12:29.126628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.211614Z digest=sha256:50897da0bce4606d7089101c878bfd9187ce8416bdfa0c7e462c1c085f07211d

Observation 4f1ecb96-0d57-4b53-a581-8d8a9af093e7 · outbound

This paper cites winter" or.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards winter" or

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.085792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.237440Z digest=sha256:ef5998efc7304a91cd87f9b281d3abbc9da4249e1800a2460b4479d81739991b

Observation cb8f08da-78c5-4932-af6a-d324e3f48bad · outbound

This paper cites Frost paints silent trees.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Frost paints silent trees

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.039194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.273350Z digest=sha256:173f4e0d9d2fe3ae604798a431249983c92233921d846af08bdf2db56df5fee2

Observation a2dc8b3e-b321-45df-ad93-d42e2c5a7526 · outbound

This paper cites Areas for Improvement:.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Areas for Improvement:

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.000972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.307975Z digest=sha256:7150f24712cb75d7e72006d02f1abcb0558062061422c009453f3eb4f842b256

Observation f5f50768-320c-4ffe-b633-fb24295d1e6a · outbound

This paper cites an unresolved cited work.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:12:28.959781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.343975Z digest=sha256:bba489db2296d65cb384de90f06fcf7eb9c0f13bf0f68f70d8bf3c43f0c40b10

Observation 4dabcec5-a205-4020-a9fb-8001799c2277 · outbound

This paper cites Hushed, the world awaits.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Hushed, the world awaits

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:28.925905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.379650Z digest=sha256:bdb2040216075d98bd81a7ebc0af7cd03fb1f772f7a1c6dcf63d6e8b5bdac3fc

Observation da3599d4-4a77-42ad-9dd7-21a63ff4dca4 · outbound

This paper cites Overall, the model’s response is a well-crafted poem that meets most of the criteria set by the user’s request.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Overall, the model’s response is a well-crafted poem that meets most of the criteria set by the user’s request

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:28.888987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.406197Z digest=sha256:7ba6dc1f2190c7763f61c1b13e90f1c7cf03c2fdd808d02b4dff4ad9a0cd6801

Observation 90a7f719-402d-4232-af37-03bd184a0d68 · outbound

This paper cites an unresolved cited work.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:12:29.360097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.442401Z digest=sha256:9408a87895c1fc308689bb30683bb8a70fce3282ef6df2001bcee76bef183549

Observation 3760ca5f-29b7-492b-bafc-66a1f6127621 · outbound

This paper cites winter" or.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards winter" or

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.326832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.474519Z digest=sha256:cd17288e3d54543985820f2fcc26afabbcbc43b4b01072079fab29953e292957

Observation 81617923-d253-4bcd-84ee-d0bd20ab08c7 · outbound

This paper cites Footsteps fade on paths.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Footsteps fade on paths

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.290702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.498619Z digest=sha256:6b37469041a327b03a0de030d9f56c34bc88ea9ca5d62c0836e87ef6cf2c8954

Observation f99abb23-5c3e-467b-9bb5-56e6d605ded9 · outbound

This paper cites Snow": While the user specifically asked to avoid the word.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Snow": While the user specifically asked to avoid the word

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.249162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.524571Z digest=sha256:98c5cd6fac82067eeb03c3bf2653554aa2d70706ec78a84a2daf0841262f9a4a

Observation 732feab5-0c52-4947-9755-1c654ff4287d · outbound

This paper cites Introducing a bit of variation in sentence structure could add to the poetic quality, such as using a question or exclamation to create a different tone or emphasis.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Introducing a bit of variation in sentence structure could add to the poetic quality, such as using a question or exclamation to create a different tone or emphasis

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:29.211427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.557622Z digest=sha256:ce2e67333e750e7d54002d3ebc48e850065cd20d5aaa85de4c62284c9fc8a5b2

Observation 18639146-0e69-446c-92e3-3d1c6358ea0f · outbound

This paper cites hushed, waiting world.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards hushed, waiting world

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:28.843893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.585953Z digest=sha256:f4dc852d3f6a6662a1611a1f5e712fb425e15e3f5982c923f38ea0c9d2021ade

Observation 085ef6eb-93af-49e5-8768-f947b4eee1c3 · outbound

This paper cites winter" or.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards winter" or

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:28.802136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.621485Z digest=sha256:e8afa9017b9aa8d70e248d1581712d0f715c1b1181db787ab6feeb047a84b74d

Observation 2f2578bb-fb20-47d0-ab72-5fd11d05d2d5 · outbound

This paper cites Frost paints silent trees.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Frost paints silent trees

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:28.759702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.656931Z digest=sha256:2fcf239b6c23b3346b57640b202f9041319f82d213e7f902d2a25ef0b433b599

Observation 30f937ea-45dd-497a-9ac0-ab57082168cb · outbound

This paper cites Areas for Improvement:.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Areas for Improvement:

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:28.728416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.691320Z digest=sha256:60bd1792569b0a42609187e8b40fa7364257d592cda2b7e4e3a4ecbce4036aef

Observation fe1a8d83-464b-4755-951f-ae615343ce38 · outbound

This paper cites For example, including different sensory details (e.g., sounds, smells) could make the poem more immersive.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards For example, including different sensory details (e.g., sounds, smells) could make the poem more immersive

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:28.673946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.712977Z digest=sha256:63e4c55655fa915baf5b02ca77af6b6515f47ec5f3e70e8370118b58e3492af5

Observation fcdebfd1-4cf9-403f-b87a-12c4d9483e16 · outbound

This paper cites For instance, a line that hints at nostalgia or anticipation could deepen the reader’s connection to the season.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards For instance, a line that hints at nostalgia or anticipation could deepen the reader’s connection to the season

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:12:28.626009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.747083Z digest=sha256:7136fab68b9413f8ea9a60c09784931466e2a57298d22c01ee96111bb9cdf705

Observation 1fe53379-46ae-4626-bee2-c05fbcf51fc7 · outbound

This paper cites an unresolved cited work.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:12:28.584414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:12:27.792351Z digest=sha256:cb670fd09ef40f6227bb3c733decd5b5d2afc5a40efa5d4b6564ba1e143bed82

Observation 88afcf16-237b-40ed-9de4-73d8854d0102 · outbound

This paper cites DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.105107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.105107Z digest=sha256:fd3858571de658f7f0fc188076da4a90ac922ff5c47d0154f9c6d6b169a22a7d

Pith citing papers

Observation 09e18a57-e8f8-4d76-8cb1-f03086c96058 · inbound

Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance cites this paper.

Think-with-Rubrics: From External Evaluator to Internal Reasoning Guidance CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:56.220230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:50:50.037018Z digest=sha256:dd94cb6e21ff2a948f27e64872dda6af864616e44e2a2a1fdd9bd8e3729a7c5f

Observation 1a759a44-cc4c-4f97-bd61-83c79b70ff22 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards

Reference 193

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:57:17.294644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:08898cb726280cd0a9c15475fe1acb8ea16d88376f79a26e6646ee4bcd9d4c51

Observation ce92f95c-8114-4434-93a2-f09526416a7c · inbound

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification cites this paper.

SCI-PRM: A Tool Aware Process Reward Model for Scientific Reasoning Verification CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:46:46.530940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T06:38:37.969788Z digest=sha256:ced85e54b069d5ce97f9da6254e26b04d544d774f160cf6f49e09a9869f83f75