Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T13:47:45.942549Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 12 inbound Pith citation observations for arXiv:2501.17195.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T13:47:45.942549Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:50.298845Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T06:49:38.206304Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5955838e-8963-437a-8e44-d5305b984b90 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Constitutional AI: Harmlessness from AI Feedback
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 924147ce-03bf-488d-8607-451c7c70fadb · outbound
Atla Selene Mini: A General Purpose Evaluation Model Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e59109-531f-47ce-bf4b-6e89b8147eda · outbound
Atla Selene Mini: A General Purpose Evaluation Model From generation to judgment: Opportunities and challenges of llm-as-a-judge
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d519f7cc-c2a7-402f-b96e-64bc54789c1f · outbound
Atla Selene Mini: A General Purpose Evaluation Model Offsetbias: Leveraging debiased data for tuning evaluators, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6c88dab8-23fa-4195-bb04-d7bf001ddc41 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Self-preference bias in llm-as-a-judge, 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cac79179-7a51-442e-b6f6-01a4143307a7 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Judging the judges: Evaluating alignment and vulnerabilities in llms-as- judges, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a5d8d982-0432-454c-835c-fff7d9de49e3 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2390a24-5fb5-4bac-a135-ebf499258bde · outbound
Atla Selene Mini: A General Purpose Evaluation Model Flow judge: An open small language model for llm system evaluations
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2869b014-b550-48a7-bd40-2d92ac4ddaf2 · outbound
Atla Selene Mini: A General Purpose Evaluation Model GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d408f67e-3710-4adc-ac61-8548b048e9de · outbound
Atla Selene Mini: A General Purpose Evaluation Model Direct Judgement Preference Optimization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79e9433a-0aff-4a2c-b88b-8e9db9716d59 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12270fef-1e97-4658-95ee-b5862f1abbf8 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Judge arena: Benchmarking llms as evaluators
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 09b82bed-2fbb-47d1-b17f-9d9bc1733cac · outbound
Atla Selene Mini: A General Purpose Evaluation Model Iterative Reasoning Preference Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1bdb19a-0c51-4eb1-9ac9-686627d8924a · outbound
Atla Selene Mini: A General Purpose Evaluation Model Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e00fff85-a96b-423f-9fbd-6aaf8643d522 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Xing, Hao Zhang, Joseph E
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b353b3fb-05f6-4310-b62c-8513b57a9903 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Flask: Fine-grained language model evaluation based on alignment skill sets, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c800eb55-f596-45a7-a26b-c4b0e6da3766 · outbound
Atla Selene Mini: A General Purpose Evaluation Model The biggen bench: A principled benchmark for fine-grained evaluation of language models with language models, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 70d8b0b6-2769-4a06-8e26-40ac8621153f · outbound
Atla Selene Mini: A General Purpose Evaluation Model Smith, and Hannaneh Hajishirzi
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50873398-0189-4be5-8d51-ad421f06022e · outbound
Atla Selene Mini: A General Purpose Evaluation Model A critical evaluation of evaluations for long-form question answering, 2023
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4bb91bf5-4d9a-4df5-ad71-464c7a70840f · outbound
Atla Selene Mini: A General Purpose Evaluation Model A general language assistant as a laboratory for alignment, 2021
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abc8f2af-4af7-4261-8d7b-c02e1e611dad · outbound
Atla Selene Mini: A General Purpose Evaluation Model Fabbri, Jiawen Chen, Yilun Zhao, Simeng Han, Shafiq Joty, Pengfei Liu, Dragomir Radev, Chien-Sheng Wu, and Arman Cohan
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6b21a0c0-0721-4d53-a68e-04f0535a4f07 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Generative judge for evaluating alignment, 2023
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 146b0000-1fe2-4c00-92bd-13c45b35aba3 · outbound
Atla Selene Mini: A General Purpose Evaluation Model InFoBench: Evaluating Instruction Following Ability in Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99354afd-5543-4c5d-8395-9c962c2ebbb4 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Minicheck: Efficient fact-checking of llms on grounding documents, 2024
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 871b8494-0aa9-4bb1-9b05-6f1697567fb6 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Craft-md: A conversational evaluation framework for comprehensive assessment of clinical llms
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d860d53e-dff1-4508-a5d6-5993a8ef79a6 · outbound
Atla Selene Mini: A General Purpose Evaluation Model FinanceBench: A New Benchmark for Financial Question Answering
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a849a26e-062b-4851-8560-102b7525e6fc · outbound
Atla Selene Mini: A General Purpose Evaluation Model Systematic evaluation of llm-as-a-judge in llm alignment tasks: Explainable metrics and diverse prompt templates, 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b807122d-c6e6-4330-8f30-2add5ee06b06 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Does prompt formatting have any impact on llm performance?, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4b79dd25-cc73-46e2-a05c-1847f9c7df36 · outbound
Atla Selene Mini: A General Purpose Evaluation Model The comparative trap: Pairwise comparisons amplifies biased preferences of llm evaluators, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dd667c32-6ed3-4b0b-aa00-338f0240b428 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46e1c7a0-247e-4f59-a187-02305a63d8f6 · outbound
Atla Selene Mini: A General Purpose Evaluation Model OpenAI o1 System Card
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa086e8d-5142-4f1e-aef1-7db168a90d00 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 05bc346a-6d67-43c4-82d9-a00004dc7779 · outbound
Atla Selene Mini: A General Purpose Evaluation Model Nomic atlas
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 198cdc71-4ff9-4f27-a514-e71f8c8d65c3 · outbound
Atla Selene Mini: A General Purpose Evaluation Model "Dear Readers, <omitted for conciseness> P.S. No garden gnomes were harmed in the writing of this book
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 10d6e181-a23c-4326-a1ea-b681e34819c8 · inbound
Reward Reasoning Model Atla Selene Mini: A General Purpose Evaluation Model
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 473fb600-3bd8-4f37-939b-d734da71ea0f · inbound
ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments Atla Selene Mini: A General Purpose Evaluation Model
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b21ac2cc-0103-47d7-a5f5-69cad4632d49 · inbound
PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Atla Selene Mini: A General Purpose Evaluation Model
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 186743ae-0d1a-452e-8e87-9a8674819b32 · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Atla Selene Mini: A General Purpose Evaluation Model
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 633b275c-4d7e-406b-9fe7-66f5b8e3a17a · inbound
Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization Atla Selene Mini: A General Purpose Evaluation Model
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8646f0f6-5d56-4f6d-8d75-6de409b7d501 · inbound
CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems Atla Selene Mini: A General Purpose Evaluation Model
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2c34bab5-fb6a-464f-8d10-984958b188a9 · inbound
VERDI: Single-Call Confidence Estimation for Verification-Based LLM Judges via Decomposed Inference Atla Selene Mini: A General Purpose Evaluation Model
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2d4c3f1a-13ab-4a46-ba7c-0ad48e2eb519 · inbound
A Finite-Calibration Regime Map for LLM Judge Panels Atla Selene Mini: A General Purpose Evaluation Model
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b61684d5-7b12-4020-adf8-8846be308705 · inbound
Decoupled Smart Contract Audits: Lightweight LLM Framework via Distillation and Aggregation Atla Selene Mini: A General Purpose Evaluation Model
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a097e6de-8e70-4f5b-b512-32b3998d4134 · inbound
Counsel: A Meta-Evaluation Dataset for Agentic Tasks Atla Selene Mini: A General Purpose Evaluation Model
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 69427740-b102-4ced-8a95-6852b925c18c · inbound
Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models Atla Selene Mini: A General Purpose Evaluation Model
Reference 167
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1072ece7-ec15-493e-843c-e6c2c3ebaf38 · inbound
Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds Atla Selene Mini: A General Purpose Evaluation Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.