Pith. sign in

Paper Citation Record · LEDGER

Atla Selene Mini: A General Purpose Evaluation Model

As of 17 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 12 inbound Pith citation observations for arXiv:2501.17195.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.17195 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T13:47:45.942549Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:50.298845Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:49:38.206304Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5955838e-8963-437a-8e44-d5305b984b90 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Atla Selene Mini: A General Purpose Evaluation Model Constitutional AI: Harmlessness from AI Feedback

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.742177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.742177Z digest=sha256:8c30b6c142875fc0b189bd2d1af2d58bb3e0456061a3b17c2591aa4998da324d

Observation 924147ce-03bf-488d-8607-451c7c70fadb · outbound

This paper cites Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision.

Atla Selene Mini: A General Purpose Evaluation Model Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.748385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.748385Z digest=sha256:02e2f163eec39cd77c0dc98112ac911a894cdd2e132e11ffd1fcd662744ca58b

Observation 15e59109-531f-47ce-bf4b-6e89b8147eda · outbound

This paper cites From generation to judgment: Opportunities and challenges of llm-as-a-judge.

Atla Selene Mini: A General Purpose Evaluation Model From generation to judgment: Opportunities and challenges of llm-as-a-judge

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.755134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.755134Z digest=sha256:a88664f22e7814b9d23216598ee7e668f134dbb9eff4235d70a8c2a565fc0393

Observation d519f7cc-c2a7-402f-b96e-64bc54789c1f · outbound

This paper cites Offsetbias: Leveraging debiased data for tuning evaluators, 2024.

Atla Selene Mini: A General Purpose Evaluation Model Offsetbias: Leveraging debiased data for tuning evaluators, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.821317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.761709Z digest=sha256:69e1ceff591d9ec55b19e1bbd5c66c0d7cebfb98be23a0216b080bd80495ab81

Observation 6c88dab8-23fa-4195-bb04-d7bf001ddc41 · outbound

This paper cites Self-preference bias in llm-as-a-judge, 2024.

Atla Selene Mini: A General Purpose Evaluation Model Self-preference bias in llm-as-a-judge, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.806539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.768791Z digest=sha256:533827c82f92f1ec1827c2ccd9896f5178836f49441c48d2db953472e98e7ed7

Observation cac79179-7a51-442e-b6f6-01a4143307a7 · outbound

This paper cites Judging the judges: Evaluating alignment and vulnerabilities in llms-as- judges, 2024.

Atla Selene Mini: A General Purpose Evaluation Model Judging the judges: Evaluating alignment and vulnerabilities in llms-as- judges, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.792124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.774292Z digest=sha256:c31b4fd7013b3a5c25bb43201f286c2363635866f49c94af823b0eda34400891

Observation a5d8d982-0432-454c-835c-fff7d9de49e3 · outbound

This paper cites Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation.

Atla Selene Mini: A General Purpose Evaluation Model Foundational Autoraters: Taming Large Language Models for Better Automatic Evaluation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.780674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.780674Z digest=sha256:4f8da0a8266b9ceb1bfefe9bc2f45b09e7d69b24ff2f24cccc2f35d8adf0fbbb

Observation a2390a24-5fb5-4bac-a135-ebf499258bde · outbound

This paper cites Flow judge: An open small language model for llm system evaluations.

Atla Selene Mini: A General Purpose Evaluation Model Flow judge: An open small language model for llm system evaluations

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.774775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.786833Z digest=sha256:ee9d4f9a07f378c5e5fd319feb55a73bdba7c424c21376d060105855f7477720

Observation 2869b014-b550-48a7-bd40-2d92ac4ddaf2 · outbound

This paper cites GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking.

Atla Selene Mini: A General Purpose Evaluation Model GLIDER: Grading LLM Interactions and Decisions using Explainable Ranking

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.791540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.791540Z digest=sha256:199e34f8b0eed61c07c5ae1e4d52e3158fe2a5d873eaed02e2e38160d7c7d43a

Observation d408f67e-3710-4adc-ac61-8548b048e9de · outbound

This paper cites Direct Judgement Preference Optimization.

Atla Selene Mini: A General Purpose Evaluation Model Direct Judgement Preference Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.796976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.796976Z digest=sha256:004154dab1c913a2d7daeb72513fb6a9d1de517b13512991bba496894e59dbe7

Observation 79e9433a-0aff-4a2c-b88b-8e9db9716d59 · outbound

This paper cites Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models.

Atla Selene Mini: A General Purpose Evaluation Model Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.802686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.802686Z digest=sha256:4daa3dadf02408e64c76f5368bbb8616ec5da96e68787c1db4de67e4fc51021b

Observation 12270fef-1e97-4658-95ee-b5862f1abbf8 · outbound

This paper cites Judge arena: Benchmarking llms as evaluators.

Atla Selene Mini: A General Purpose Evaluation Model Judge arena: Benchmarking llms as evaluators

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.756546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.809130Z digest=sha256:e89f618831f1ae1cd832afc2852af759ff6e0c73ec457752ea16cad9d02182cc

Observation 09b82bed-2fbb-47d1-b17f-9d9bc1733cac · outbound

This paper cites Iterative Reasoning Preference Optimization.

Atla Selene Mini: A General Purpose Evaluation Model Iterative Reasoning Preference Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.815883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.815883Z digest=sha256:b0e4550a9088f5cd900f7e9372fc51f9f1af1663be5dd9a0dedf1cc08229a06b

Observation b1bdb19a-0c51-4eb1-9ac9-686627d8924a · outbound

This paper cites Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts.

Atla Selene Mini: A General Purpose Evaluation Model Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.821883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.821883Z digest=sha256:dfbe9355eb696d550508fa74ba4249814d9ccd4335a499a376ceecb4352e7e5a

Observation e00fff85-a96b-423f-9fbd-6aaf8643d522 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Atla Selene Mini: A General Purpose Evaluation Model Xing, Hao Zhang, Joseph E

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.830201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.830201Z digest=sha256:5b345d1d86bd2c10bff6d5f68e395e9c77ccf476ff5ca04e98bbcd6b4c5a8de7

Observation b353b3fb-05f6-4310-b62c-8513b57a9903 · outbound

This paper cites Flask: Fine-grained language model evaluation based on alignment skill sets, 2024.

Atla Selene Mini: A General Purpose Evaluation Model Flask: Fine-grained language model evaluation based on alignment skill sets, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.723017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.835083Z digest=sha256:794ac23d47e4758c3bd3efefc3750e47761bc43f3c019807e1dec30994d3dc1b

Observation c800eb55-f596-45a7-a26b-c4b0e6da3766 · outbound

This paper cites The biggen bench: A principled benchmark for fine-grained evaluation of language models with language models, 2024.

Atla Selene Mini: A General Purpose Evaluation Model The biggen bench: A principled benchmark for fine-grained evaluation of language models with language models, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.705680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.840992Z digest=sha256:89141820d83b46c38efb7a1c7c54ff81dce9fce49043408ccbe3b95348c7fbe9

Observation 70d8b0b6-2769-4a06-8e26-40ac8621153f · outbound

This paper cites Smith, and Hannaneh Hajishirzi.

Atla Selene Mini: A General Purpose Evaluation Model Smith, and Hannaneh Hajishirzi

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.846708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.846708Z digest=sha256:363d4e449a1aab5dc46f1e9d8f32cd4d1f5107e4f6449bd44602e7670a7267e5

Observation 50873398-0189-4be5-8d51-ad421f06022e · outbound

This paper cites A critical evaluation of evaluations for long-form question answering, 2023.

Atla Selene Mini: A General Purpose Evaluation Model A critical evaluation of evaluations for long-form question answering, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.673556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.852047Z digest=sha256:5e8d374c01721138e1847cd6bdb991003d9345c66f8a7ab7feb3ed1b45e66de3

Observation 4bb91bf5-4d9a-4df5-ad71-464c7a70840f · outbound

This paper cites A general language assistant as a laboratory for alignment, 2021.

Atla Selene Mini: A General Purpose Evaluation Model A general language assistant as a laboratory for alignment, 2021

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.857961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.857961Z digest=sha256:405ae13b875a7e539dc0ec983bf741e61ef012ed40267a69d2a3ff7dc15d42f4

Observation abc8f2af-4af7-4261-8d7b-c02e1e611dad · outbound

This paper cites Fabbri, Jiawen Chen, Yilun Zhao, Simeng Han, Shafiq Joty, Pengfei Liu, Dragomir Radev, Chien-Sheng Wu, and Arman Cohan.

Atla Selene Mini: A General Purpose Evaluation Model Fabbri, Jiawen Chen, Yilun Zhao, Simeng Han, Shafiq Joty, Pengfei Liu, Dragomir Radev, Chien-Sheng Wu, and Arman Cohan

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.645533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.866016Z digest=sha256:92f0dd920c9e7097a6ab50dd60a44dda6ba3c2688a7cf517cc5c3863a211a078

Observation 6b21a0c0-0721-4d53-a68e-04f0535a4f07 · outbound

This paper cites Generative judge for evaluating alignment, 2023.

Atla Selene Mini: A General Purpose Evaluation Model Generative judge for evaluating alignment, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.629644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.870979Z digest=sha256:60845203630c84b21e6560a928db8b9bb099c1bb3f8553798d7b1cd72d019030

Observation 146b0000-1fe2-4c00-92bd-13c45b35aba3 · outbound

This paper cites InFoBench: Evaluating Instruction Following Ability in Large Language Models.

Atla Selene Mini: A General Purpose Evaluation Model InFoBench: Evaluating Instruction Following Ability in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.877232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.877232Z digest=sha256:1c582bd85f84922ab77567fb333cf46b4147645f1401b495c378a990fa3363f2

Observation 99354afd-5543-4c5d-8395-9c962c2ebbb4 · outbound

This paper cites Minicheck: Efficient fact-checking of llms on grounding documents, 2024.

Atla Selene Mini: A General Purpose Evaluation Model Minicheck: Efficient fact-checking of llms on grounding documents, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.613654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.883625Z digest=sha256:ae8e32bfaaec53a2050c6af65df557c78c316568c169cb1f7c457d1da4710346

Observation 871b8494-0aa9-4bb1-9b05-6f1697567fb6 · outbound

This paper cites Craft-md: A conversational evaluation framework for comprehensive assessment of clinical llms.

Atla Selene Mini: A General Purpose Evaluation Model Craft-md: A conversational evaluation framework for comprehensive assessment of clinical llms

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.584750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.888820Z digest=sha256:fa9fcc9f549763672110c01651c400b83c225417205a7212b5c9fd5f357c2bf7

Observation d860d53e-dff1-4508-a5d6-5993a8ef79a6 · outbound

This paper cites FinanceBench: A New Benchmark for Financial Question Answering.

Atla Selene Mini: A General Purpose Evaluation Model FinanceBench: A New Benchmark for Financial Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.898085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.898085Z digest=sha256:ddd4cf697f2aa02d49dbc53a9d7096615e6cf01196c88fd7cd3b9b33d0dfc994

Observation a849a26e-062b-4851-8560-102b7525e6fc · outbound

This paper cites Systematic evaluation of llm-as-a-judge in llm alignment tasks: Explainable metrics and diverse prompt templates, 2024.

Atla Selene Mini: A General Purpose Evaluation Model Systematic evaluation of llm-as-a-judge in llm alignment tasks: Explainable metrics and diverse prompt templates, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.484769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.904118Z digest=sha256:03cb96e6d1f8e05bed54eba2895fb36dedf991be569f88d10ea78a7ac55a0a71

Observation b807122d-c6e6-4330-8f30-2add5ee06b06 · outbound

This paper cites Does prompt formatting have any impact on llm performance?, 2024.

Atla Selene Mini: A General Purpose Evaluation Model Does prompt formatting have any impact on llm performance?, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.429889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.909136Z digest=sha256:3f7d9175d50a40ed3b8be25af9ac4ca8ad00f40b81a17274b70919a47e220138

Observation 4b79dd25-cc73-46e2-a05c-1847f9c7df36 · outbound

This paper cites The comparative trap: Pairwise comparisons amplifies biased preferences of llm evaluators, 2024.

Atla Selene Mini: A General Purpose Evaluation Model The comparative trap: Pairwise comparisons amplifies biased preferences of llm evaluators, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.403937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.914857Z digest=sha256:14ef749b5644526f3cd6409f6b6082325142267533434ec27bf3533ade44cddf

Observation dd667c32-6ed3-4b0b-aa00-338f0240b428 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Atla Selene Mini: A General Purpose Evaluation Model Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.920196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.920196Z digest=sha256:fdff93ca919223dbadc4628557375226da11eab55d476afd1fd72ff23d7d52f2

Observation 46e1c7a0-247e-4f59-a187-02305a63d8f6 · outbound

This paper cites OpenAI o1 System Card.

Atla Selene Mini: A General Purpose Evaluation Model OpenAI o1 System Card

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T13:47:45.926401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:47:45.926401Z digest=sha256:45c4a5f33c603c36b6396539bef29654ab8645ec9aa26170b3472d722a916760

Observation fa086e8d-5142-4f1e-aef1-7db168a90d00 · outbound

This paper cites an unresolved cited work.

Atla Selene Mini: A General Purpose Evaluation Model Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T13:47:46.386336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.932717Z digest=sha256:c18f5f40c1e1dc775b4eb314eed9e17c0012f3ef98ccf7a19afca5dcd190a591

Observation 05bc346a-6d67-43c4-82d9-a00004dc7779 · outbound

This paper cites Nomic atlas.

Atla Selene Mini: A General Purpose Evaluation Model Nomic atlas

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.369292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.937602Z digest=sha256:738aee3d968399c09ab178ec906ab201a01492031e9771e2b338d68ae45f49d6

Observation 198cdc71-4ff9-4f27-a514-e71f8c8d65c3 · outbound

This paper cites "Dear Readers, <omitted for conciseness> P.S. No garden gnomes were harmed in the writing of this book.

Atla Selene Mini: A General Purpose Evaluation Model "Dear Readers, <omitted for conciseness> P.S. No garden gnomes were harmed in the writing of this book

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T13:47:46.348249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-10T13:47:45.942549Z digest=sha256:c4beba46f915ee6dca6ab14790c364455ced284245b091c069a6bff57ad94b62

Pith citing papers

Observation 10d6e181-a23c-4326-a1ea-b681e34819c8 · inbound

Reward Reasoning Model cites this paper.

Reward Reasoning Model Atla Selene Mini: A General Purpose Evaluation Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:50.298845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:50.298845Z digest=sha256:ca135a593fe888bc02ae5c553ef6e469b85d6687b2007e37434fd55dc6b612cc

Observation 473fb600-3bd8-4f37-939b-d734da71ea0f · inbound

ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments cites this paper.

ReliableEval: A Recipe for Stochastic LLM Evaluation via Method of Moments Atla Selene Mini: A General Purpose Evaluation Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.293260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:20:16.293260Z digest=sha256:8f67be0a24549bab7d531dc45543d8e03af34cfa1184824fdf8babb0e9b7fa52

Observation b21ac2cc-0103-47d7-a5f5-69cad4632d49 · inbound

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling cites this paper.

PaTaRM: Bridging Pairwise and Pointwise Signals via Preference-Aware Task-Adaptive Reward Modeling Atla Selene Mini: A General Purpose Evaluation Model

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:20:49.325679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T03:15:57.744706Z digest=sha256:1951e00787678a59a178acd1cd0835f19dfc55f2a5eb9250345cc276490ad210

Observation 186743ae-0d1a-452e-8e87-9a8674819b32 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Atla Selene Mini: A General Purpose Evaluation Model

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.002417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.002417Z digest=sha256:023b827b69fd6364a7e28784b155c8f5e23f48b95e239452d313316bf714854b

Observation 633b275c-4d7e-406b-9fe7-66f5b8e3a17a · inbound

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization cites this paper.

Toward Robust LLM-Based Judges: Taxonomic Bias Evaluation and Debiasing Optimization Atla Selene Mini: A General Purpose Evaluation Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T05:56:17.811091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:56:17.811091Z digest=sha256:41209c6750ca29aab22d168a8911515731ed9388c3e2f83468544ab26aca801c

Observation 8646f0f6-5d56-4f6d-8d75-6de409b7d501 · inbound

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems cites this paper.

CompliBench: Benchmarking LLM Judges for Compliance Violation Detection in Dialogue Systems Atla Selene Mini: A General Purpose Evaluation Model

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:26:02.445512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T14:56:00.449776Z digest=sha256:1ea657a482b5a10d7e93ac4a46bcaca63f9f5581555b701b7ff2574ab69e4efa

Observation 2c34bab5-fb6a-464f-8d10-984958b188a9 · inbound

VERDI: Single-Call Confidence Estimation for Verification-Based LLM Judges via Decomposed Inference cites this paper.

VERDI: Single-Call Confidence Estimation for Verification-Based LLM Judges via Decomposed Inference Atla Selene Mini: A General Purpose Evaluation Model

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:52:05.633574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-13T01:48:29.633350Z digest=sha256:df27705cc18bed1e69b60015670c5fb41b689d9c339502b00189d4cf39a79a36

Observation 2d4c3f1a-13ab-4a46-ba7c-0ad48e2eb519 · inbound

A Finite-Calibration Regime Map for LLM Judge Panels cites this paper.

A Finite-Calibration Regime Map for LLM Judge Panels Atla Selene Mini: A General Purpose Evaluation Model

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:06:13.157300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T17:34:28.258222Z digest=sha256:b61a9a457d6637f0871a79ec961bba742bd1e5fe3e2a324c610d597908936b48

Observation b61684d5-7b12-4020-adf8-8846be308705 · inbound

Decoupled Smart Contract Audits: Lightweight LLM Framework via Distillation and Aggregation cites this paper.

Decoupled Smart Contract Audits: Lightweight LLM Framework via Distillation and Aggregation Atla Selene Mini: A General Purpose Evaluation Model

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:29.997093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T09:58:30.611774Z digest=sha256:389e83ea345dd7082314eb15fd22315c153b00ed2bc9c7bff6937c1aa39bc446

Observation a097e6de-8e70-4f5b-b512-32b3998d4134 · inbound

Counsel: A Meta-Evaluation Dataset for Agentic Tasks cites this paper.

Counsel: A Meta-Evaluation Dataset for Agentic Tasks Atla Selene Mini: A General Purpose Evaluation Model

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:49:38.208141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T14:07:59.446478Z digest=sha256:519d83da063184dc9487aac41c73939734160661791804a2c752f4af613e66fe

Observation 69427740-b102-4ced-8a95-6852b925c18c · inbound

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models cites this paper.

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models Atla Selene Mini: A General Purpose Evaluation Model

Reference 167

Resolution
unresolved
no resolver link, observed 2026-08-01T19:02:46.756858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:02:46.756858Z digest=sha256:67c5740f581211c065019d55cd1e3ba6f797fdb631e83573913b40e8221f691e

Observation 1072ece7-ec15-493e-843c-e6c2c3ebaf38 · inbound

Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds cites this paper.

Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds Atla Selene Mini: A General Purpose Evaluation Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T17:29:12.599582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T17:29:12.599582Z digest=sha256:e78f170cf6b4b722a6ca336a0f51a340de6fc9596f51dacf4303983d079cf849