Pith. sign in

Paper Citation Record · LEDGER

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta

As of 16 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2501.00257.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00257 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:01:05.175338Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T10:13:22.927828Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:07:56.328267Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy17
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88aeec0e-04dc-43dc-a02e-2aa1fe330a44 · outbound

This paper cites Easy problems that llms get wrong.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Easy problems that llms get wrong

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.147951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:04.991198Z digest=sha256:f546fe9217ba94de0f9533019e194db066b9ad67a668b003f884c46bc429e9c9

Observation 98cc8109-6d03-4cbf-92e9-25e880f9240e · outbound

This paper cites Wider and deeper llm networks are fairer llm evaluators.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Wider and deeper llm networks are fairer llm evaluators

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.135629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:04.995596Z digest=sha256:15a1397b19231b6ba05d01cd36363e6a430c7524328b61b3427d1205dfee9902

Observation c38bcac9-4dd5-4ac1-b8c2-da444060f07b · outbound

This paper cites Likelihood-based mitigation of evaluation bias in large language models,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Likelihood-based mitigation of evaluation bias in large language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:04.999522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:04.999522Z digest=sha256:598d4eeb6120bb0c9ccd983cfa85e346cefb4377aa4682678b318dd1c9fb13de

Observation 4e4ffac1-4a65-4adb-92fa-df9d93044a48 · outbound

This paper cites Eliminating Position Bias of Language Models: A Mechanistic Approach.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Eliminating Position Bias of Language Models: A Mechanistic Approach

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.004248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.004248Z digest=sha256:628f6826015f75db96a9c21a094281c477b432beeaf51dd04862fa192d128aaf

Observation 919b7864-f95b-4f15-8e20-127e29523506 · outbound

This paper cites Is Reference Necessary in the Evaluation of NLG Systems? When and Where?.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Is Reference Necessary in the Evaluation of NLG Systems? When and Where?

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:05.708641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.009442Z digest=sha256:7abd877c21590011ed794eb1f8940d8a22fecdae55ab0af452ffa482a5fab10e

Observation 71448988-0eea-4d2b-af94-4d6b91ca0f03 · outbound

This paper cites Style Over Substance: Evaluation Biases for Large Language Models.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Style Over Substance: Evaluation Biases for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.014038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.014038Z digest=sha256:5cfd7a6e683b2dba37ac7427093134043a773cb42ecbb4e6554ad63febcb1fcc

Observation b1994f4a-8592-4206-8661-7595b712b042 · outbound

This paper cites Developing Safe and Responsible Large Language Model : Can We Balance Bias Reduction and Language Understanding in Large Language Models?.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Developing Safe and Responsible Large Language Model : Can We Balance Bias Reduction and Language Understanding in Large Language Models?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.019358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.019358Z digest=sha256:31491f1f91eb5b673fb433b314ab7354e2adbee994b89b0b6dd11515d5d77fb2

Observation 30aa11da-13e1-4ed8-ab8d-c1c7a56d4f0e · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Can Large Language Models Be an Alternative to Human Evaluations?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.024146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.024146Z digest=sha256:d59d26d97db484a588ccc24704a141b177d2cba8048926d4726add079d67b470

Observation 5181b783-d8b3-4886-b614-6d16f5062f3b · outbound

This paper cites Calibrating Reasoning in Language Models with Internal Consistency.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Calibrating Reasoning in Language Models with Internal Consistency

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.028500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.028500Z digest=sha256:46d09f61fe66671efe3f7f282ff76b370f614d9d167970afbf6d90a6ca80e2e5

Observation dc3a4243-4793-4e0f-adb9-efc3d43ec620 · outbound

This paper cites Investigating automatic scoring and feedback using large language models,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Investigating automatic scoring and feedback using large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.120911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.032816Z digest=sha256:1e9e8e0492f2f4a52375805b776d2748cfb415d467de11dd22ecc77b00181706

Observation ad6b61c0-c0bb-4514-ba2b-36edc90e2bc3 · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.036834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.036834Z digest=sha256:c0983f4ddf92ff1650a51b3695eb461aeffd5b38a1068041bf03dd062b427bc8

Observation d537cba4-f00a-4fdc-8a62-f1585623981b · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.041156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.041156Z digest=sha256:0ec597afb1ec5b0a102be1fa728b2029adf345623adeaa8bf7b5f920827ce5aa

Observation 0091f0e4-15eb-4ed8-8a54-b22ce563682f · outbound

This paper cites Multiple-Choice Questions are Efficient and Robust LLM Evaluators.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Multiple-Choice Questions are Efficient and Robust LLM Evaluators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.046176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.046176Z digest=sha256:83a6421858036a484c31f12b368d5f68c99b9caf73c214b3919692a8d9bf3f85

Observation 95a6f582-434a-4cc3-a5f1-a2e516786e7c · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.050418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.050418Z digest=sha256:010f9754fac7e5254a5925cc023f46e43c75cfcb7bed7f353fca502e95b58441

Observation a6f32267-84db-447c-9f6a-117e6592ba28 · outbound

This paper cites HybridRAG: Integrating Knowledge Graphs and Vector Retrieval Augmented Generation for Efficient Information Extraction.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta HybridRAG: Integrating Knowledge Graphs and Vector Retrieval Augmented Generation for Efficient Information Extraction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.054412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.054412Z digest=sha256:08e206c48a7cf2b049b9b1ea56b3bc49cd241c52dfefb170bf0334d178b23602

Observation 02f17650-ef9a-4586-8f96-c346e3425b57 · outbound

This paper cites Evaluations are critical for understanding the capabilities of large language models (llms),.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Evaluations are critical for understanding the capabilities of large language models (llms),

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.107389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.058834Z digest=sha256:842bb8f0ab4aec4fe958ff202993a9e566bec5a459aa6d354ccb4885c09f1dda

Observation a4bdfffe-f263-4c9a-9483-dfe02acdf823 · outbound

This paper cites Exploring bias and prediction metrics to characterise the fairness of machine learning for equity-centered public health decision-making: A narrative review,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Exploring bias and prediction metrics to characterise the fairness of machine learning for equity-centered public health decision-making: A narrative review,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.095920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.063244Z digest=sha256:2cc989eb4f5ba0d270f69886222da1bff3ec54651d9fab9ae44a5c9303898287

Observation 160e2bd3-fa5b-4f0f-969c-8db522327a6a · outbound

This paper cites Fake news detection: Comparative evaluation of bert-like models and large language models with generative ai-annotated data,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Fake news detection: Comparative evaluation of bert-like models and large language models with generative ai-annotated data,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.084624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.067062Z digest=sha256:098bdbb62f55fe85b207b787bbcf4d31957e5616c35827c87a7b8b21c58df98c

Observation 28f090c3-1cf8-4b2d-b148-b1467cdaf159 · outbound

This paper cites Gender Bias in Contextualized Word Embeddings.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Gender Bias in Contextualized Word Embeddings

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.070741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.070741Z digest=sha256:700cd1cbcba35eb23977fc546de070343cdab189a029e40c3267a18183db268f

Observation f7c52f11-703d-46f4-a9d4-c3cc1ed79500 · outbound

This paper cites Vilbias: A framework for bias detection using linguistic and visual cues,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Vilbias: A framework for bias detection using linguistic and visual cues,

Reference 20

Resolution
verified exact
raw_fallback, observed 2026-08-10T23:01:05.445888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.074678Z digest=sha256:ec92fd270077c59e2b131d7971f542df48f2776169e87f3c170dd20972c21c13

Observation d062cd46-0c09-44ba-9e1f-b33e2882fbcf · outbound

This paper cites Usability maturity model: Human centredness scale,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Usability maturity model: Human centredness scale,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.072497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.079004Z digest=sha256:3766e713466629dd92083928ff99ac353bfa647a2f505fb02ba3b9bc2906241a

Observation 3cf6f5e5-8b38-41e1-97f7-0d1bff4a7367 · outbound

This paper cites Cohen, Statistical Power Analysis for the Behavioral Sciences.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Cohen, Statistical Power Analysis for the Behavioral Sciences

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.058887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.083143Z digest=sha256:b535617d47e842becd2baab77a6ca0536ebeeb2413736e0e0dc74dc66b0589ef

Observation d8523b20-6e6a-42f0-a7d6-ad7a2fe4a815 · outbound

This paper cites Cohen, Statistical Power Analysis for the Behavioral Sciences.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Cohen, Statistical Power Analysis for the Behavioral Sciences

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.046042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.087412Z digest=sha256:4c5b847ce3ebf881705b18c35b3c2d1ed31d3c8d5abde7dd2a0b750c8ce1c5cb

Observation 3326e2ab-d9a2-43eb-b656-12ac69f29f96 · outbound

This paper cites New effect size rules of thumb,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta New effect size rules of thumb,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.024271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.092038Z digest=sha256:34450e238bddfc5fc1d22018f9de655fae9e330f1ecf73ad6d9e101c906955c7

Observation ada8d743-b219-42e6-a95d-f0288407b618 · outbound

This paper cites Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and anovas,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and anovas,

Reference 25

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T23:01:06.012329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.096178Z digest=sha256:0eaa6c13bf1471156dced42b01d16d5c5aa499803c80f59bee88af87f62b29d5

Observation f7eeab6a-e473-45ac-9484-ca0408855362 · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.998912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.101515Z digest=sha256:2043a590efefbb780b55829cafd27032d73514a0c02a17d9d0c997605057a1cb

Observation 97674471-872d-4a43-8fd7-cfc1a24921b5 · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.985984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.105726Z digest=sha256:e4c400d9e919fd8070d4b6ee155319d709360d726f7e6fb4cf2ed2c611b4bbe4

Observation 3e13059f-c86d-4c40-98c1-c5521b4cae4b · outbound

This paper cites Human-Eval.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Human-Eval

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.974031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.110234Z digest=sha256:6b1972ac2f9e29a39a7c18bccc3c7bc01a2a2f2079e8e3e5f6eb6f86dbec4101

Observation 578eca70-1d96-4042-8efe-da6bf14b93dc · outbound

This paper cites • Length of the Bar: The longer the bar, the greater the effect size, meaning the model’s performance deviates more significantly from the EQUATOR evaluator.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta • Length of the Bar: The longer the bar, the greater the effect size, meaning the model’s performance deviates more significantly from the EQUATOR evaluator

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.961202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.114552Z digest=sha256:b17b52558fa7e42b137c21b823423efcb627872978bc1230d5b620df6fa5befb

Observation e0b85857-014a-4e37-a229-9e8124b00c8f · outbound

This paper cites – Higher values mean a larger difference between the mean scores, standardized by the pooled variability.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta – Higher values mean a larger difference between the mean scores, standardized by the pooled variability

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.948640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.118676Z digest=sha256:d7a030a0c90445b68788cdcd09913b89b12906b6771cda1520835ba3be5f1577

Observation f43caacc-ba0d-42e1-9dca-4c2f796bcbfc · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.936795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.122512Z digest=sha256:2575ad6b03087c8a041febd3256be2a9c4c1b9cc35c2a1ea5205eec4b2c2a707

Observation bab3a378-174a-4712-bd00-2973322a4400 · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.925327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.128120Z digest=sha256:3ea10dd506196e0ddcc8c4939f6f19d42822675560c7b0c95a0114982783cf30

Observation 47203ee1-7968-4860-910c-561dd076de8a · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.914179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.132175Z digest=sha256:837cb7952807080987958617c9481aa5efc654ec95881dfe6b1e0a3021d6caa9

Observation eac7ffff-97f9-4208-ab5e-c19e3cafc5be · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.901853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.136389Z digest=sha256:7823fafd7e38909e0e641a06322a05903c45790dc3ce665a402192a24222d699

Observation a9548bf7-494d-44ed-a4e9-d6151f2ecb7e · outbound

This paper cites Larger effects require fewer questions.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Larger effects require fewer questions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.889126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.143033Z digest=sha256:25f80b618acb102ae4e62d56da21ea4a7beb463e040735ada7cbd50098d6d81c

Observation f86df44c-5829-41a8-81d6-0b05a06f1a07 · outbound

This paper cites • These models are relatively aligned with the EQUATOR evaluator, suggesting they produce similar results.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta • These models are relatively aligned with the EQUATOR evaluator, suggesting they produce similar results

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.876754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.149377Z digest=sha256:b45a28ecf582ca61c4ac8dc3eba50be0618695792cb5446c99930d3d93ac4469

Observation e7d2daf7-0ddb-4958-b39e-b523e8786954 · outbound

This paper cites • These models perform significantly worse under the given evaluation, as indicated by large positive Cohen’s d values.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta • These models perform significantly worse under the given evaluation, as indicated by large positive Cohen’s d values

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.863436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.154981Z digest=sha256:6298405063fd35cd32095b21d8987ee6a6df4d7c2b3fd9028e8a35d66d0a31ff

Observation 0785b479-33f6-4e8d-b79f-7a8c26d733ff · outbound

This paper cites • A wide range of Cohen’s d values (e.g., from˜0.3 to >2.0) indicates substantial variability in model performance.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta • A wide range of Cohen’s d values (e.g., from˜0.3 to >2.0) indicates substantial variability in model performance

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.850675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.160027Z digest=sha256:88d367f7032097a844fba51163c29be8be718eeac886c69e91a245de80eecdd5

Observation 7b4aa07f-5509-4225-8e08-401831fa7131 · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.836770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.165784Z digest=sha256:29970884c7e3d63d63936e77f596c9c43109575c2e2c06426b77374dd7b1cfd1

Observation acd28fd1-4a70-420b-9280-b816ee0352ff · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.819994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.170926Z digest=sha256:e48208df4c698dd24b652b35cf0ca19a5563f9da11dfe8168cca6eecced1049a

Observation d9836045-3ec6-4a4c-8783-ccf30441411d · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.807268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-10T23:01:05.175338Z digest=sha256:14f1c64e86807b8b94516a4e45322a84549968e63a647ade5cc9d109d1ee8214

Pith citing papers

Observation 42c8eb86-374a-40d3-8b79-4964b2a39105 · inbound

Hybrid LLM-based Intelligent Framework for Robot Task Scheduling cites this paper.

Hybrid LLM-based Intelligent Framework for Robot Task Scheduling EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:22:39.541189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T16:21:19.722476Z digest=sha256:c403a9ab9c8b466102f5fff1e2ef3568598670b65afe6b2da60c4db11997a7cb

Observation 0478b8ac-07c5-4e89-9319-ea126317238f · inbound

Improving Cross-Format Robustness in Language Models with Multi-Format Training cites this paper.

Improving Cross-Format Robustness in Language Models with Multi-Format Training EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:07:56.329771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T10:13:22.927828Z digest=sha256:525635de5a640b50a8db50d5fb9674faec317c3e5e12d6a248fecb96d7fc0608