Pith. sign in

Paper Citation Record · LEDGER

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta

As of 16 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 2 inbound Pith citation observations for arXiv:2501.00257.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00257 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:01:05.175338Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T10:13:22.927828Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T10:07:56.328267Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact2
  • verified fuzzy17
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88aeec0e-04dc-43dc-a02e-2aa1fe330a44 · outbound

This paper cites Easy problems that llms get wrong.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Easy problems that llms get wrong

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.147951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:04.991198Z digest=sha256:aa77e0003ec607f3fc5810eecefcc00b165a6d1d00dec2280c4f03ccf877adbb

Observation 98cc8109-6d03-4cbf-92e9-25e880f9240e · outbound

This paper cites Wider and deeper llm networks are fairer llm evaluators.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Wider and deeper llm networks are fairer llm evaluators

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.135629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:04.995596Z digest=sha256:698c54e90ec63acd4b48d450c527bcb07625d092ecc21533f20e3bbb6d16f2d2

Observation c38bcac9-4dd5-4ac1-b8c2-da444060f07b · outbound

This paper cites Likelihood-based mitigation of evaluation bias in large language models,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Likelihood-based mitigation of evaluation bias in large language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:04.999522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:04.999522Z digest=sha256:598d4eeb6120bb0c9ccd983cfa85e346cefb4377aa4682678b318dd1c9fb13de

Observation 4e4ffac1-4a65-4adb-92fa-df9d93044a48 · outbound

This paper cites Eliminating Position Bias of Language Models: A Mechanistic Approach.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Eliminating Position Bias of Language Models: A Mechanistic Approach

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.004248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.004248Z digest=sha256:628f6826015f75db96a9c21a094281c477b432beeaf51dd04862fa192d128aaf

Observation 919b7864-f95b-4f15-8e20-127e29523506 · outbound

This paper cites Is Reference Necessary in the Evaluation of NLG Systems? When and Where?.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Is Reference Necessary in the Evaluation of NLG Systems? When and Where?

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:01:05.708641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.009442Z digest=sha256:08a7e8c8df7976cbe58b27823f73fb24ee7381fed1b5b2484113c3187565c889

Observation 71448988-0eea-4d2b-af94-4d6b91ca0f03 · outbound

This paper cites Style Over Substance: Evaluation Biases for Large Language Models.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Style Over Substance: Evaluation Biases for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.014038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.014038Z digest=sha256:5cfd7a6e683b2dba37ac7427093134043a773cb42ecbb4e6554ad63febcb1fcc

Observation b1994f4a-8592-4206-8661-7595b712b042 · outbound

This paper cites Developing Safe and Responsible Large Language Model : Can We Balance Bias Reduction and Language Understanding in Large Language Models?.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Developing Safe and Responsible Large Language Model : Can We Balance Bias Reduction and Language Understanding in Large Language Models?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.019358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.019358Z digest=sha256:31491f1f91eb5b673fb433b314ab7354e2adbee994b89b0b6dd11515d5d77fb2

Observation 30aa11da-13e1-4ed8-ab8d-c1c7a56d4f0e · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Can Large Language Models Be an Alternative to Human Evaluations?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.024146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.024146Z digest=sha256:d59d26d97db484a588ccc24704a141b177d2cba8048926d4726add079d67b470

Observation 5181b783-d8b3-4886-b614-6d16f5062f3b · outbound

This paper cites Calibrating Reasoning in Language Models with Internal Consistency.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Calibrating Reasoning in Language Models with Internal Consistency

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.028500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.028500Z digest=sha256:46d09f61fe66671efe3f7f282ff76b370f614d9d167970afbf6d90a6ca80e2e5

Observation dc3a4243-4793-4e0f-adb9-efc3d43ec620 · outbound

This paper cites Investigating automatic scoring and feedback using large language models,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Investigating automatic scoring and feedback using large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.120911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.032816Z digest=sha256:b51d85c36cda835735521b1c3050ab7b9f2f123e403db59a4444412da874eb15

Observation ad6b61c0-c0bb-4514-ba2b-36edc90e2bc3 · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.036834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.036834Z digest=sha256:c0983f4ddf92ff1650a51b3695eb461aeffd5b38a1068041bf03dd062b427bc8

Observation d537cba4-f00a-4fdc-8a62-f1585623981b · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.041156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.041156Z digest=sha256:0ec597afb1ec5b0a102be1fa728b2029adf345623adeaa8bf7b5f920827ce5aa

Observation 0091f0e4-15eb-4ed8-8a54-b22ce563682f · outbound

This paper cites Multiple-Choice Questions are Efficient and Robust LLM Evaluators.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Multiple-Choice Questions are Efficient and Robust LLM Evaluators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.046176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.046176Z digest=sha256:83a6421858036a484c31f12b368d5f68c99b9caf73c214b3919692a8d9bf3f85

Observation 95a6f582-434a-4cc3-a5f1-a2e516786e7c · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.050418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.050418Z digest=sha256:010f9754fac7e5254a5925cc023f46e43c75cfcb7bed7f353fca502e95b58441

Observation a6f32267-84db-447c-9f6a-117e6592ba28 · outbound

This paper cites HybridRAG: Integrating Knowledge Graphs and Vector Retrieval Augmented Generation for Efficient Information Extraction.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta HybridRAG: Integrating Knowledge Graphs and Vector Retrieval Augmented Generation for Efficient Information Extraction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.054412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.054412Z digest=sha256:08e206c48a7cf2b049b9b1ea56b3bc49cd241c52dfefb170bf0334d178b23602

Observation 02f17650-ef9a-4586-8f96-c346e3425b57 · outbound

This paper cites Evaluations are critical for understanding the capabilities of large language models (llms),.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Evaluations are critical for understanding the capabilities of large language models (llms),

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.107389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.058834Z digest=sha256:c46c196930c529342a372d41ddfd312184fc866bc350e96f2e80aa63bd04f738

Observation a4bdfffe-f263-4c9a-9483-dfe02acdf823 · outbound

This paper cites Exploring bias and prediction metrics to characterise the fairness of machine learning for equity-centered public health decision-making: A narrative review,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Exploring bias and prediction metrics to characterise the fairness of machine learning for equity-centered public health decision-making: A narrative review,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.095920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.063244Z digest=sha256:5ee74f34701355dfc7aa17aa8d9e8b760a89395f1ac5642d4142554193aa4881

Observation 160e2bd3-fa5b-4f0f-969c-8db522327a6a · outbound

This paper cites Fake news detection: Comparative evaluation of bert-like models and large language models with generative ai-annotated data,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Fake news detection: Comparative evaluation of bert-like models and large language models with generative ai-annotated data,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.084624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.067062Z digest=sha256:39603ff97a49f5062b359da49183bbd857fd53bbb74986f775c6635cbbb39a1c

Observation 28f090c3-1cf8-4b2d-b148-b1467cdaf159 · outbound

This paper cites Gender Bias in Contextualized Word Embeddings.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Gender Bias in Contextualized Word Embeddings

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:01:05.070741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:01:05.070741Z digest=sha256:700cd1cbcba35eb23977fc546de070343cdab189a029e40c3267a18183db268f

Observation f7c52f11-703d-46f4-a9d4-c3cc1ed79500 · outbound

This paper cites Vilbias: A framework for bias detection using linguistic and visual cues,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Vilbias: A framework for bias detection using linguistic and visual cues,

Reference 20

Resolution
verified exact
raw_fallback, observed 2026-08-10T23:01:05.445888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.074678Z digest=sha256:9ac51258ddb0e1f8b44a23aed5572e776312de9b1dfeeaa0415328ea18482f19

Observation d062cd46-0c09-44ba-9e1f-b33e2882fbcf · outbound

This paper cites Usability maturity model: Human centredness scale,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Usability maturity model: Human centredness scale,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.072497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.079004Z digest=sha256:464ae5d4f738700fcf6baacecadc63c309e226a591b4cd7b9e92298645693610

Observation 3cf6f5e5-8b38-41e1-97f7-0d1bff4a7367 · outbound

This paper cites Cohen, Statistical Power Analysis for the Behavioral Sciences.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Cohen, Statistical Power Analysis for the Behavioral Sciences

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.058887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.083143Z digest=sha256:bd4da69c38d5ec2e15298c3dfa52c0a1dd7e52b40bdec67be43af31cdca9e297

Observation d8523b20-6e6a-42f0-a7d6-ad7a2fe4a815 · outbound

This paper cites Cohen, Statistical Power Analysis for the Behavioral Sciences.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Cohen, Statistical Power Analysis for the Behavioral Sciences

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.046042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.087412Z digest=sha256:44aca056ab61e5d538ec01920ddd5403ddc86f6e064c62c03ff3c9de07f24f8a

Observation 3326e2ab-d9a2-43eb-b656-12ac69f29f96 · outbound

This paper cites New effect size rules of thumb,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta New effect size rules of thumb,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:06.024271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.092038Z digest=sha256:910e29c6432ec11ee7dc8fc4ab64f3f67c847c947bda89927099c20050d0c228

Observation ada8d743-b219-42e6-a95d-f0288407b618 · outbound

This paper cites Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and anovas,.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Calculating and reporting effect sizes to facilitate cumulative science: A practical primer for t-tests and anovas,

Reference 25

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T23:01:06.012329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.096178Z digest=sha256:3968b6bc422a96e5c54a1d99e40d64d68225fd56cde5bb35707f28bd0bb144e7

Observation f7eeab6a-e473-45ac-9484-ca0408855362 · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.998912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.101515Z digest=sha256:6fb95cbefee9655475bb9cb9b312825b343af955d2a3f2168f48702d05f7c7f3

Observation 97674471-872d-4a43-8fd7-cfc1a24921b5 · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.985984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.105726Z digest=sha256:5683594641c145fb7aa91ff61ae3616106eaf55ba62b49f54d11695346684ebc

Observation 3e13059f-c86d-4c40-98c1-c5521b4cae4b · outbound

This paper cites Human-Eval.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Human-Eval

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.974031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.110234Z digest=sha256:b87c82c86109827b0c6ae749419504fdd4d2dda9ad9dd8b812dfd9ed11c0e3ca

Observation 578eca70-1d96-4042-8efe-da6bf14b93dc · outbound

This paper cites • Length of the Bar: The longer the bar, the greater the effect size, meaning the model’s performance deviates more significantly from the EQUATOR evaluator.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta • Length of the Bar: The longer the bar, the greater the effect size, meaning the model’s performance deviates more significantly from the EQUATOR evaluator

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.961202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.114552Z digest=sha256:be07bd5fde002e193807ea013efdc7d4c9bde977432a11d23a6bf6f93d6051a8

Observation e0b85857-014a-4e37-a229-9e8124b00c8f · outbound

This paper cites – Higher values mean a larger difference between the mean scores, standardized by the pooled variability.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta – Higher values mean a larger difference between the mean scores, standardized by the pooled variability

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.948640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.118676Z digest=sha256:638bf4b9aa7c2aef6ea2b06843f65ef27a7dba0dbe39e03cb9c725e303abedbf

Observation f43caacc-ba0d-42e1-9dca-4c2f796bcbfc · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.936795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.122512Z digest=sha256:3d8e0b85e3560fbdb559b788c31469bfc0f3dc1985f429db2f40e30233803b4f

Observation bab3a378-174a-4712-bd00-2973322a4400 · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.925327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.128120Z digest=sha256:bbd05d2d08c2be9b76c45acfb51777cefcecf80fd6f79ddc94ee9cc7c46407b7

Observation 47203ee1-7968-4860-910c-561dd076de8a · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.914179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.132175Z digest=sha256:96fabb9b37873cf12aba7a190229f852a4f8e5943f5825d5caf80004497c6a06

Observation eac7ffff-97f9-4208-ab5e-c19e3cafc5be · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.901853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.136389Z digest=sha256:515d01289890884a93ffcde8c4806a9f6d5c141d60020704500558398e34195a

Observation a9548bf7-494d-44ed-a4e9-d6151f2ecb7e · outbound

This paper cites Larger effects require fewer questions.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Larger effects require fewer questions

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.889126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.143033Z digest=sha256:9ba077f2383561edfc8fd8cb687b865d0145ed7d1f31eb2c2ebb3a4b6ff09fd0

Observation f86df44c-5829-41a8-81d6-0b05a06f1a07 · outbound

This paper cites • These models are relatively aligned with the EQUATOR evaluator, suggesting they produce similar results.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta • These models are relatively aligned with the EQUATOR evaluator, suggesting they produce similar results

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.876754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.149377Z digest=sha256:35443103891e1ca8114fe2388fe8fbc8ed42a9e5d4c560352f0b204c65dc7ceb

Observation e7d2daf7-0ddb-4958-b39e-b523e8786954 · outbound

This paper cites • These models perform significantly worse under the given evaluation, as indicated by large positive Cohen’s d values.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta • These models perform significantly worse under the given evaluation, as indicated by large positive Cohen’s d values

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.863436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.154981Z digest=sha256:e497e7d9f5c57b4631e179bd36d26852c7bda5d0bd15752310cc3025d3342d75

Observation 0785b479-33f6-4e8d-b79f-7a8c26d733ff · outbound

This paper cites • A wide range of Cohen’s d values (e.g., from˜0.3 to >2.0) indicates substantial variability in model performance.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta • A wide range of Cohen’s d values (e.g., from˜0.3 to >2.0) indicates substantial variability in model performance

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:01:05.850675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.160027Z digest=sha256:843bb8973c93b3b01686b7a2b2378de20f59c901ebcd3fea4da154eacdecca50

Observation 7b4aa07f-5509-4225-8e08-401831fa7131 · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.836770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.165784Z digest=sha256:e9069e3e1daf96ea4bd1cff2a89464e16db8ebb5e9049bb48e805f4bd1f09a36

Observation acd28fd1-4a70-420b-9280-b816ee0352ff · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.819994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.170926Z digest=sha256:89be78200e657484b0654f0657f335a40b1f2f16f9588a7dc1125627a46602f4

Observation d9836045-3ec6-4a4c-8783-ccf30441411d · outbound

This paper cites an unresolved cited work.

EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:01:05.807268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T23:01:05.175338Z digest=sha256:670eaf3b8e44970de1785c87f8c13e2578b6abac2fd3f333050561399515cc5e

Pith citing papers

Observation 42c8eb86-374a-40d3-8b79-4964b2a39105 · inbound

Hybrid LLM-based Intelligent Framework for Robot Task Scheduling cites this paper.

Hybrid LLM-based Intelligent Framework for Robot Task Scheduling EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:22:39.541189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T16:21:19.722476Z digest=sha256:5d846a2ab462a57a3f26903768dc6d93d68aa2a605a046f4880f993c04f71a2b

Observation 0478b8ac-07c5-4e89-9319-ea126317238f · inbound

Improving Cross-Format Robustness in Language Models with Multi-Format Training cites this paper.

Improving Cross-Format Robustness in Language Models with Multi-Format Training EQUATOR: A Deterministic Framework for Evaluating LLM Reasoning with Open-Ended Questions. # v1.0.0-beta

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:07:56.329771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T10:13:22.927828Z digest=sha256:9fc718f567700780a15e6ce9a12737a25a839ee00e1d4c25b23dbcce9c076cd4