Pith. sign in

Paper Citation Record · LEDGER

Answer Matching Outperforms Multiple Choice for Language Model Evaluation

As of 13 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 11 inbound Pith citation observations for arXiv:2507.02856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02856 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:23:08.655335Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:13:47.447302Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:00:01.346988Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2da168eb-6191-4d7f-8310-64499fb7d47c · outbound

This paper cites • Reversible: Expansion is infinitesimally slow, maintaining equilibrium.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation • Reversible: Expansion is infinitesimally slow, maintaining equilibrium

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:10.621361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:23:07.892717Z digest=sha256:6ec4d55fad0aada48385259053b09f617aa4d8e8d1fc152a06c38690f8e22148

Observation 498c18bb-1daf-4009-a314-41e5477a80ee · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:10.418925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:23:07.958785Z digest=sha256:7bb50dd04c57d8cd53e19592a6d4f74fd27038d3bade0c8624e0091f43eb191f

Observation 478e6010-5de2-44f0-a451-1b74e04be053 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:10.185150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:23:08.044400Z digest=sha256:d32266acdaeb9a4caa5f443f36c9ae3820631ed5551a4efd4d66fcdcc08cdda6

Observation 5b9ed96b-4f81-40a1-ba5a-cf33897bef0b · outbound

This paper cites Calculate the total change in entropy, when a sample of nitrogen gas of mass 14 g at 298 K and 1.00 bar doubles its volume in an isothermal reversible expansion.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Calculate the total change in entropy, when a sample of nitrogen gas of mass 14 g at 298 K and 1.00 bar doubles its volume in an isothermal reversible expansion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:09.978120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:23:08.111058Z digest=sha256:be4bf3f3eafd7d30aabbc4af6cad4f2e301275f891e409fa655d0050bb5be1e1

Observation a63bb3a2-be4b-4674-aba5-836b42a28dc2 · outbound

This paper cites We find that MCQ estimates the highest accuracy followed by LLM-judges.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation We find that MCQ estimates the highest accuracy followed by LLM-judges

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:10.801232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:23:07.805474Z digest=sha256:95c6b1eb427bb18e2015c39533b32859e6419a46d9dfc3b9b3187c38f3bf3764

Observation 7992e3db-e97f-4297-a46e-d39ff3bb736c · outbound

This paper cites " " # The response can have more information than the ,→ ground− t r u t h . # I t can be more s p e c i f i c ( f o r example ,.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation " " # The response can have more information than the ,→ ground− t r u t h . # I t can be more s p e c i f i c ( f o r example ,

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T20:23:08.975801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:23:08.655335Z digest=sha256:b2236c6e28c19bbcbe8ee1f01de7c878acb607a89fc6a4eeca8e2b1d95d94850

Observation 06cffa44-e03c-4bf0-b34e-c30b8d3f07e2 · outbound

This paper cites • Reversible: Expansion occurs slowly enough to maintain equilibrium.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation • Reversible: Expansion occurs slowly enough to maintain equilibrium

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:09.822721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:23:08.199589Z digest=sha256:488a87a4347ad24832d7d13840f32bb5c0ec88830b205ae8a27f62ce1dd81b3a

Observation 6d569b6e-02ca-400c-9298-cf15e9353434 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.674329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:23:08.289554Z digest=sha256:30c09422768287e258b99e86c5a2b8959156e36ec346733d9188aa890fbb0886

Observation 1bf6c3ef-6f46-4469-b995-a7d238a450d7 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.489007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:23:08.356038Z digest=sha256:a19d4392573f06ae5ce2fb918d52f1b48f2e8c5687b1cdcca9432ba8e7f422c9

Observation 67b44c01-64bb-4470-99fb-410eb0e0a570 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.334697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:23:08.450396Z digest=sha256:4479e7b1d991ce6fe2e7b05f19860a18791d41a8c22b91d13c71d346ba027ac1

Observation 7b28d7db-4c2a-41b7-8754-b16ac26cda0a · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.125448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T20:23:08.546381Z digest=sha256:b02229258e285bac72ec16fb3adff7059ae54db22b6daa69892be779d0985ee0

Observation dc0d4a58-c0d3-4146-bf3f-70b7017a9a22 · outbound

This paper cites cloze procedure.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation cloze procedure

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:07.761432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:07.761432Z digest=sha256:57486d66e676426db55f0deefd6ee865060626e2ca0b4761a4c3b70a660d4737

Pith citing papers

Observation 7c5363e6-c712-4c37-aa28-cbdd340f86bd · inbound

Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions cites this paper.

Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:11.483898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:11.483898Z digest=sha256:0874e4d0355744c44a06bb678ef356cb20e8a7731b95c6f2890fa0adc268c050

Observation 6e6db20a-4026-483b-a1bd-cca37b5d2f2d · inbound

Entropy After </Think> for reasoning model early exiting cites this paper.

Entropy After </Think> for reasoning model early exiting Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:52:35.599288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T11:51:58.579048Z digest=sha256:3b89141977841e8de2d26e17bb04c5732a1ff3f0c37aa9ba55123eed59499fca

Observation 12c11f98-a00a-4ed0-ba99-d3424204496e · inbound

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants cites this paper.

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:12:22.093472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-18T03:12:03.909181Z digest=sha256:f9d7af03fb6420fafb1db912288cb5ed1a6169a5960d2bb39cb3e57d8b362f98

Observation 14106616-9b6b-493c-8290-c9e419e0e622 · inbound

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety cites this paper.

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-15T13:17:48.274611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:17:48.274611Z digest=sha256:19d272edfc31313cc1b176588426dd9d136f091a3d59538624fcd0bca7a75436

Observation 6b6c605c-69f9-42ce-8b0b-e98af22fd748 · inbound

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings cites this paper.

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:31.152523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T16:32:23.001110Z digest=sha256:88cf8e198d48bc63607126f28536ded5d739d1fe327928fb49c9aa34bfcd2672

Observation 9a7f5675-c8a4-4ee0-8a6c-a4089dc377a3 · inbound

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval cites this paper.

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:27:40.323764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T13:13:21.687950Z digest=sha256:7bbabfaee1cc610726758f165674d34f41727acb3a3d3833682ab654c9e1b99e

Observation 0104e497-a442-4600-88db-9684d3bc0a6a · inbound

Improving Cross-Format Robustness in Language Models with Multi-Format Training cites this paper.

Improving Cross-Format Robustness in Language Models with Multi-Format Training Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:07:56.335157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T10:13:22.927828Z digest=sha256:9192e5104757ff1bf1b2bd88976467d71d3a8b4068d266bd28ee7ae5a605729d

Observation caacba0b-b620-4009-8718-1170947ae885 · inbound

Storyline Trees: Hierarchical Representations for Long-Form Narratives cites this paper.

Storyline Trees: Hierarchical Representations for Long-Form Narratives Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:19:33.787019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T17:10:05.604537Z digest=sha256:4410949b786b141bbe1fae09262b8a6b468069ebc9b35da305f61dc1c0d47676

Observation 22d4eb0e-8361-4029-849e-ce002bb215c0 · inbound

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice cites this paper.

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:00:01.348676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-25T23:08:27.995672Z digest=sha256:692ecdd19307b79f166a3f4e8f3665f0cffb5a9907ca1c59fab0e8c18c105427

Observation c967abae-1ab3-4368-b4b6-5466d675666d · inbound

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets cites this paper.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:56.011737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:56.011737Z digest=sha256:f967a0927c568ad4623ea143670f5a1f0d1725ae5611d704eb518d1eade8620c

Observation d1c55097-9c9a-49dc-bbdf-d047c057a294 · inbound

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs cites this paper.

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:47.447302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:47.447302Z digest=sha256:0e029eb516235a0f87fb24da4ae28edc644aac18938c24856df24a74e0fc1ae0