Pith. sign in

Paper Citation Record · LEDGER

Answer Matching Outperforms Multiple Choice for Language Model Evaluation

As of 9 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 11 inbound Pith citation observations for arXiv:2507.02856.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02856 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:23:08.655335Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:13:47.447302Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:00:01.346988Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved7
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2da168eb-6191-4d7f-8310-64499fb7d47c · outbound

This paper cites • Reversible: Expansion is infinitesimally slow, maintaining equilibrium.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation • Reversible: Expansion is infinitesimally slow, maintaining equilibrium

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:10.621361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:23:07.892717Z digest=sha256:d06c95213e8c841a0a8db3c653cc866e5ca38fdc7ed367e3031e5c30d3fd9d56

Observation 498c18bb-1daf-4009-a314-41e5477a80ee · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:10.418925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:23:07.958785Z digest=sha256:3c8507f479ea829562fa271c7e28f535e1ea515a55ddefaf6e6777627abf92e2

Observation 478e6010-5de2-44f0-a451-1b74e04be053 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:10.185150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:23:08.044400Z digest=sha256:8b82726bf9960412e1ba9327b03c19899dcc55023cf27868c5a729d8554cec76

Observation 5b9ed96b-4f81-40a1-ba5a-cf33897bef0b · outbound

This paper cites Calculate the total change in entropy, when a sample of nitrogen gas of mass 14 g at 298 K and 1.00 bar doubles its volume in an isothermal reversible expansion.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Calculate the total change in entropy, when a sample of nitrogen gas of mass 14 g at 298 K and 1.00 bar doubles its volume in an isothermal reversible expansion

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:09.978120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:23:08.111058Z digest=sha256:8303cce74fd0a583489b8262f07b72e9a5756994dbddb08a95b1a550e7848eba

Observation a63bb3a2-be4b-4674-aba5-836b42a28dc2 · outbound

This paper cites We find that MCQ estimates the highest accuracy followed by LLM-judges.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation We find that MCQ estimates the highest accuracy followed by LLM-judges

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:10.801232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:23:07.805474Z digest=sha256:60b9c5a783a2047f17ed5613625d9413b8d0d9a49085f03cef59e3187910ddc2

Observation 7992e3db-e97f-4297-a46e-d39ff3bb736c · outbound

This paper cites " " # The response can have more information than the ,→ ground− t r u t h . # I t can be more s p e c i f i c ( f o r example ,.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation " " # The response can have more information than the ,→ ground− t r u t h . # I t can be more s p e c i f i c ( f o r example ,

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T20:23:08.975801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:23:08.655335Z digest=sha256:2667d8cf08d073bb935e168dc3d10203031ebccb11f4ce34c882e3da9febd9a6

Observation 06cffa44-e03c-4bf0-b34e-c30b8d3f07e2 · outbound

This paper cites • Reversible: Expansion occurs slowly enough to maintain equilibrium.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation • Reversible: Expansion occurs slowly enough to maintain equilibrium

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:23:09.822721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:23:08.199589Z digest=sha256:cbc24fe582eb49cba218d6db1f0fc7e43524565fb9eb960e838fd1be4af884af

Observation 6d569b6e-02ca-400c-9298-cf15e9353434 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.674329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:23:08.289554Z digest=sha256:2167dc5c18cf1a520f5d68f579b2705dd0b22136d053dafb4a207d60940aa47c

Observation 1bf6c3ef-6f46-4469-b995-a7d238a450d7 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.489007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:23:08.356038Z digest=sha256:f14e793b9d1c1e416cd9ffd754071d7378edfb04fdd91fb8fa9a96b00cfe707a

Observation 67b44c01-64bb-4470-99fb-410eb0e0a570 · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.334697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:23:08.450396Z digest=sha256:c69f09c69055f5302018e20368f785d71218028e28b2aa85e540750c9c9f98af

Observation 7b28d7db-4c2a-41b7-8754-b16ac26cda0a · outbound

This paper cites an unresolved cited work.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T20:23:09.125448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T20:23:08.546381Z digest=sha256:ad0189310c2883b0c9db3d735d642797d54b0e7053189459577d26b438ddeac2

Observation dc0d4a58-c0d3-4146-bf3f-70b7017a9a22 · outbound

This paper cites cloze procedure.

Answer Matching Outperforms Multiple Choice for Language Model Evaluation cloze procedure

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:07.761432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:07.761432Z digest=sha256:89c4267fbb1807e783bebd61199aa0db3bb15fdfc12ebc6cfe3bc380a3f8afc7

Pith citing papers

Observation 7c5363e6-c712-4c37-aa28-cbdd340f86bd · inbound

Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions cites this paper.

Bridging the Knowledge-Prediction Gap in LLMs on Multiple-Choice Questions Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:43:11.483898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:43:11.483898Z digest=sha256:3502a3e91c7b3f9dbeee999b1791aca8e47806ba2e1498186b7898c3ae0ed5ad

Observation 6e6db20a-4026-483b-a1bd-cca37b5d2f2d · inbound

Entropy After </Think> for reasoning model early exiting cites this paper.

Entropy After </Think> for reasoning model early exiting Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:52:35.599288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T11:51:58.579048Z digest=sha256:438b9b3f57732928768e11f43e463670af3965875673ed78c50071f1a69a00a4

Observation 12c11f98-a00a-4ed0-ba99-d3424204496e · inbound

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants cites this paper.

Beyond MCQ: An Open-Ended Arabic Cultural QA Benchmark with Dialect Variants Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:12:22.093472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T03:12:03.909181Z digest=sha256:0b209fb602a6810c0e15b01dce706effd17d8da5702d9cb6d7332221b1e76762

Observation 14106616-9b6b-493c-8290-c9e419e0e622 · inbound

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety cites this paper.

Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-15T13:17:48.274611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:17:48.274611Z digest=sha256:d30d3d0fd5d16768708e241861130ef6ca628fcd97397ff3c768a11cb9e064ef

Observation 6b6c605c-69f9-42ce-8b0b-e98af22fd748 · inbound

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings cites this paper.

Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:31.152523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T16:32:23.001110Z digest=sha256:791810633048bcce884749f32e6ff23c627604773611785d883d66c0d49f5900

Observation 9a7f5675-c8a4-4ee0-8a6c-a4089dc377a3 · inbound

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval cites this paper.

Are We Evaluating Knowledge or Phrasing? Mitigating MCQA Sensitivity with ParaEval Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:27:40.323764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T13:13:21.687950Z digest=sha256:32cf4e762ebd1061d9fc93ea75b18639a18289c0fb3f93bb8de2d608c1fe1cf5

Observation 0104e497-a442-4600-88db-9684d3bc0a6a · inbound

Improving Cross-Format Robustness in Language Models with Multi-Format Training cites this paper.

Improving Cross-Format Robustness in Language Models with Multi-Format Training Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:07:56.335157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T10:13:22.927828Z digest=sha256:9979a3c02a37df74f89ad95871ee15b28f8a09ae8b7e5ece0d2cc939eb7790cf

Observation caacba0b-b620-4009-8718-1170947ae885 · inbound

Storyline Trees: Hierarchical Representations for Long-Form Narratives cites this paper.

Storyline Trees: Hierarchical Representations for Long-Form Narratives Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T04:19:33.787019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T17:10:05.604537Z digest=sha256:9defaf0ce7bf7dad8cba253e6924824187cbd447d0414815e6638787ddc73422

Observation 22d4eb0e-8361-4029-849e-ce002bb215c0 · inbound

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice cites this paper.

HelpBench: Assessing the Ability of LLMs to Provide Privacy, Safety, and Security Advice Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:00:01.348676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:08:27.995672Z digest=sha256:7866f140033a5b422f842bf7022b98961874d146ceadd05a63af51c831270143

Observation c967abae-1ab3-4368-b4b6-5466d675666d · inbound

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets cites this paper.

Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:42:56.011737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:42:56.011737Z digest=sha256:0e66d4df1ce6591bff0753ffb1a4afb64629e6860a5b3fcebd78319e6830760f

Observation d1c55097-9c9a-49dc-bbdf-d047c057a294 · inbound

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs cites this paper.

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Answer Matching Outperforms Multiple Choice for Language Model Evaluation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:47.447302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:47.447302Z digest=sha256:f958a2a5f66529e5ea20961f8fd97e7af04282937cb3ffe4356757a2cc6947d9