Pith. sign in

Paper Citation Record · LEDGER

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

As of 16 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 4 inbound Pith citation observations for arXiv:2507.21476.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21476 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:48:09.853658Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:54.823963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T23:08:25.106938Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a953727c-fb1c-4391-b553-3239a5d0b3bf · outbound

This paper cites Phi-4-reasoning Technical Report.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Phi-4-reasoning Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.539696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.539696Z digest=sha256:02246567ae99f96b2e14f66a3a6092dbd1b767a04e96472f6222b1d5c83fe39f

Observation 12470ab8-9ba2-4b65-8aff-82ea03bf0790 · outbound

This paper cites https: //blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench https: //blog.google/technology/google-deepmind/ gemini-model-thinking-updates-march-2025/

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:12.204493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:48:07.793387Z digest=sha256:19286b1b90f70dd7bb28900e37ce978ee2e524a7526a2450f21a36c1a21e1298

Observation 683f9eb8-8304-4b22-9cc8-ac4ddc7001b2 · outbound

This paper cites In Working Notes of CLEF 2025 - Conference and Labs of the Evaluation F orum.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Working Notes of CLEF 2025 - Conference and Labs of the Evaluation F orum

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:12.130059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:48:07.887911Z digest=sha256:ae772a24c90b169fa849ef864f4d5c8eff775adaf5b0c605636f2716321f91bf

Observation 62bbde9c-ac16-4bc1-92d8-1c737c8671eb · outbound

This paper cites BIG-Bench Extra Hard.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench BIG-Bench Extra Hard

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.971609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.971609Z digest=sha256:16edb1022919c101b824c3707f5a6dad4a6e99b7c5de200f94eae20694dd7b50

Observation 80ac73db-40d3-45a8-86ed-31cbec686418 · outbound

This paper cites When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench When 'YES' Meets 'BUT': Can Large Models Comprehend Contradictory Humor Through Comparative Reasoning?

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:11.099671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:48:08.109209Z digest=sha256:4872ba0d7eab10fdb58ade5cc74195b210cea5723121e336fa2c7a8754b25562

Observation f28ad304-3e3b-4fcb-a08e-165d092eff2d · outbound

This paper cites Let's Verify Step by Step.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Let's Verify Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.158360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.158360Z digest=sha256:88528d62dcda5e2155ca79e9fa3e34992dc1d67422ca4bd2c8c971063575c59d

Observation 7cc49c60-75f7-49d8-b82f-f6234cad64d4 · outbound

This paper cites In Proceed- ings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 11069–11081.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Proceed- ings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 11069–11081

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:11.956878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:48:08.199060Z digest=sha256:741f54a434ded51ba91633169fc3853622e80fd323ffa614c0892cf0274420d1

Observation ad9be2f5-65c2-4393-bae4-9f82ba291152 · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.248926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.248926Z digest=sha256:acf838eed23a95b4e2208bacb22442b7ff9df7517b9c246dfb7cd543b750fbe7

Observation b084aec8-6fee-4087-a879-fb1e266d68de · outbound

This paper cites Inverse Scaling: When Bigger Isn't Better.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Inverse Scaling: When Bigger Isn't Better

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.300856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.300856Z digest=sha256:2c9ca5912df68ccca2f1f281d51c144c047b1ea06f66fe48bf9a0ad41477e6e6

Observation 75acb1a9-c225-43a2-9ff5-3e318ecf04b2 · outbound

This paper cites CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.423730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.423730Z digest=sha256:7bc3693a14bc7895238e85515f9bcea10a6042810905c9674736765ee83214a9

Observation 1c712a67-ee2b-4547-b5fb-db7e402e4fd9 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.472560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.472560Z digest=sha256:24c3532eade13baec48606e00397cd8ce09ad5a0ad59ba8756adc84f7f39a8a8

Observation 06d8d6d1-29d0-44c1-8461-11d8fd80da06 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.489572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.489572Z digest=sha256:7749c7a0bd17575af5144f5bae36061fdcb5d23ef8d1d2f1466939b137ddda3a

Observation c7ae82a2-ee3b-4401-854d-6be08571409b · outbound

This paper cites Phishing Awareness via Game-Based Learning.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Phishing Awareness via Game-Based Learning

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:10.816118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:48:08.645463Z digest=sha256:25a9f7a545476db47bac822bb6b00f1c3fe2c4aef543f7115d555056aa685e70

Observation f61fd422-528b-483d-81a9-585fc134065a · outbound

This paper cites Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.790009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.790009Z digest=sha256:220b459fac80468cfaf5e39e196de20aa109d4aa65974fe4aadec77dc8258579

Observation a26467da-b9b2-4995-ac92-11b90ac777cb · outbound

This paper cites Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.957774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.957774Z digest=sha256:285cb7547f4eb61469861b927196ca1944b5f9d58ba8e3d7ad345c552bd71e29

Observation 96465f1a-7000-4334-ae1f-ad044ac3a944 · outbound

This paper cites In Proceedings of the 2022 Confer- ence on Empirical Methods in Natural Language Processing, pages 2866–2879.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Proceedings of the 2022 Confer- ence on Empirical Methods in Natural Language Processing, pages 2866–2879

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:11.719843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:48:09.095642Z digest=sha256:3176c0c4e7bde0dcdceacb184f53a4da33a9bad6a42b118237460c17f399c720

Observation a0c34740-672c-4533-a901-2407262fa64c · outbound

This paper cites Signatures of room-temperature superconductivity emerging in two-dimensional domains within the new Bi/Pb-based ceramic cuprate superconductors at ambient pressure.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Signatures of room-temperature superconductivity emerging in two-dimensional domains within the new Bi/Pb-based ceramic cuprate superconductors at ambient pressure

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:10.523027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:48:09.185449Z digest=sha256:2abba78009560954636358ea6facafbfb0719dc41de5e31404bf07f0c3c655e5

Observation b61d379f-95df-47da-bca1-d3318f3f7d37 · outbound

This paper cites In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, pages 4213–4228.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologies, pages 4213–4228

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:48:11.432296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:48:09.311008Z digest=sha256:d677c06190aa85932d2f12f4eca1e080926b00271c2190b98d52ea6d5f14d3cb

Observation 656c4b5c-aa05-4ebe-8311-f2e126843037 · outbound

This paper cites Uniformly rotating vortices for the lake equation.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Uniformly rotating vortices for the lake equation

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T12:48:10.276133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:48:09.407105Z digest=sha256:8299212597d2190001d6f1271eb6387689beac97d33e60765947e0fd128dc845

Observation 9c75b147-50f5-4294-a4eb-4c281786a722 · outbound

This paper cites arXiv preprint arXiv:2502.18080.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench arXiv preprint arXiv:2502.18080

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:09.550871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:09.550871Z digest=sha256:be8a2d64359b4d2f25487a9af2b3a137707d285f836875da8be19f032f879f6c

Observation b202ac87-475a-49a2-a998-8a4f9e111b8b · outbound

This paper cites Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:09.697591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:09.697591Z digest=sha256:db3ae6165c0d742cf824157cab7af27d351d15b8fe471dba94383579fb04f2f1

Observation 2720ae89-6683-4c2a-a731-9dbeb5c2d960 · outbound

This paper cites Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Bridging the Creativity Understanding Gap: Small-Scale Human Alignment Enables Expert-Level Humor Ranking in LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:09.853658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:09.853658Z digest=sha256:bb7fbdc33f6eb423bfabf3b3c1147b352562a868e2966c57f8eb5945af15683d

Observation 66882ed1-a3c6-4913-8e54-06cd911d2f7f · outbound

This paper cites AmbigQA: Answering Ambiguous Open-domain Questions.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench AmbigQA: Answering Ambiguous Open-domain Questions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.379481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.379481Z digest=sha256:27deaaa5cec3e8f6b0e70a3d49f10a62098855bb43d18323f8572b9bbfa952db

Observation 5a21e266-f982-45af-b54a-c35634c4fe36 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:08.040390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:08.040390Z digest=sha256:c06a6bedeee759148b96c7205631e82b5d8a780e8a6c1cea53ce241029e42727

Observation 1c50d8ef-ea44-4d1b-a529-5caa53466861 · outbound

This paper cites ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench ToRA: A Tool-Integrated Reasoning Agent for Mathematical Problem Solving

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.929176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.929176Z digest=sha256:281c6f716b4da00951f49e7bed414a2645702001f36a498d78d3cc205b47a2f2

Observation 64403be8-ce55-4562-a051-38c6ef8c9e4b · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.648384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.648384Z digest=sha256:c8f52f17459d45a7c065f1bddbecc0562b44f879cc8099058e3aaf778fa61e4d

Observation 3118d44f-b9c8-4fe4-8ee5-0a339bc9a225 · outbound

This paper cites ARC Prize 2024: Technical Report.

Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench ARC Prize 2024: Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:48:07.697172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:48:07.697172Z digest=sha256:86953d75a6d319bc97cb56270b95569f422199e6ee04460f2e05f18e3047db6a

Pith citing papers

Observation 7da47090-9728-4baa-8a07-054361c8b7ce · inbound

Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care cites this paper.

Relationship-Centered Care: Relatedness and Responsible Design for Human Connections in Mental-Health Care Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T20:22:12.729190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:22:12.729190Z digest=sha256:e8ac8fc0b1b81d96643dd144a4e688b7ebf3f111df0f82368847f0e60ff3372b

Observation 0dc073df-7454-4b5b-beb0-9f15dbc037eb · inbound

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models cites this paper.

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:08:25.110234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T23:07:38.352198Z digest=sha256:042badb35acf8cd84bd78f9890397b54388833588b94226e6b4e591210af5902

Observation 6af44ac6-42c1-4c0b-87bb-15bc9f946ce6 · inbound

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models cites this paper.

HumorRank: A Tournament-Based Leaderboard for Evaluating Humor Generation in Large Language Models Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T17:05:10.690456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T17:05:10.690456Z digest=sha256:0cdcb6f82acc6093121eb3b0ed7bb09a7e22581a15c5f1a09e734a66376d0e50

Observation 3d3eebc3-281f-4c1e-9d41-a48f9b77f061 · inbound

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges cites this paper.

Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges Which LLMs Get the Joke? Probing Non-STEM Reasoning Abilities with HumorBench

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:54.823963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:54.823963Z digest=sha256:3c27bb7ef3861e6d3b1a939394739da8d283a94f222570f2639e94952904a7bd