Pith. sign in

Paper Citation Record · LEDGER

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge

As of 14 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 4 inbound Pith citation observations for arXiv:2412.17032.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17032 v3

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:55:14.902112Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:19:14.992316Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact3
  • verified fuzzy4
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c3fa33c9-9fff-4ece-937b-ff73f8c67a48 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.648021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.648021Z digest=sha256:400105fb1fb286f2a8d3db6414fa302c2d44148e5743d1b08ce8a5b2c8146860

Observation 1ccf4b38-807d-498b-ae17-62afec84d6b8 · outbound

This paper cites Learning to Recover Reasoning Chains for Multi-Hop Question Answering via Cooperative Games.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Learning to Recover Reasoning Chains for Multi-Hop Question Answering via Cooperative Games

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:55:15.509178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.654283Z digest=sha256:d03e5c9859cb0aa3af1e4c2d61cdb737ec30b9c7bf01bec941d795060ae0579b

Observation 15774753-2e12-4a10-9856-cf60599fbd4a · outbound

This paper cites The Llama 3 Herd of Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.659377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.659377Z digest=sha256:1c08c2611b88f9d6eb925d5ac32eaf75fec1d1676f763ca4515558c4aa8b8240

Observation 04c5b463-0714-4264-a42a-5631638fdc3a · outbound

This paper cites Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Constructing A Multi-hop QA Dataset for Comprehensive Evaluation of Reasoning Steps

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.664341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.664341Z digest=sha256:0f55d51366863c09b644343d919b2774c7022befc95dcbca69924c750ae7aa3f

Observation 500c559b-11b4-4c74-8f0e-2d9a29db9ef0 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.669252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.669252Z digest=sha256:864819f441326c996c530395ef85434199b2f4cf341c19a0225b93e175ea823a

Observation 06437c8e-b0eb-469b-93d5-ccc896513f73 · outbound

This paper cites Qwen2.5-Coder Technical Report.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Qwen2.5-Coder Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.674445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.674445Z digest=sha256:9ec01ba21b267344e4b41dd1a2cdceca0d2aebfc59a15bea1a0bce3799ef7ab3

Observation 3b02736f-ad64-4f45-9448-9412ecd0ec97 · outbound

This paper cites Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Open-RAG: Enhanced Retrieval-Augmented Reasoning with Open-Source Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.679583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.679583Z digest=sha256:a4d3bdf690d59bed1e2d87e9babb21516285b1263ddc44de44c5517d19dd74c3

Observation cd17f135-a0c9-4e03-b441-1586e96e3c1f · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.684533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.684533Z digest=sha256:3700173cc0a22688a3c5360ea81d31f91c5d307d3ae5bd3393c27c46f707242f

Observation 5c031ec2-67df-42d5-852b-7bb713bb8a1a · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.815764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.688361Z digest=sha256:568f80e6d01ea75e8b734bd1e99fa1993e8d341dad3c78a77e383bb21d18c4fe

Observation fbad84af-98c5-481e-bdbb-9f151e4b04d1 · outbound

This paper cites Mistral 7B.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Mistral 7B

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.692321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.692321Z digest=sha256:129eba34d87897a8df8fa3777f7efa6f9c455fcf56f5533750883ac8e46144bb

Observation 3a00bcdd-b499-4755-a9fd-2bb4ff41f467 · outbound

This paper cites Evaluating Open-Domain Question Answering in the Era of Large Language Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Evaluating Open-Domain Question Answering in the Era of Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.696605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.696605Z digest=sha256:a96ff9b654fe5bfb12d87919ee97740030be4cce84db240b32053cb2a483c78a

Observation 8851b154-f772-4be2-92ea-668b71633222 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.802024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.700990Z digest=sha256:19f134375d0000ea295120e389f55d61c63fecf41b7b46b6d1786d147b0ba2fd

Observation 0955d653-7425-44d3-94e7-ab48c022ca2f · outbound

This paper cites Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.705536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.705536Z digest=sha256:f7b1e8f0530af2e2a61e9b274cf3d8e9379fdecaf288569144de3d106af9b5ae

Observation d43c95b4-fa1f-4398-8c3a-b04ee565d55b · outbound

This paper cites Gonzalez, Haotong Zhang, and Ion Stoica.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Gonzalez, Haotong Zhang, and Ion Stoica

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.709849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.709849Z digest=sha256:429bbb8b7ae0ea45078fcaeaa6bd978ab62ae68e11a409745e630d6052463cb7

Observation bb878223-5a60-4e33-b24d-03755f0d0a05 · outbound

This paper cites Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.714204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.714204Z digest=sha256:6ec2581ab1155c721038ed71f1e9ab55186758d8d701fe3dd158ed2fbae3d3bf

Observation 02b048db-2b15-4816-8818-97996dcbb148 · outbound

This paper cites Large Language Models with Controllable Working Memory.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Large Language Models with Controllable Working Memory

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.719947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.719947Z digest=sha256:f8c084c951c3312f6a704eb93efd021993240db601a959c921c1a56ff850c0ab

Observation 55862471-7f16-4987-a417-96e3dfdfbf31 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.770632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.724903Z digest=sha256:2b0c91f03d8f4ad99a9860e9c1d18be0f36dc0cb844cc540dc0506637388b4ef

Observation c8441089-dcbf-443e-ace2-e2863a413b03 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.757538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.732120Z digest=sha256:0ef722a77710e726ce66f80c3c94530211e57472fc4a882f40c3de26fad35ccf

Observation 6e0d918c-de0a-447f-9e7d-c2d6a31eef0f · outbound

This paper cites Retrieval Helps or Hurts? A Deeper Dive into the Efficacy of Retrieval Augmentation to Language Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Retrieval Helps or Hurts? A Deeper Dive into the Efficacy of Retrieval Augmentation to Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:55:15.381405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.735869Z digest=sha256:972b8270a10d8afce34c32ed87b8e25d63490eaa9493a2d3ef283c205f710d61

Observation ef7d9c5f-36a5-43c4-8394-17cb944a296f · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.743955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.739708Z digest=sha256:4aed14fc14afb120f515ef90cf87268930027315d7e82b919b79fb7175cdcbf2

Observation 0fa58127-4eac-4b8f-ac01-1566fc47bb09 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.730181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.743987Z digest=sha256:fc676db8ca9f404c574c9c361f47022e396f537ab83a88b305710d9fd424f1cf

Observation 2c7b9b02-5790-4afd-8d9b-fe30b38ab018 · outbound

This paper cites Large Dual Encoders Are Generalizable Retrievers.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Large Dual Encoders Are Generalizable Retrievers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.748838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.748838Z digest=sha256:42510d1ce31f4d364a481ac2796a3797214094c0be375c1bf6fd0f9c5e21b09c

Observation 3ec105f1-b80f-4935-bc89-fca35185e6d7 · outbound

This paper cites Guo, and Xueqi Cheng.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Guo, and Xueqi Cheng

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:55:15.716433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.754170Z digest=sha256:ba3a365cee0284eceabeee552d680ab4d8ddb94dc67a6f930cc606aff059a782

Observation a05b6d6b-a9dc-4058-8d9b-a2fb603eff54 · outbound

This paper cites Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Investigating the Factual Knowledge Boundary of Large Language Models with Retrieval Augmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.758418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.758418Z digest=sha256:4d80f578cbd253384de6a10741b9ca7c0802e91d8cc808efb2ebe43fd98d9dcb

Observation 7e030baa-9c5a-4ac3-a45e-0d9f5340b4e0 · outbound

This paper cites Robertson and Hugo Zaragoza.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Robertson and Hugo Zaragoza

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.762778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.762778Z digest=sha256:934470bf2356e6e84e845acc1cd8cc9809c5b9c0db3da6ece0956c3fc3144a85

Observation 1afe389f-61dd-4943-85fd-8433934924f2 · outbound

This paper cites Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Mintaka: A Complex, Natural, and Multilingual Dataset for End-to-End Question Answering

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.767197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.767197Z digest=sha256:92bc84486eae5a1e55f9299aa64dd4347ac195ead1333ac79f68f2855080649f

Observation 219b7b2a-d79e-4d49-a115-73da98a3c1d0 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.694290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.771553Z digest=sha256:7303259fc4c47aafa0382bbb859a2e5e8060223bb71b1d6dd0df43093209cfe2

Observation 968fa1e3-3d56-43a0-af76-4db2c8f9dc94 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.681754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.776241Z digest=sha256:68262ad7b10eb7de948ad6c6cec5d6bdcb1e6d0c42c75750510a3239a3fe6216

Observation 5ad18645-bd1e-461f-8481-c266fa38f5a0 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.780421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.780421Z digest=sha256:aa2b393299019fd23ad37caa618c016fc7992b65d2baa55bfd993043773597dd

Observation 2632d840-75c1-46a7-80a1-555e5e21e3dd · outbound

This paper cites One Embedder, Any Task: Instruction-Finetuned Text Embeddings.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge One Embedder, Any Task: Instruction-Finetuned Text Embeddings

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.784568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.784568Z digest=sha256:5c45b0ed15778931acd17bcd6535bcc02fea0eca55df2b9ed769ba3dc6883c8c

Observation 550c42f2-7140-49fe-bea9-de6241d3e40c · outbound

This paper cites Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Head-to-Tail: How Knowledgeable are Large Language Models (LLMs)? A.K.A. Will LLMs Replace Knowledge Graphs?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.789259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.789259Z digest=sha256:c23061ab6f40935add093e1ecba9deb391dc6bbd24a839c9817b4687205cc92f

Observation 2c5cc6f1-7512-4a28-93d6-cc3f8b5542ca · outbound

This paper cites MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge MultiHop-RAG: Benchmarking Retrieval-Augmented Generation for Multi-Hop Queries

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.793714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.793714Z digest=sha256:7507aa4045a0520c2e559e70a2c05be77b29b1f78fb0aa06fff4f6f759a9b447

Observation bd6bcbcd-51f6-4416-ba5d-e533e5e6260f · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Gemma 2: Improving Open Language Models at a Practical Size

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.798033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.798033Z digest=sha256:0faf3327a16b11e9a0a333e3fac9926bef1d8f3df63e8e163a8f6867cf9c1659

Observation f6b3bcba-d562-4e67-904f-6eae0c253f65 · outbound

This paper cites Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:55:15.668466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.802177Z digest=sha256:b7572f95fa4aa1d8fe580060878d39dd97b412bb3f15f1f0b325c4be9c67eef2

Observation 8f51e6fb-cd0f-4d0a-8aa2-06f12269e604 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.805847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.805847Z digest=sha256:ee117c0aedd8a9e67f603dedb12f94ad83a7719d88f7aa816fab4c86bf1f9fbb

Observation 034e0e7b-aedc-43f0-9c26-b3fce9532a38 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.809422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.809422Z digest=sha256:c6b7eaf1d939c9e8b26e64d0b06ef5d3a970056133b2f5f8c68d017d5cbe70bb

Observation dac56400-d0a1-4af8-9c5c-c8eca28261bd · outbound

This paper cites Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop Reasoning.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Self-prompted Chain-of-Thought on Large Language Models for Open-domain Multi-hop Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.814055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.814055Z digest=sha256:976424fd80ef11ea5a92de2eba95fb3d07ed42c4492033c1d421832be973f0bf

Observation 06dc3a3d-0616-4137-9848-230da6870e7c · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.645885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.818452Z digest=sha256:049670a0ac123fefa68678ede8c0d7b1dfd27e59df3025f9ecfc02a08a2eb6bf

Observation c120a452-40e0-4e3c-b2a9-211633a44dbf · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.632440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.823098Z digest=sha256:7e41e5967b6c7fb80e5d9cf7d7eaa54f6161230b952b4b0d9306e775e3dc901e

Observation d376a72f-5f67-4854-83dc-c55de4e41a2f · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.827330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.827330Z digest=sha256:7a637509a126069bd890598ee08e452bb104b0c12a54bc4b2945765e5241c852

Observation 96b9467a-0c67-42a6-bbeb-f1efb39208df · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.831416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.831416Z digest=sha256:96354226589170ace99898dffd82a9c131817741f9313d7b7cf164ec2f9c88e0

Observation b0cfbe42-b200-4ed5-a25b-255c5e4a5cf4 · outbound

This paper cites Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Promptriever: Instruction-Trained Retrievers Can Be Prompted Like Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.836046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.836046Z digest=sha256:581869f2d736a716f8015d48085aa150cfd7afee9c3e7260bfa61348f7698c39

Observation b86cd3b5-149c-4dd1-aa52-7d208e5707f3 · outbound

This paper cites LM-Cocktail: Resilient Tuning of Language Models via Model Merging.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge LM-Cocktail: Resilient Tuning of Language Models via Model Merging

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.841141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.841141Z digest=sha256:97eaeb8bf0fe8a0b5ec4b78ff33ee825a3da9e5bd6b03c7cc24417ad3e2d46a0

Observation 2e64a726-be0e-4ea9-8590-bacd53e343a1 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 44

Resolution
verified exact
doi, observed 2026-08-11T05:55:14.939744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.845683Z digest=sha256:5bf1273c604d8299025f2bb8daea308acb580474c43f6ce45584dcfb085c2ed5

Observation 59b3c53c-3d12-4743-acb5-a38b61c4ffb7 · outbound

This paper cites Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Answering Complex Open-Domain Questions with Multi-Hop Dense Retrieval

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.850001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.850001Z digest=sha256:b915d850c804e3d888d7adba59f7f4139ef04ead9d915b27a0b1744e52695f6b

Observation 79a03d29-3262-41a0-b30c-5d87dfe99e0d · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.609094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.854280Z digest=sha256:b032bebcf60970eae0ed347b460c6a36b2da94d8eba053f3c1e69497fa915a40

Observation 69b744c9-ad5f-4f5b-ae1d-60f4bf7c28c2 · outbound

This paper cites Cohen, Ruslan Salakhutdinov, and Christopher D.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Cohen, Ruslan Salakhutdinov, and Christopher D

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:55:15.595699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.858421Z digest=sha256:ac8b7da71f9621f9f539506b9cc52aba90504f7c9762e33ddd844c1d13caf79a

Observation 07a26606-07bf-4f5e-a3d6-63e8dd4602e7 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.582715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.862741Z digest=sha256:bc16d050460ec10305892c3f44990f5907685eecc75e19164626f7010398ca74

Observation cb29f5ca-c30b-4377-8b3e-2a55bd819688 · outbound

This paper cites Yu, Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Yu, Wenhao Yu, Chenguang Zhu, Zaitang Li, Zhiting Hu, Qingyun Wang, Heng Ji, and Meng Jiang

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:55:15.569316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.867077Z digest=sha256:3dbd26f67eee3944b207e10e9edbf434051ccf28fe6c289ea1eb5542c98f275b

Observation 0004aca3-7805-4f9e-966e-2ea550cf9e33 · outbound

This paper cites RetrievalQA: Assessing Adaptive Retrieval-Augmented Generation for Short-form Open-Domain Question Answering.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge RetrievalQA: Assessing Adaptive Retrieval-Augmented Generation for Short-form Open-Domain Question Answering

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.871403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.871403Z digest=sha256:f683fe401fb8f7d9264e7131ecf37e6967adbf9d834eb32a10d858e7da634abb

Observation e7161233-33ae-4eeb-af2f-c2a0691163f1 · outbound

This paper cites Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought Framework.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Verify-and-Edit: A Knowledge-Enhanced Chain-of-Thought Framework

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.875132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.875132Z digest=sha256:534961b627778cb44f5a200dd51f440a5b19cbd13b967ef29229c3668bcd1e30

Observation 45b8792c-1047-4b93-9b26-99ed99da4b89 · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T05:55:15.555225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T05:55:14.878899Z digest=sha256:538cf9168eae6bb4538f64cc74e2724199a75db7ab1ffd6b820e8e847086f1ce

Observation 4954c2a4-3254-45a3-9c22-9d1d0b6196c0 · outbound

This paper cites Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Accelerating Inference of Retrieval-Augmented Generation via Sparse Context Selection

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.883477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.883477Z digest=sha256:9e9ea2ab83e37e02b6e18d8f4226f14f53c7ca0dbc6359daf58991d00fb51b49

Observation d3217fc6-19cf-4db7-b636-3b5574086b6c · outbound

This paper cites an unresolved cited work.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.888105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.888105Z digest=sha256:9d2afab0307b74596ca518ce36abb78000fb4ca5391a95d5e835cf4e2f439149

Observation c182206e-d2c1-45f7-a7a2-daa558bbfe16 · outbound

This paper cites EfficientRAG: Efficient Retriever for Multi-Hop Question Answering.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge EfficientRAG: Efficient Retriever for Multi-Hop Question Answering

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.892561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.892561Z digest=sha256:e59a7e0c8f9d3252a599ae62e5e077df5dc1bcfece9cc2ed9e4f744af7eb4d81

Observation 0b4d1837-57d9-4514-a222-7d94ae615d56 · outbound

This paper cites online" 'onlinestring :=.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge online" 'onlinestring :=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.897152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.897152Z digest=sha256:41e71a4f9ab9002245df001e429eb77e8f37d3a27c4db097c39d1d1f2fddfaea

Observation 4af56952-683f-4c5e-a41d-42f68a081186 · outbound

This paper cites write newline.

MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge write newline

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T05:55:14.902112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:55:14.902112Z digest=sha256:0f85de09adc25bb45443b6a571dd97b87f56f8a938f292935163933c0a959b13

Pith citing papers

Observation 0eee85ad-94c3-4cb4-99d2-2ce6b8a33120 · inbound

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation cites this paper.

InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:19:14.992316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:19:14.992316Z digest=sha256:abd9b5ff0178901a94ec8ea9789b4cc314db8726a8086b657a1baad293679eda

Observation 24d508a4-30e9-4d4d-8781-b77d9e3ff792 · inbound

Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs cites this paper.

Towards Agentic RAG with Deep Reasoning: A Survey of RAG-Reasoning Systems in LLMs MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T18:00:38.663920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:00:38.663920Z digest=sha256:e65045df90cfe5b9320994303b0ff0d06b58b75100d4f8bb5c903d3aa7067522

Observation 9c9cb728-9e24-481d-9046-4501a99e1a5d · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:00:56.135207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:de511a5db8a6821d67873a3293516e25a97dc3f2710545cd79bea4fba8228968

Observation e47dcf84-bf68-4a0b-ac11-8d4ce0bb4789 · inbound

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents cites this paper.

FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Tail Knowledge

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:27:56.385466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T10:01:45.332920Z digest=sha256:0079490d449d9c9cabf883746728c8ec5fb35d62808c37a111e5b20353d95600