Pith. sign in

Paper Citation Record · LEDGER

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations

As of 16 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 2 inbound Pith citation observations for arXiv:2505.23121.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23121 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:56:45.083869Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:56:44.139662Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T15:07:03.943418Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 39eb7bad-7e14-4a37-b432-a3ba92efe39a · outbound

This paper cites To further expand the capabilities of large language models, multi-modal models are developed to in- corporate various types of input beyond text.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations To further expand the capabilities of large language models, multi-modal models are developed to in- corporate various types of input beyond text

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:46.288581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:56:44.104608Z digest=sha256:eafa16d99e1d49fc91407edc2a5e884219c3e2e36bc63bbaf99fbbb28adb56ec

Observation 042632f9-227d-48b1-a99f-1d6aafaba300 · outbound

This paper cites an unresolved cited work.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:56:46.018188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:56:44.307320Z digest=sha256:add925e0c72a26598b49088f4f17ea1f6dd44f2df4c204d415020c400bfc7d97

Observation ec812042-398a-4186-89c5-39076d662674 · outbound

This paper cites The first part comprises the data used for training, including multi-modal pretraining and instruction tuning.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations The first part comprises the data used for training, including multi-modal pretraining and instruction tuning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:45.915927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:56:44.368951Z digest=sha256:2e41120913c9d88d575ae3e1f75e5e6c1cea7d8d31f68b7508cc5c772134487d

Observation af2a5c7d-82d0-4f5d-90c7-866643a5b361 · outbound

This paper cites Statistics show that dialogues in TMDialog are much longer, with more relevant questions and answers, than other datasets.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations Statistics show that dialogues in TMDialog are much longer, with more relevant questions and answers, than other datasets

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:46.160579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:56:44.218545Z digest=sha256:19594f1534d615adb40614112119ef871a3c083bde17597d21260ca890f1e0b9

Observation 59980d4c-04ee-4f79-a176-d0aaa23b871e · outbound

This paper cites Experimental Setup During the pre-training and fine-tuning phases, we execute a total of 40,000 and 10,000 training itera- tions, respectively.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations Experimental Setup During the pre-training and fine-tuning phases, we execute a total of 40,000 and 10,000 training itera- tions, respectively

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:45.670532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:56:44.464050Z digest=sha256:9a6600391a2ff4922f465bc53350196ec36a3fe3c43c39f58e57bb832607de03

Observation 7e6c37bc-011f-4b5b-aed4-3af9d028278f · outbound

This paper cites an unresolved cited work.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:56:45.768378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:56:44.423514Z digest=sha256:527919dfe9ce85bdf58c51d4275fdb07706dd4a411f29ef6c12b6d2c0f088796

Observation be50690c-f9fa-400e-a3ee-2fc800bc2d10 · outbound

This paper cites We intro- duce a method that effectively utilizes the GPT-4 API for multi-modal multi-turn dialogue data gen- erating.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations We intro- duce a method that effectively utilizes the GPT-4 API for multi-modal multi-turn dialogue data gen- erating

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:45.313019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:56:44.622812Z digest=sha256:d41d5c9371e343549b94e94f1707eabacc3ee944854277922b7fb069c384745c

Observation de84d191-a7c3-4199-bb90-ee50ebbc5a4f · outbound

This paper cites While we applied prompt en- gineering techniques and performed data filtering andmodification,itispossiblethatsomelow-quality dialogues remain in the resulting dataset.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations While we applied prompt en- gineering techniques and performed data filtering andmodification,itispossiblethatsomelow-quality dialogues remain in the resulting dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:56:45.437717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:56:44.540767Z digest=sha256:9779c47a9c7c9f86ab29346062486d79d3b5238c3dc934d800cd06b8e5166c7e

Observation 605d3b08-b56e-45ca-a171-22c35b1ff878 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:44.677381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:44.677381Z digest=sha256:bd08aafd80b746c2335fbad10e450f08b3658497d79367256ab0dc8ac7c80d95

Observation 3893c571-7756-419b-a713-d39188771d29 · outbound

This paper cites Addressing Some Limitations of Transformers with Feedback Memory.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations Addressing Some Limitations of Transformers with Feedback Memory

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:44.790395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:44.790395Z digest=sha256:c6e0fe83aa5459d1dce6db1cdfe879a6ef50acd127117ea20a946fee9b4e4130

Observation 2a2f0c2b-c766-4e57-b8e3-b7387a072e7c · outbound

This paper cites The Curse of Recursion: Training on Generated Data Makes Models Forget.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations The Curse of Recursion: Training on Generated Data Makes Models Forget

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:44.958460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:44.958460Z digest=sha256:10b3876529f5953c7c8af3d8e6b27b6ed295c636eac8ca368a4c2279e6704417

Observation 807fbcf5-bb4d-4d4c-ae49-df28125ce891 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:45.083869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:45.083869Z digest=sha256:c9faac3ac36a5851b9cb38124ffdd8248882beb677f2bef975cc5b24d1403e14

Observation 9a3f4199-17f5-434f-bf6d-41922d510f10 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:44.868908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:44.868908Z digest=sha256:a0603a3fea2c27eef1ae52c54e8121951253ad3053223c2a66a999f7411f07b7

Observation 9a6a28ac-9720-4e62-adf8-f5cd80f3dc34 · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:44.719903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:44.719903Z digest=sha256:799ad92777d1f48a1fdfa698e0b0d43a68c5b14c42a859d60c0b29226303550e

Observation 67dff097-a4ef-4016-8731-8c3e97a36b34 · outbound

This paper cites ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:44.139662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:44.139662Z digest=sha256:375a0f1e224e86190335abe9befdb112ca6b948bdb1c84432f35fe1fd6ba3d64

Observation b97d02d6-97ab-4242-a303-e1dcd79285c4 · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:44.913208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:44.913208Z digest=sha256:a39ca1ec3be66a74fb1635cf5df8d0440cf051ee6c4fc8d081b9f317957e7df4

Pith citing papers

Observation 67dff097-a4ef-4016-8731-8c3e97a36b34 · inbound

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations cites this paper.

ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T12:56:44.139662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:56:44.139662Z digest=sha256:375a0f1e224e86190335abe9befdb112ca6b948bdb1c84432f35fe1fd6ba3d64

Observation d446a591-5606-4e75-86c4-16112f074709 · inbound

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning cites this paper.

StochasT: Learning with Stochastic Turn Depth for Visual Instruction Tuning ContextQFormer: A New Context Modeling Method for Multi-Turn Multi-Modal Conversations

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:07:03.944796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-02T15:04:18.880016Z digest=sha256:7472e5b779adc6e233d2d450c88f3fe0a9d9c538a2e53c5d0e3171be8308b929