Pith. sign in

Paper Citation Record · LEDGER

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval

As of 9 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2506.08887.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08887 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:04:38.641568Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact1
  • verified fuzzy48
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69266413-019c-46c6-9327-ef3c60df9862 · outbound

This paper cites Localizing mo- ments in video with natural language.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Localizing mo- ments in video with natural language

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.101686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.504246Z digest=sha256:e4267a21e2905c1f5c3f553b2d11244d438d31e36b9725da8c94cfb679de4b5f

Observation 8819666a-0375-46b3-b42a-42f9bb34c443 · outbound

This paper cites Vqa: Visual question answering.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Vqa: Visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.095804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.507071Z digest=sha256:15a8e4ff4599ea980e5dd63c39b2c02789e5639396f018cca3cfbce98e1683ca

Observation 4628a9ca-b838-4b54-b7ce-56bd9964f1c3 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.089994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.509532Z digest=sha256:d6cdad0e18ce268ab011080d17aa7793aed0b3a8e00fe71118c4d26272e9b7b6

Observation 0b308772-8625-4b7a-b330-03ca9c3b1748 · outbound

This paper cites Cross modal retrieval with querybank normalisation.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Cross modal retrieval with querybank normalisation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.083821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.512142Z digest=sha256:554e40e73737fce8429ff329c10bf51e0caf6adb153836c2540a0ec6679280d5

Observation dc198174-2b83-4f14-a585-9fe2ece5a579 · outbound

This paper cites RAP: Efficient text-video retrieval with sparse-and- correlated adapter.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval RAP: Efficient text-video retrieval with sparse-and- correlated adapter

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.077014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.514579Z digest=sha256:e6af461d58a6c0cbf8dbc6e52c38d535ac53eeb516d03b35b54c711f02d3f5ad

Observation b5ee4d76-0774-4ee6-9b80-77e60927dd13 · outbound

This paper cites Adaptformer: Adapt- ing vision transformers for scalable visual recognition.Ad- vances in Neural Information Processing Systems, 2022.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Adaptformer: Adapt- ing vision transformers for scalable visual recognition.Ad- vances in Neural Information Processing Systems, 2022

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.070737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.516855Z digest=sha256:cf90c7e008fb350c7b456efb497919eb0d3a7f06bb379c87104d527c4dd9ff2b

Observation 23c14efb-88db-425d-b8a0-770b3f7f91d7 · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.Advances in Neural Information Processing Sys- tems, 2023.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.Advances in Neural Information Processing Sys- tems, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.064336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.519656Z digest=sha256:06367dafc1d6d71492799818661f7be45eb01d1271ce5bd14d789284f87cae6a

Observation 32d8a49a-cd82-4d9a-ac3e-fd9236325f46 · outbound

This paper cites Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.521872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.521872Z digest=sha256:f467bbf771f9e9967138dbf99b331c519590c4ea08a8251a58468523230fb875

Observation f48e58cb-3cbe-4924-ba15-f4b2fe443d8c · outbound

This paper cites Prompt switch: Efficient clip adaptation for text-video re- trieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Prompt switch: Efficient clip adaptation for text-video re- trieval

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.058173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.524699Z digest=sha256:e680e4cd09a983beefb808451cbc5ff44540a2cfe171e7db66e12381cac0b946

Observation 97dfb5ab-823f-42b0-a475-1ea517437e7a · outbound

This paper cites Acnet: Strengthening the kernel skeletons for powerful cnn via asymmetric convolution blocks.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Acnet: Strengthening the kernel skeletons for powerful cnn via asymmetric convolution blocks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.051684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.526766Z digest=sha256:baa01a1d0d448457727f71e6628910f32add18d4f61cccdb2b0b6e7d65154af9

Observation 41a91c59-7797-47da-890b-a138a5bbbdcf · outbound

This paper cites Repvgg: Making vgg-style convnets great again.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Repvgg: Making vgg-style convnets great again

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.044688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.529008Z digest=sha256:7023fa27d4ed4d0903cadca4c9febe97345509579a08e5b5c75da0c451b307f5

Observation 1a91173e-b2c3-4f50-a64c-20c3fbbf3af1 · outbound

This paper cites Improving clip training with language rewrites.Advances in Neural Information Processing Sys- tems, 2024.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Improving clip training with language rewrites.Advances in Neural Information Processing Sys- tems, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.038445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.531373Z digest=sha256:b33b3be8769879f137db216a5af51636f4a49deae72ca392fa90440cfdbbfbe5

Observation 32c04091-3f0f-4b2c-8cd5-ef0c393c6c98 · outbound

This paper cites Multi-modal transformer for video retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Multi-modal transformer for video retrieval

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.031867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.533670Z digest=sha256:4fd1a1b97052576b39fe809ef520934f554ea3b3f04b5cdc872193367f1f4ca5

Observation 9f8a736e-eac1-4dc4-acbb-27e3224e3cfe · outbound

This paper cites X-pool: Cross-modal language-video attention for text- video retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval X-pool: Cross-modal language-video attention for text- video retrieval

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.025277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.535876Z digest=sha256:1ec47ae7fd3ba3b4bcdbd601a4cc08387e39b3efcea65d61f343034de310b75c

Observation 093d98e5-9f5f-4bc0-af90-1ad33cb684ec · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.538033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.538033Z digest=sha256:0d963c8ed5a8cb48dcfc81f6fb28f3288173fe61b5f9ff9b63844ddf970d6989

Observation 55a5bf61-fbbc-45a9-8b9b-848a786f3dde · outbound

This paper cites Framewise phoneme classification with bidirectional lstm and other neural net- work architectures.Neural Networks, 2005.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Framewise phoneme classification with bidirectional lstm and other neural net- work architectures.Neural Networks, 2005

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.018752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.540500Z digest=sha256:389d732fd49e24ea9b11dc7bf391249b40d9b81ba52bde823ab349723a8d98e2

Observation 57571e0b-2ffa-47b9-996f-bb3704489bd4 · outbound

This paper cites Towards a unified view of parameter-efficient transfer learning.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Towards a unified view of parameter-efficient transfer learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.542644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.542644Z digest=sha256:4a97e293d2c1e8f37083d80b3c95a52d403a1efc69039decfc5121005e6d301f

Observation 63af92d2-9708-40f8-bd12-91613cb213f3 · outbound

This paper cites Secret: Self-consistent pseudo label refinement for unsupervised domain adaptive person re-identification.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Secret: Self-consistent pseudo label refinement for unsupervised domain adaptive person re-identification

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.008141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.544754Z digest=sha256:0d2dcdd8560f59ee4927f53d1d001ecd5064b0ee248e13ff4935e054e61b0eee

Observation 3e4554c0-31bb-48be-b182-ab87fffcf3f1 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Gaussian Error Linear Units (GELUs)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.546950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.546950Z digest=sha256:a2b87058b1a6d47e08a56b97d652f1d727289e226e394bfe84615c1f5cba6cd3

Observation 82e4b745-76e2-4cf8-b463-6c323e630f29 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Parameter-efficient transfer learning for nlp

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:39.001674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.549128Z digest=sha256:b8f8704e22eeaeca891db2a8242d25abae37d8d4d68cbd0dd181ec90d9caed57

Observation 6625c00c-b459-4507-a833-94e918b277d1 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval LoRA: Low-Rank Adaptation of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.551012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.551012Z digest=sha256:796c83a9ebcc4d6f82af7280a4c0208a78be42c49d83311f9d46d225b3673b0a

Observation 0329a248-4f29-4b20-8274-a3a2a8f55253 · outbound

This paper cites V op: Text-video co- operative prompt tuning for cross-modal retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval V op: Text-video co- operative prompt tuning for cross-modal retrieval

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.994442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.553091Z digest=sha256:d9b24cda72c3fc6ec55cb3dde4e3c0d4b6ca20d3d074c47f353b568674952346

Observation 1c1c649c-d85c-4583-90dc-8678cbbb7fbd · outbound

This paper cites Vi- sual prompt tuning.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Vi- sual prompt tuning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.987864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.555136Z digest=sha256:f5311b50edb9ca1186dc6da004c9bf3522123c9d85557c5af624b14e703a1bb1

Observation e48bacc2-7389-4021-ac23-3c6a801563a9 · outbound

This paper cites Video- text as game players: Hierarchical banzhaf interaction for cross-modal representation learning.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Video- text as game players: Hierarchical banzhaf interaction for cross-modal representation learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.981161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.557171Z digest=sha256:22a5f554ce188847f6a2ab41a9a397b5d3d3b7616d8c9d901500dc42c9f1c3f1

Observation c79fb09d-18d4-46c3-8db5-a85b3247de15 · outbound

This paper cites Mv-adapter: Multimodal video transfer learning for video text retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Mv-adapter: Multimodal video transfer learning for video text retrieval

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.973336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.559599Z digest=sha256:8dcae843e9f4e4a3eb6990352f8c41cf6dd6312de1c5aa8f381566e0c5caad92

Observation 903f48d1-5eee-4799-b8b3-5b30427eb340 · outbound

This paper cites Deep visual-semantic align- ments for generating image descriptions.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Deep visual-semantic align- ments for generating image descriptions

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.966127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.561792Z digest=sha256:057d07b75d64e82454008902b3ff6a409b1e98f855a27705465bb03ae8eef35b

Observation 66598c33-e1f8-43e7-8b1e-b96415f926ba · outbound

This paper cites Maple: Multi-modal prompt learning.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Maple: Multi-modal prompt learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.959199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.563869Z digest=sha256:01cb018f5726ff70c40bdd0462fc17b377320f48535a5c71b555af58cf4581ff

Observation e1b60c54-ab64-4a65-b8c8-ddbdc45a349a · outbound

This paper cites Self-regulating prompts: Foundational model adaptation without forgetting.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Self-regulating prompts: Foundational model adaptation without forgetting

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.952290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.566045Z digest=sha256:e157c26a7af1e3dcd87046316152096b153c3f0c8e9f1ec4f4f32ab8b906c310

Observation be6e81ff-b0a4-4c84-96dd-753df8724479 · outbound

This paper cites Dense-captioning events in videos.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Dense-captioning events in videos

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.945206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.568125Z digest=sha256:6a353d9004b37c766d38f66d6c0896c63008d8956fc0433cac257e095895b2ff

Observation 37d14fd7-c852-4144-b417-c8a86a2965dd · outbound

This paper cites Courier Corporation, 1997.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Courier Corporation, 1997

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.570005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.570005Z digest=sha256:2858e88b80a3a0970f0a6813c6f0e5ab593fdb545b6b12e28e632c8ab957f8a9

Observation 0a4cfc7b-c815-4033-bf1b-934abdf1bb8f · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Less is more: Clipbert for video-and-language learning via sparse sampling

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.932510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.571998Z digest=sha256:7bc752dcac920347d07a8d93f16d33cc52e67fa38a76535193231b83262ca7ba

Observation 92789e07-17c7-4199-bb2a-94da68badea5 · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in Neural Infor- mation Processing Systems, 2021.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in Neural Infor- mation Processing Systems, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.925580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.573993Z digest=sha256:3a5c35007f52582b76f95bdd348649ed5f5c1eb426aa519d1f87b9bd1640bb74

Observation 6669b5e9-2724-497d-9548-25bcd1d9e577 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.576455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.576455Z digest=sha256:a9dfae6e98816bff45e8292abf0cd98ce5bb335a9b5048781a473290c1dba3f2

Observation 472df5ca-6f3a-4ae8-8f94-535037418fd3 · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Unmasked teacher: Towards training-efficient video foundation models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.914431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.578756Z digest=sha256:22f4c57ee7fa35a02cb51a9772a26c86e6cf0f3a91f104044a068ff035419107

Observation 2470c477-442b-4551-8d82-6811e1f69e24 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.907185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.580758Z digest=sha256:db92bc51e8e21a3509d1f716cf21f9904adc87e10854173488fb5078797f4bcf

Observation c4330bd7-a4c8-4b56-b714-198727c40189 · outbound

This paper cites Sgdr: Stochastic gradient descent with warm restarts.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Sgdr: Stochastic gradient descent with warm restarts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.900181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.583189Z digest=sha256:70d5d1768ee7ade23dbd493f87f69cfff51ab10fe85c7fcfde2831c2f44ab290

Observation e6db3627-e858-42d0-b83d-3907c87c4145 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 2022.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.Neu- rocomputing, 2022

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.893052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.585284Z digest=sha256:9c5b4100fd0cad89d61cdf5e8576c7e79cbfad08851b3f511ee77d7c730680e4

Observation d1a8c765-90e9-45e0-8f5a-0a79c9e6f4c4 · outbound

This paper cites Ea-vtr: Event-aware video-text retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Ea-vtr: Event-aware video-text retrieval

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.886181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.587649Z digest=sha256:7bca86b663aa92103c55ed44a96fff2273f49eafa21d117c53d0599099cdb580

Observation 675359d4-1a39-4b9f-b23d-8ee148482d73 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.879247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.589815Z digest=sha256:53c1cf5e9b34c4c0127e874165ac6a3e6db3bb84b04e8a983e8dce95f4bce483

Observation cb69a46c-7eb6-4721-ab43-ac139614fd12 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Representation Learning with Contrastive Predictive Coding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.591729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.591729Z digest=sha256:7bc7447558d1c00838f21d6eafe0c56a4e9c3d236c2dec10538ee046d23c59ef

Observation e4603b17-f9ba-4280-bf00-2e3f29eb956a · outbound

This paper cites Language models are unsu- pervised multitask learners.OpenAI blog, 2019.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Language models are unsu- pervised multitask learners.OpenAI blog, 2019

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.594112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.594112Z digest=sha256:8011ce6f0e9c53b8b8a9b1cb14dc8277516473e584515aae71c1b417ef511683

Observation f5a8f1e4-476f-4dbb-aadf-fa421504ecd9 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Learn- ing transferable visual models from natural language super- vision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.868194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.595898Z digest=sha256:06cc507155e6d4fd96de8999e12b837129fea7ef07db207be3b8d6a1269ccbfa

Observation 3ddc66e8-4b1c-40f1-a540-51cc0dec58eb · outbound

This paper cites The long-short story of movie description.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval The long-short story of movie description

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.859860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.597925Z digest=sha256:e5b5e2996c5a0bca9c05f4ecd8407d7e98f4c0f3c2d93b008dc7902f564709ab

Observation 20943e42-6839-4b62-ad48-600cbe32f003 · outbound

This paper cites TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.599960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.599960Z digest=sha256:42f559a46644d05608c3b91dab1aa7f1f3c4001ef17baa28329ba772827883f7

Observation 5050387a-a682-48d0-87e4-b2319976e873 · outbound

This paper cites X-reid: Cross-instance transformer for identity-level person re- identification.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval X-reid: Cross-instance transformer for identity-level person re- identification

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.852437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.602395Z digest=sha256:6c4d2700770c868cf65dc9b751b20b3d4da3e049c59be52d06ba064a1cfb2702

Observation a487edaf-4706-4cd6-8dc1-127138457d33 · outbound

This paper cites Zerocap: Zero-shot image-to-text generation for visual- semantic arithmetic.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Zerocap: Zero-shot image-to-text generation for visual- semantic arithmetic

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.845005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.604375Z digest=sha256:779a2aa55eed108b4a7acd558de2543c3891387253ae869f39fb60c52aed712b

Observation f644fda7-d10d-4b96-8176-32ee03ac7e91 · outbound

This paper cites Yolov10: Real-time end-to-end object de- tection.Advances in Neural Information Processing Systems,.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Yolov10: Real-time end-to-end object de- tection.Advances in Neural Information Processing Systems,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.837245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.606186Z digest=sha256:db32a4c81eed8244f2a84a12222cdfbf07dbdab5bfdbd6c61005defd71cdad95

Observation aaef0ede-80b5-4deb-9093-06115b7bdf9d · outbound

This paper cites Text is mass: Modeling as stochastic embedding for text-video retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Text is mass: Modeling as stochastic embedding for text-video retrieval

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.828037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.608355Z digest=sha256:6bea87efb1ce9b5c26987280fd629a024960075387b5060d1d91c71766ba7c3e

Observation 318d79f3-c194-4bb6-a2ea-277328ae9909 · outbound

This paper cites Disentangled Representation Learning for Text-Video Retrieval.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Disentangled Representation Learning for Text-Video Retrieval

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.610482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.610482Z digest=sha256:97c297957364a8876a5f768e639dc987bd4ed08f57cdf9738c62c0d562bc425d

Observation 300d6d22-2ed4-4a14-89bc-5e3f9fe69607 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.612750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.612750Z digest=sha256:f3e931745ab2cc73aedf3b8e678f8469acba8216f6a33aebd139836a685107c7

Observation 13d9f64f-a526-4fb3-8666-890cc69a6917 · outbound

This paper cites Cap4video: What can auxiliary captions do for text-video retrieval? InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition, 2023.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Cap4video: What can auxiliary captions do for text-video retrieval? InProceedings of the IEEE Confer- ence on Computer Vision and Pattern Recognition, 2023

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.820295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.614769Z digest=sha256:21e871160d32d953ea8966417f0a1a5c9b9002521280c62688afc144d264a1c2

Observation d3bb257c-fec5-4d6a-b455-d1824b8df63d · outbound

This paper cites Demystifying clip data.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Demystifying clip data

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.810675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.616890Z digest=sha256:fa902d1fc4fed16191ad1791ac9cb15fb38f9e4641986d444655b59e04b9d76f

Observation 62a76fcd-6b96-400d-a602-ea4d4ecd08fa · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Msr-vtt: A large video description dataset for bridging video and language

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.802277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.618838Z digest=sha256:e857aa55a034e9d1b9fc58723cbd2321325b76423e88873c1f4b786bbd9c4052

Observation a2d4fc0f-88d0-4624-877e-0ea3b83a0e31 · outbound

This paper cites Show, attend and tell: Neural image caption gen- eration with visual attention.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Show, attend and tell: Neural image caption gen- eration with visual attention

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.794086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.620852Z digest=sha256:3a2e3b6bc3ad9059a610afcad1fafba4f692b6bb946b5d2635c6c5b640341890

Observation ac8193da-961f-4068-b1fb-a10e92076d7c · outbound

This paper cites Clip-vip: Adapting pre-trained image-text model to video-language alignment.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Clip-vip: Adapting pre-trained image-text model to video-language alignment

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.785496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.623017Z digest=sha256:13412a1d173ed7e6f436d374c2ee9fa05169d4e1b528e20faf647f1a6babbc1d

Observation 4429bc67-e6b9-413d-8cfc-0eeb2161aa64 · outbound

This paper cites LLMI3D: MLLM-based 3D Perception from a Single 2D Image.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval LLMI3D: MLLM-based 3D Perception from a Single 2D Image

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:04:38.682071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.625078Z digest=sha256:61c36b3b3ccd0fc888f7d50206f3dce2a173040419cbd7f034be5c57737d2443

Observation c7252e5e-663e-4109-b549-6fbcafb3c3dd · outbound

This paper cites HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval HEIE: MLLM-Based Hierarchical Explainable AIGC Image Implausibility Evaluator

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.627382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.627382Z digest=sha256:700270c05e7490c9cdda94f6c89a611271fa65abe7d7b36a06a75767ef77844c

Observation cbdea981-e643-4f0b-9a62-f98e4d62c816 · outbound

This paper cites Dgl: Dynamic global-local prompt tuning for text-video re- trieval.Proceedings of the AAAI Conference on Artificial Intelligence, 2024.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Dgl: Dynamic global-local prompt tuning for text-video re- trieval.Proceedings of the AAAI Conference on Artificial Intelligence, 2024

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.777722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.629860Z digest=sha256:d9c5549e7a69f973923ae9316ccd43cb887e198b7deb9f94f0f5c808f8de29a4

Observation e338126f-3b20-4a2c-813f-89a4f1771003 · outbound

This paper cites Cross-modal and hierarchical modeling of video and text.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Cross-modal and hierarchical modeling of video and text

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.770886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.631911Z digest=sha256:c6d5f2aaab2c6d04f09ee00d5a5f198716fc606096520092a06860bd14e699f0

Observation d77fac82-44a4-4399-8190-8b237b154d11 · outbound

This paper cites Neural Prompt Search.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Neural Prompt Search

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:04:38.634022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:04:38.634022Z digest=sha256:97d73fd7ac21d220b066a20264509b76ccbebb6e25df195eed607a7b79474c22

Observation 7e438381-88ec-420d-8564-76c6862d6121 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Conditional prompt learning for vision-language mod- els

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.763763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.636387Z digest=sha256:e0a9270003957dc390f03b0f7e3c1e7525b9b72f2049fae7c32a64b85d6626b4

Observation 60c6db4a-d10a-4ccb-b7c0-3814677297bb · outbound

This paper cites Learning to prompt for vision-language models.Inter- national Journal of Computer Vision, 2022.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval Learning to prompt for vision-language models.Inter- national Journal of Computer Vision, 2022

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.755584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.638481Z digest=sha256:77f80094bd8701cbe34f7c04128fe705ddc899097d2561260877c6ba3bbbfac5

Observation 207f0633-d482-4615-9278-a64b4d1f201a · outbound

This paper cites 14,αandβare set to0.3and1.0, respectively.

DiscoVLA: Discrepancy Reduction in Vision, Language, and Alignment for Parameter-Efficient Video-Text Retrieval 14,αandβare set to0.3and1.0, respectively

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:04:38.747510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:04:38.641568Z digest=sha256:097cf474963579c47dcecb85f491ed8b444a11152e873528706c94c87d7ec4e8

Pith citing papers

No inbound Pith citation observations are available.