Pith. sign in

Paper Citation Record · LEDGER

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval

As of 15 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 0 inbound Pith citation observations for arXiv:2601.20597.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.20597 v2

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T10:40:01.902254Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

81 of 81 outbound references displayed

  • verified exact2
  • verified fuzzy69
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1116804f-5659-4cc0-9b45-1d9667496262 · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.709794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:2ebb2c79d879755961c1b3fdedf89b4f4c08751c3470203594be1dcd5d159c71

Observation 4f82f5e7-7c65-4c5d-9e40-dee810632c0d · outbound

This paper cites Ashok, K.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Ashok, K

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.714676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:4fc2d8727dea964c7abaf61d5f10cef7e824d0c5a20282365e1919c401d64116

Observation 6e0b33d6-b5fd-43b8-860c-b96508f0c9e3 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.707392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:9aa9ddf46548c1c1b604f3d4b8449d9036d9fb7bcfdf0c5e2ad5892b16aea2d3

Observation 37f15509-ad8c-4d42-87e0-b2800cb4290c · outbound

This paper cites Non-autoregressive cross- modal coherence modelling.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Non-autoregressive cross- modal coherence modelling

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.716828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:1b12c905ec202ec32ed5332e061d214dcccb3ae88428d69921142f0cc2fd68f8

Observation 1982bc2e-ebc8-4b69-9931-0a75c947a475 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Activitynet: A large-scale video benchmark for human activity understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.705324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:23a1fbe89fe8bcf575fdf6195c672247b5ea8f5f63f22687089e11f47707256c

Observation fdc776cf-03f6-418a-a7fd-e847250d8444 · outbound

This paper cites Online fast adaptation and knowledge accumulation (osaka): A new approach to continual learning.Advances in Neural Information Processing Systems, 33:16532–16545.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Online fast adaptation and knowledge accumulation (osaka): A new approach to continual learning.Advances in Neural Information Processing Systems, 33:16532–16545

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.709995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:5bc4350de6133cebb386f7b4633320049a130b831f34cd278b97c9d14b15397e

Observation dcab8dd2-25f9-4a20-9dff-5d406f0fad9e · outbound

This paper cites Clumo: Cluster- based modality fusion prompt for continual learning in visual question answering.Journal of Artificial Intelligence Research, 83.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Clumo: Cluster- based modality fusion prompt for continual learning in visual question answering.Journal of Artificial Intelligence Research, 83

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.712837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:c82fec95773f15e9b2b46591f295fd495aedd9f99175e72c31c1452b6142e5b1

Observation 7681fd1e-64b4-4899-8472-6a0cfa999907 · outbound

This paper cites Castro, Manuel J.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Castro, Manuel J

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.701475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:2eb32b013e09e0599fd14cd221c5b323a33b83e4d286ecbcd4e46884f7effecc

Observation 3b2f0bbb-149a-4ba9-8c8a-369115acf023 · outbound

This paper cites Fine-grained video-text retrieval with hierarchical graph reasoning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Fine-grained video-text retrieval with hierarchical graph reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.694266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:1e669a66e2a05809cc4be811d634eb9c53d5aeedda6a9a25e46edb072c234c1f

Observation da9974f5-19a1-4cfb-9e61-c1b7a53ee94f · outbound

This paper cites Vision- sensor attention based continual multimodal egocentric ac- tivity recognition.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Vision- sensor attention based continual multimodal egocentric ac- tivity recognition

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.703508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:6605a98d70de3027d0fdcab95251e800909b32c34caab19c5c62f8b748081252

Observation 169e921f-6c01-41e7-abb1-4d21d640f6a5 · outbound

This paper cites Teachtext: Cross-modal generalized distillation for text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Teachtext: Cross-modal generalized distillation for text-video retrieval

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.705551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:71561b90114128697ecbcd7e6e87ba40246bfe2eae44104cc3d744a4069504b2

Observation 8c686c0e-6338-41a2-9249-55373b401d8d · outbound

This paper cites Don't Stop Learning: Towards Continual Learning for the CLIP Model.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Don't Stop Learning: Towards Continual Learning for the CLIP Model

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:40:50.904706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:5ef881e1a9f8a7d3dfc7c4a04fed1bad59eada4b90127564756fbfe8e2ccce39

Observation 227e82ff-e973-409c-b3cb-f3a67b75ef71 · outbound

This paper cites Podnet: Pooled outputs distillation for small-tasks incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Podnet: Pooled outputs distillation for small-tasks incremental learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.696497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:0595e143ebb4a0243c9ccd171f215f926a603831e92b1290e1a0b3d73480486a

Observation d136fbec-b734-4c12-aa50-905633e73a58 · outbound

This paper cites A feature- space multimodal data augmentation technique for text- video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval A feature- space multimodal data augmentation technique for text- video retrieval

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.598654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:eeab5c0099647c59928e1cebdfee6705361168ad4cd913f2f3b4c917018d544f

Observation cd9416df-fac6-48d9-91d0-cdce30f56e10 · outbound

This paper cites Uatvr: Uncertainty-adaptive text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Uatvr: Uncertainty-adaptive text-video retrieval

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.600733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:bffb56f07a9549ed5b73b287c42166c488b83049d914e7fd15c7bf3ea21c24b2

Observation c99db1a2-c20f-4648-a484-33184fd37811 · outbound

This paper cites Transferring image-clip to video-text retrieval via temporal relations.IEEE Transactions on Multimedia, 25:7772–7785.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Transferring image-clip to video-text retrieval via temporal relations.IEEE Transactions on Multimedia, 25:7772–7785

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.618257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:8cde2e3da0a6134a95ee3ad26c9afd971747a41ef9d39d649dfae09f9708784a

Observation c0b5fd60-b1d3-4dee-b8fa-af287fae4eef · outbound

This paper cites Multi-modal transformer for video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Multi-modal transformer for video retrieval

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.620652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:4623e7383538c71aa4788973072200313f5f8c2a19451015a73c3a63ae36e81d

Observation 52c14299-a180-40fa-93af-00ca30203a9a · outbound

This paper cites X-pool: Cross-modal language-video attention for text- video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval X-pool: Cross-modal language-video attention for text- video retrieval

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.625377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:b60bdabd7b7609a2ae10a6596c40c6b64c12e9d503b7f27b67a968238e737146

Observation ea1e9ff3-71d0-499b-b885-fedd0577eae9 · outbound

This paper cites Dyson: Dynamic feature space self- organization for online task-free class incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Dyson: Dynamic feature space self- organization for online task-free class incremental learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.630233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:b4b5885e24c03b6a9e00ab73d9591e36e8d48d877a4631e630b7d01bd23c7300

Observation 235bf544-8942-41bd-9e04-cc4ce4c8aa7b · outbound

This paper cites Learning a unified classifier incrementally via rebalancing.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Learning a unified classifier incrementally via rebalancing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.639797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:4c15893fc7200f8566bd4cf83c8709c46b773858f261218644c3daedc1ae7115

Observation 23e57aa6-d95f-4629-a46d-21e504b01840 · outbound

This paper cites Curiosity-driven class- incremental learning via adaptive sample selection.IEEE Transactions on Circuits and Systems for Video Technology, 32(12):8660–8673.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Curiosity-driven class- incremental learning via adaptive sample selection.IEEE Transactions on Circuits and Systems for Video Technology, 32(12):8660–8673

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.635083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:fcd368d8e0cc29f1858d4d11aab368f886de7fd91b7cf2e1c4a646593cbefdf3

Observation e604ae7e-8b2e-4f53-a92d-f059bf6d3837 · outbound

This paper cites Distilling causal effect of data in class-incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Distilling causal effect of data in class-incremental learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.675445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:994b59e99987f4219071f0c0e393e851cae8b811a53aa10e325bb7cef31e84ff

Observation 8480298d-d87b-4d0b-b8ea-1a0e5e4af4bf · outbound

This paper cites Neural collapse inspired federated learning with non-iid data.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Neural collapse inspired federated learning with non-iid data

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.609558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:741d27aba56c90d885a43d03692b8dca89106be2399f3989381baa10842f14b0

Observation 4dfa2c9d-afc7-46a9-8162-c85451acfb4f · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.582087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:b91b8da7edbe003030e890fcae73511648da2025f0121badcbef12c93b85b5c5

Observation 8b9fcda3-507a-41ba-92c2-0993e30f8931 · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.573147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:a99d057d4f2121c339da7a64e793e63d1e1f6482e62514a65cb62fde38e06f12

Observation ddb1b485-f960-4d80-b661-c6fd3649bfd3 · outbound

This paper cites Hybrid-tower: Fine-grained pseudo-query interaction and generation for text-to-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Hybrid-tower: Fine-grained pseudo-query interaction and generation for text-to-video retrieval

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.567018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:0c024b8b6adf4a7c4aa2534557eab5f9319a80c22cab5e37bf97538a2a5ecc74

Observation ab85f800-dd7a-4da5-a188-4f29c50e368d · outbound

This paper cites Bakker, Nicu Sebe, and Michael S.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Bakker, Nicu Sebe, and Michael S

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.569168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:8c9a1e5ae80e806b63ea72be4b232c15e3c2b07ffee13110a53f8f0dd93f6c5a

Observation afe372b3-fb8c-47ef-b4ee-d415d6312775 · outbound

This paper cites Dynamic integration of task-specific adapters for class incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Dynamic integration of task-specific adapters for class incremental learning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.576420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:d422b26eaa57b42cd212ee04a24d85c3c4a1026e4252e31e004411041bc0eb6d

Observation a2dd7117-829c-4ad6-ba6f-a74700987012 · outbound

This paper cites Multi-modal inductive framework for text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Multi-modal inductive framework for text-video retrieval

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.658278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:46eb9a6b22973d79a152e6c61a61db1b15b4f3036ddf507f343c70a2e0430c0b

Observation 3bcd1dc6-93b9-493d-8c62-a0ebf4ce32be · outbound

This paper cites Learning without forgetting.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Learning without forgetting

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.663084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:1d0e9389b0faffed4b0a7476407642e892aab42c77095a0b2bca938eee35a16f

Observation edebad7e-2da5-4a8d-9e83-d68f5976bc11 · outbound

This paper cites Anchor assisted experience replay for online class- incremental learning.IEEE Transactions on Circuits and Systems for Video Technology, 33(5):2217–2232.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Anchor assisted experience replay for online class- incremental learning.IEEE Transactions on Circuits and Systems for Video Technology, 33(5):2217–2232

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.667584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:0252e12ba69b84017aeb6f7e7fbf6d9b58fbc3bb66d2e21aca3a4a7e6f389ad5

Observation b3b04c33-f98c-4d7e-b06a-1c2d3b4ab5ff · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM Computing Surveys, 55(9):1–35.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.ACM Computing Surveys, 55(9):1–35

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.687741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:17a1ef580bb0b06b0cc8dde4eb2cd1ea6a3a1cf6689ffef6acdc6614ece74b04

Observation 6f9179a4-001e-4107-967b-5669b4a120f1 · outbound

This paper cites Use what you have: Video retrieval using representations from collaborative experts.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Use what you have: Video retrieval using representations from collaborative experts

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.645249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:795bab9cb7d39c3d95a68c25e5677fb27322e6dfa2696a02afca5b1616b97330

Observation eb99e289-a5c8-44d6-bc01-1f01e42901f9 · outbound

This paper cites Adaptive aggregation networks for class-incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Adaptive aggregation networks for class-incremental learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.647251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:69128c06e351ef84473fd70de1d5c85aa9b8c2ad03fa1b55ac6dfca956dcad80

Observation dd4eb987-5ff5-4cc8-8321-8b6609bd0d7d · outbound

This paper cites Ts2-net: Token shift and selection transformer for text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Ts2-net: Token shift and selection transformer for text-video retrieval

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.640901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:e464239e9fc986803d2c7a7afe5afa82e45f13045738685178e8e99494976211

Observation a06aaf53-a0db-4a1f-8699-a4c9bf11e4e1 · outbound

This paper cites Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Clip4clip: An empirical study of clip for end to end video clip retrieval and captioning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.643129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:9431a1801d542c4820148db1d036ace4563c81d3bf8e8c760dad962b63f34a5b

Observation b9c1695b-e4ae-4135-9907-1c55d730415a · outbound

This paper cites X-clip: End-to-end multi-grained contrastive learning for video-text retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval X-clip: End-to-end multi-grained contrastive learning for video-text retrieval

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.649523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:caccb1fda4b6673070b03603f487377460dd6a30dc8f119ca0676264ed7104d1

Observation cc713bd8-6722-4803-9dcb-d562db5fcfc9 · outbound

This paper cites Packnet: Adding multiple tasks to a single network by iterative pruning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Packnet: Adding multiple tasks to a single network by iterative pruning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.674326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:458c3c30e6f51e2ff46fd4aff585efc274e0c1ebf771cfe607c706c3c4bdcf70

Observation 69ef110c-35c8-4964-b321-bec612227c23 · outbound

This paper cites Prevalence of neural collapse during the terminal phase of deep learning training.Proceedings of the National Academy of Sciences, 117(40):24652–24663.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Prevalence of neural collapse during the terminal phase of deep learning training.Proceedings of the National Academy of Sciences, 117(40):24652–24663

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.661076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:62630bb965827abcca856eb5300ceb6f948e9ad2bd84a09a4f59a84344205e9f

Observation ca335b31-e5ac-4caf-8464-f3fcb3b3b882 · outbound

This paper cites Prabhu, P.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Prabhu, P

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.591003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:17719ef712a81998d3ac54bf4e6bff08d79331282b15a7198764f5d2c1861305

Observation 396540dc-c96c-4a1c-9ffd-0ff774ac9832 · outbound

This paper cites Learning transferable visual models from natural language supervision.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Learning transferable visual models from natural language supervision

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.605292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:0e742bacac3c7fcd9c50bb79707771da25e98f881a609e98c35b99d57d05314e

Observation c2124e43-fd82-4aff-86a1-65bc36a5367d · outbound

This paper cites De Melo, Benjamin Van Durme, and Rama Chellappa.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval De Melo, Benjamin Van Durme, and Rama Chellappa

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.620155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:12e6f4cf4d879685777fb3b4306338d704a122cc0b1393133af0441a49c8e6f4

Observation db425f6b-aa98-4063-b043-ce955592cef6 · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.591571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:7a623f9fcbb5f4631851ab62a8c4519b2ca16189972194c9578b51aa7d2d7318

Observation 01acc3f3-6af0-4b52-971d-ed6d2bd58277 · outbound

This paper cites Relation triplet construction for cross-modal text-to-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Relation triplet construction for cross-modal text-to-video retrieval

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.661864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:175554027c1ac96e348a627d4214468f0e5f45fd84668893a552eb7dc699cc50

Observation 1e61f864-8c9c-446a-8fd5-297b26d80f5d · outbound

This paper cites Spatial-temporal graphs for cross-modal text2video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Spatial-temporal graphs for cross-modal text2video retrieval

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.666346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:1e86904579cbc52be9bc0e52298968acd58ba39791df7a3bc9cdf7b27ef12b30

Observation 1921058a-4e86-402b-8514-441de59d9813 · outbound

This paper cites Learning endogenous attention for 12 incremental object detection.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Learning endogenous attention for 12 incremental object detection

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.648600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:894d101df901a8ce986d31203829132be51ea691e35dbfa5df25d7cc2702c58f

Observation a58a7f6e-660d-4dac-b416-4111b8a21b09 · outbound

This paper cites Multimodal continual learning using online dictionary updating.IEEE Transactions on Cognitive and Developmental Systems, 13(1):171–178.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Multimodal continual learning using online dictionary updating.IEEE Transactions on Cognitive and Developmental Systems, 13(1):171–178

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.651198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:d4d716f59981e7d7c2c7cd3bf6276c12fe65e4023eec0516ea4c7ed73d58ea6a

Observation 3fdd9fa0-b3d5-4eb6-b708-802f7c1f237c · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.669820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:0f2aaa6260266b4271fd516212f886bb9aa21a657d39ba97a9987aa7df26f10a

Observation 8f1985a2-8ea5-4538-ae4f-cee317aeadb9 · outbound

This paper cites Topology-preserving class-incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Topology-preserving class-incremental learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.670974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:b7500f2ab6c71fcbee60df114bbed7d3649c9ea6ea20ed87e75b1d89cf95e6c0

Observation b7428cfb-c325-49d1-b8a4-7d36a9a4d9cf · outbound

This paper cites new” while consolidating “known.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval new” while consolidating “known

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.682933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:f60fcf4a8568d29d995a52146177bef61219e1fa428cc7c7090e73525c8d00c8

Observation e983ac38-969c-421f-ace1-8e98a8a855b8 · outbound

This paper cites Holistic features are almost sufficient for text-to- video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Holistic features are almost sufficient for text-to- video retrieval

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.615710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:5a123c0e13f3121140e667d74b9830e3c06cf2002264b32cd69ed0c03da58354

Observation c3ecd1c6-91d7-406b-82a4-28246964eb74 · outbound

This paper cites Dualcp: Rehearsal- free domain-incremental learning via dual-level concept prototype.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Dualcp: Rehearsal- free domain-incremental learning via dual-level concept prototype

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.636271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:3ae41e05f4b2596008376f649c3eaa21224fca5363c81592e90b78d82c409c42

Observation 2ef25a4f-1dcc-43b9-b5cd-c8d759a4a279 · outbound

This paper cites Semantic knowledge guided class-incremental learning.IEEE Transactions on Circuits and Systems for Video Technology, 33(10):5921–5931.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Semantic knowledge guided class-incremental learning.IEEE Transactions on Circuits and Systems for Video Technology, 33(10):5921–5931

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.679224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:4a77c105f32fcb63cfefe9afc466a821f0fe73afadaadb4c751927f88e9f0483

Observation 5be729ff-f5be-4411-86c5-2238436c0c81 · outbound

This paper cites Non-exemplar class-incremental learning via adaptive old class reconstruction.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Non-exemplar class-incremental learning via adaptive old class reconstruction

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.676623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:289b7134f803fe456712c887e1f9f526a644ae926978649ad0b6f8987e98e0dc

Observation b24ac00e-0264-4a54-9f7b-0aca9746794b · outbound

This paper cites T2vlad: Global-local sequence alignment for text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval T2vlad: Global-local sequence alignment for text-video retrieval

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.663904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:9ba5ad7fd4454fcefca694737b289474aee09412f650c24c460964d5fca1b675

Observation 2a1936ad-5e4f-43f7-9ece-dedbc41a08e5 · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.609246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:2af90d95e5a7d85da112593fa956ba3f20f79621986725438b287e37df7cd759

Observation 37d9e1c2-d717-4481-b2c1-fdb53fc4fe02 · outbound

This paper cites Unified coarse-to-fine alignment for video-text retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unified coarse-to-fine alignment for video-text retrieval

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.624099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:cac3f840d0d9687e511bcb7993d0842029cfe40fc5443f37fb06869f7ddd9b33

Observation 72e1b749-ee1a-482f-8266-452eb054661d · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.683137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:3e807b350136c9d9b471233d7cf63c14a07486a2a9f04e2239c2fb50adc377db

Observation f599089f-2aec-42db-b05b-dcf40d1976cc · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.681235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:f968fe020f9469b905890ebe33048412e00aa1ed8a3902631336774948777bf5

Observation 8575d487-c113-46d4-b50c-953a9fb3053d · outbound

This paper cites Striking a balance between stability and plasticity for class-incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Striking a balance between stability and plasticity for class-incremental learning

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.622234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:d61c63b3868388114f3bdb80ae259d849e22d5dc65519eba795d774ab53a8cb2

Observation bf5cbcf8-e89d-44a7-b88a-955fb23350a6 · outbound

This paper cites Large scale incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Large scale incremental learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.628196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:94ef08f4baf91cbe5410ccb890c1298f1b259a77cc79b806cdcafe2711b6f8e1

Observation 76bf1e52-9479-4f74-ad57-51d64cfcb98e · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Msr-vtt: A large video description dataset for bridging video and language

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.638408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:6ae59bc61fa8dcd176815eaf7e431781a20207771a128c661c5641a8a2bafb92

Observation 323b25a9-ab67-4122-84d9-172ef450bf81 · outbound

This paper cites Clip-vip: Adapting pre- trained image-text model to video-language alignment.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Clip-vip: Adapting pre- trained image-text model to video-language alignment

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.618052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:3db17cef3c986449ed26d1d6899eea310dd105881e0a3e445987441580ed776e

Observation 3a3e8bb0-5807-46a3-bcf9-5b05ae7d85a1 · outbound

This paper cites Der: Dynamically expandable representation for class incremental learning.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Der: Dynamically expandable representation for class incremental learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.632232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:50f1f4d859b6a869931eaedd7f7d65321b6ac79fe903010138dbb52fc8da01fc

Observation 293d60e1-c208-42ac-b50e-fd430be27fa1 · outbound

This paper cites Low-rank prompt interaction for continual vision-language retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Low-rank prompt interaction for continual vision-language retrieval

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.685321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:74ddc5b83fd9a4632a274e8c6e0598f21a0de9f667e2e5926c8a869ac0014b9b

Observation e20f0040-64d1-4218-bc02-81aef12cc4ae · outbound

This paper cites Dynamic support network for few-shot class incremental learning.IEEE Transactions on Pattern Analysis and Machine Intelligence.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Dynamic support network for few-shot class incremental learning.IEEE Transactions on Pattern Analysis and Machine Intelligence

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.673508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:311055c51e01d69ebb776a76a5b92012d1e8945b11cefb19b78adbfd26bb22ff

Observation e9a3d12a-91c7-4c8e-ab7f-6b50dee6a5ed · outbound

This paper cites Taco: Token-aware cascade contrastive learning for video-text alignment.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Taco: Token-aware cascade contrastive learning for video-text alignment

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.680932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:cdf2c3d568866b49784dd7e2c29ed9502271d8cd043a53933743cd756cc675a6

Observation e30ce547-9178-4dd1-8db6-b74f4d7e132e · outbound

This paper cites Recent advances of multimodal contin- ual learning: A comprehensive survey.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Recent advances of multimodal contin- ual learning: A comprehensive survey

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:40:50.901410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:1dd86e63853cfdbe178bed0770f907e95fcdcf4fa26ce5c33f53b3b0f9c573b7

Observation 254ee880-1289-427b-92b8-a25f65be6876 · outbound

This paper cites Boosting continual learning of vision-language models via mixture-of-experts adapters.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Boosting continual learning of vision-language models via mixture-of-experts adapters

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.659474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:456dd6466d545068b33ff413be2d5b8f64f4db034d28257df75ba6d2fe28edd4

Observation 08d621ff-ad31-4363-b78a-363b35155a17 · outbound

This paper cites A joint sequence fusion model for video question answering and retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval A joint sequence fusion model for video question answering and retrieval

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.602712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:3b80721324d5ff563ed44bc2a537035ad1e0361bc9a4770279581be9451bf65b

Observation 30180881-0e62-4084-8fc1-e23f12b15cf6 · outbound

This paper cites Quantifying and narrowing the unknown: Interactive text-to-video retrieval via uncertainty minimization.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Quantifying and narrowing the unknown: Interactive text-to-video retrieval via uncertainty minimization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.611791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:43a6a0a7938c4b6ee1b1d4a9cd6cf8696b58636423f394e30d7202d63ec1e83b

Observation b71c631a-5f11-4577-89f4-5fa686a4dc94 · outbound

This paper cites Mpt: Multi-grained prompt tuning for 13 text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Mpt: Multi-grained prompt tuning for 13 text-video retrieval

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.655518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:126962fcea4032dc0762a5f88e83c0d8fdee1a5d594f93bd2db0c3d10da39aaf

Observation 01de61a7-4be2-4976-a52d-d3be55779950 · outbound

This paper cites Vqacl: A novel visual question answering continual learning setting.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Vqacl: A novel visual question answering continual learning setting

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.602711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:84cab1ee93b18697be01b209be2fd3a94dfcf4414fd226c77371613e2829d415

Observation ecf872ac-51ad-464d-8f45-838e029c0fbf · outbound

This paper cites Mgsvf: Multi-grained slow versus fast framework for few-shot class-incremental learning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(3):1576–1588.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Mgsvf: Multi-grained slow versus fast framework for few-shot class-incremental learning.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(3):1576–1588

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.690142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:785dbb38144a2c953466ed1920d30ed97c57a5761474781316e09edbe54a865d

Observation c018a695-a164-4cdd-a9e8-4969c9698d9b · outbound

This paper cites Centerclip: Token clustering for efficient text-video retrieval.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Centerclip: Token clustering for efficient text-video retrieval

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.672177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:314ad225b1a3525687508efa157acebf6b803d15ef0de43a032683fb98fbd993

Observation 5282bee4-e656-420a-bff7-aebc91fd4343 · outbound

This paper cites Continual text-to-video retrieval with frame fusion and task-aware routing.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Continual text-to-video retrieval with frame fusion and task-aware routing

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.651540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:3a94dd6adeac45ef2642f37caf1319a6d6c5d3b138aee1a415a3af814cd375b0

Observation 30e8075a-d419-4b5c-8dd1-9e0ea72e66f0 · outbound

This paper cites Preventing zero-shot transfer degradation in continual learning of vision-language models.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Preventing zero-shot transfer degradation in continual learning of vision-language models

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.653601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:304812c365846e7e20cc4ba6cbaeff674b33dfb3b62bb9441eae5f3214f70314

Observation 0e8f7948-cea6-4260-ad51-2af72f224ec2 · outbound

This paper cites Understanding imbalanced semantic segmentation through neural collapse.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Understanding imbalanced semantic segmentation through neural collapse

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.665178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:b58ddaa8a1922c34041abcb6ca8ce978a5980b5fb805c1857d6a9a45835468a6

Observation fc92ad12-59cd-4713-ab6a-5ddf29ecbce1 · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.678753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:0454d1be5ed32bd5eb26e07f26649f5fc7e4e44e56f8cf2d8315db9b2e7a4eec

Observation f406e808-57cf-4780-88ec-ec31b8bf628e · outbound

This paper cites an unresolved cited work.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-16T10:40:51.626212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:2fd3beb6f357e1b1c0e9514ec150b753d4c6f287c8a5e1309eb56250c5886e08

Observation 68c220ba-d509-46e9-b739-cd48f2e79ba2 · outbound

This paper cites Complementarity- aware space learning for video-text retrieval.IEEE Transactions on Circuits and Systems for Video Technology, 33(8):4362–4374.

StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval Complementarity- aware space learning for video-text retrieval.IEEE Transactions on Circuits and Systems for Video Technology, 33(8):4362–4374

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T10:40:51.614174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T10:40:01.902254Z digest=sha256:02535cf877b7f59d8c775063403c6d9c974660ad156eb88815ac9527148feac3

Pith citing papers

No inbound Pith citation observations are available.