Pith. sign in

Paper Citation Record · LEDGER

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval

As of 9 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2510.15470.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.15470 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:27:33.138881Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0c0aa15-cc85-4c71-a286-484f31741ef5 · outbound

This paper cites Overview and current status of remote sensing applications based on unmanned aerial vehicles (uavs),.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Overview and current status of remote sensing applications based on unmanned aerial vehicles (uavs),

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:25.249563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:25.249563Z digest=sha256:9f3c6866e0bceab86eecd7cc5bb263b2bdfa76e766951f1cc3726cc6c0ab27c5

Observation 9a324bb2-db7c-440f-824b-071e0ab3a1d3 · outbound

This paper cites Unmanned aerial systems for photogram- metry and remote sensing: A review - sciencedirect,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Unmanned aerial systems for photogram- metry and remote sensing: A review - sciencedirect,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:25.347496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:25.347496Z digest=sha256:37bebf7ab20ba05a23192d5d012848c7b750b661150e7648f50326a6e5ae6527

Observation 2dc5533e-70c2-4138-9b5c-7f339c05fd3e · outbound

This paper cites Uavs challenge to assess water stress for sustainable agriculture,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Uavs challenge to assess water stress for sustainable agriculture,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:25.479710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:25.479710Z digest=sha256:b29f9b1db75b7f597d111e92d3baddb90247453e80db0ccb39a6d78fa9e9f83f

Observation b8765432-905a-40af-8f8a-1a490c021249 · outbound

This paper cites Sustainable agriculture by increasing nitrogen fertilizer efficiency using low-resolution camera mounted on unmanned aerial vehicles.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Sustainable agriculture by increasing nitrogen fertilizer efficiency using low-resolution camera mounted on unmanned aerial vehicles

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:25.621007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:25.621007Z digest=sha256:d753d7b31fe30f6e50410c719a3315cfa33ebb5cb7072ef4bc8421e58a3725fd

Observation 3e896907-0e14-4f67-8316-2bcf1ac1e3ba · outbound

This paper cites an unresolved cited work.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:25.820139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:25.820139Z digest=sha256:7b33de140b281717b5291e48bd6b67740e052bfbd773b3fc290ed377d3aaa329

Observation 90b5870b-ae9f-4d53-ac0c-6c1efecb6d73 · outbound

This paper cites Visible-thermal UA V tracking: A large-scale benchmark and new baseline,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Visible-thermal UA V tracking: A large-scale benchmark and new baseline,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:25.881936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:25.881936Z digest=sha256:43e0feabc22ced6316b242c2fbcfa4cf0098c59061d98eb30ec50771e1725d9a

Observation 4b1acb93-97e2-4dce-84b4-eaff330f7baa · outbound

This paper cites High-resolution feature pyramid network for small object detection on drone view,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval High-resolution feature pyramid network for small object detection on drone view,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:25.983511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:25.983511Z digest=sha256:12661894bf1d795cfa42f0db2b14348326bf3c66c816c02ff97c19ec980e5257

Observation 43aab4db-a5c4-463e-bdf8-2b815cc304d3 · outbound

This paper cites EarthNets: Empowering AI in Earth Observation.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval EarthNets: Empowering AI in Earth Observation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:26.139458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:26.139458Z digest=sha256:927fd5188f3ab3b46c961158f020b9f60e51553552542c9f9aa936faeff787f8

Observation a8e3df58-a671-41e1-b783-544495bf167e · outbound

This paper cites Sdanet: Semantic- embedded density adaptive network for moving vehicle detection in satellite videos,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Sdanet: Semantic- embedded density adaptive network for moving vehicle detection in satellite videos,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:26.317328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:26.317328Z digest=sha256:d13fb48d49979b4a534fac5216bacab745bc035c0a722fb5c96b2a9dcee6bf20

Observation e5ee1637-38b1-4c68-bf0c-e4ad8138e9e0 · outbound

This paper cites Pareto refocusing for drone-view object detection,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Pareto refocusing for drone-view object detection,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:26.493314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:26.493314Z digest=sha256:9f2a455fcf7b7872b3586a5dce4ee4da7ebd1a95568fa7ba852d5a0a3fb25283

Observation 4436b924-fa26-430b-96de-42e4be59cc48 · outbound

This paper cites Temporal-spatial feature interac- tion network for multi-drone multi-object tracking,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Temporal-spatial feature interac- tion network for multi-drone multi-object tracking,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:26.657760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:26.657760Z digest=sha256:94358fa19abd8a76f61f98df2e979203c25d034dd003de16bc3d044332d5251e

Observation 7c74a6b5-40d5-44cb-b2e5-442fca50f8a0 · outbound

This paper cites Cross-drone transformer network for robust single object tracking,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Cross-drone transformer network for robust single object tracking,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:26.801948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:26.801948Z digest=sha256:dec4b0008843237fd0e066870d2f250419cae3d19fdf0097be891c458372803f

Observation 2ef94245-84c3-4cb3-a8e7-e4fe1b743173 · outbound

This paper cites Transformer- based spatio-temporal unsupervised traffic anomaly detection in aerial videos,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Transformer- based spatio-temporal unsupervised traffic anomaly detection in aerial videos,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:26.921765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:26.921765Z digest=sha256:8a378e997383a090a8cb324bcf1c6dde2ee6bb9df640a8a13b4e80390f467ab9

Observation f9d4e355-cd8d-4900-8728-6c95121940a6 · outbound

This paper cites Visual contextual semantic reasoning for cross-modal drone image-text retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Visual contextual semantic reasoning for cross-modal drone image-text retrieval,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:27.051963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:27.051963Z digest=sha256:b49ae585540e92bee6ede1d41ff4ebbcf73436b145fbe784ff30b3edcbc64071

Observation 9441b470-f639-4d38-9f0a-322d8002046a · outbound

This paper cites Deep saliency smoothing hashing for drone image retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Deep saliency smoothing hashing for drone image retrieval,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:27.245911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:27.245911Z digest=sha256:9f7ead434bdc7bf9191965154d7fce42234cb7ce0fbe4ca28ec807020c08b9cd

Observation 1bb93069-d2d8-40de-92e8-c24abffa9d10 · outbound

This paper cites Uav-human: A large benchmark for human behavior understanding with unmanned aerial vehicles,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Uav-human: A large benchmark for human behavior understanding with unmanned aerial vehicles,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:27.420894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:27.420894Z digest=sha256:cfe22ea8e776889f1423101a5439a24ae123d20ca526465775c350ce406bbf06

Observation 47de33d5-058a-4de1-ad42-b20c3d5df04c · outbound

This paper cites Multi-modal transformer for video retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Multi-modal transformer for video retrieval,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:27.545556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:27.545556Z digest=sha256:f88749128aa533e95dd9f765e8a5438a2ee8db3552fcc4e0f528b1dfd0ad81f6

Observation 6ee3c2bb-ee3c-4523-b65a-165f8d75eb13 · outbound

This paper cites Towards Video Anomaly Retrieval from Video Anomaly Detection: New Benchmarks and Model.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Towards Video Anomaly Retrieval from Video Anomaly Detection: New Benchmarks and Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:27.637380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:27.637380Z digest=sha256:0ff7070b06e3624054c752ac2f1a859d82de8c800d81083a51cc958bb4e49bfa

Observation da72a306-d5d3-4c31-b58a-599f17bb7e8a · outbound

This paper cites Multilevel semantic interaction alignment for video-text cross-modal retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Multilevel semantic interaction alignment for video-text cross-modal retrieval,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:27.694022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:27.694022Z digest=sha256:a8d92ba0b8cb9c6293a96e4a2402af8492047952da214b509fa44ed673bffc38

Observation 60aa7b8b-24f2-4e87-a95e-39345f62f1d4 · outbound

This paper cites Videoclip: Contrastive pre- training for zero-shot video-text understanding,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Videoclip: Contrastive pre- training for zero-shot video-text understanding,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:27.786681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:27.786681Z digest=sha256:e440d5c3bccf819be976387f98348ecff90a7f4fea27071d11977754f032d0c8

Observation 93c71b1d-601f-4056-8acf-fcd1f85765aa · outbound

This paper cites T2VLAD: global-local sequence alignment for text-video retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval T2VLAD: global-local sequence alignment for text-video retrieval,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:27.859844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:27.859844Z digest=sha256:0aa0b3aebba1368ed1d79c5e329c4ec94e8e9b6f229fac96c65536c967e5400c

Observation f26efacf-4cd4-448b-b021-45248dea36ad · outbound

This paper cites CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:27.954552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:27.954552Z digest=sha256:b4220439fec04bd31362fcd912b5b12a25d82fa1f6130f380a8dc062652af75d

Observation d77365a2-c24e-47af-8c97-2dfa1d583892 · outbound

This paper cites Centerclip: Token clustering for efficient text-video retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Centerclip: Token clustering for efficient text-video retrieval,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:28.055281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:28.055281Z digest=sha256:578dc3b6bffc8c16b0e0b458ad6f21cc9288c7edf511b106a792ae19e270af17

Observation b3ed3792-c2ec-42b0-9593-42dd6fb4828e · outbound

This paper cites X-pool: Cross-modal language-video attention for text-video retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval X-pool: Cross-modal language-video attention for text-video retrieval,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:28.138599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:28.138599Z digest=sha256:a5a0629b10d9786b76b84ee142613c0d36f801b49205e698833d61a64c83fcf6

Observation 6c5e0c6e-f8a6-4630-a73f-a9b81851972e · outbound

This paper cites X-CLIP: end-to- end multi-grained contrastive learning for video-text retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval X-CLIP: end-to- end multi-grained contrastive learning for video-text retrieval,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:28.203192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:28.203192Z digest=sha256:d778838eaa966c60677629362a62220aede5674b16f74c8207dd159639f2e604

Observation d5c3f83a-82f4-4d47-98b6-37f333a9780c · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Less is more: Clipbert for video-and-language learning via sparse sampling,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:28.288426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:28.288426Z digest=sha256:0adfc8482baf4119cd6043a84b8c71a11105923c0f5dd6f7cc81271d2f194f6f

Observation c9803a14-b0eb-430d-a80d-598ef9db6fcb · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:28.419476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:28.419476Z digest=sha256:12ed913e97954313c9431294e3bff0976bf1630ae6c777f9af51add62ac30125

Observation 951f7b00-d963-43fd-98f4-5b50d3095524 · outbound

This paper cites Align and tell: Boosting text-video retrieval with local alignment and fine-grained supervision,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Align and tell: Boosting text-video retrieval with local alignment and fine-grained supervision,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:28.518305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:28.518305Z digest=sha256:787b3b61582589343dceca10ef94f47cdabeb8d29e9751e606ea99ee3112f6e7

Observation c2c5b1f2-781e-4669-9972-6dbc07a29a33 · outbound

This paper cites Ts2-net: Token shift and selection transformer for text-video retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Ts2-net: Token shift and selection transformer for text-video retrieval,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:28.605498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:28.605498Z digest=sha256:7f169e2e04dd28ac814f555e7e46c30ee44026874219e942321250bdeea01935

Observation 8f527a65-d759-4557-84a5-aa1fdfbb1fb7 · outbound

This paper cites an unresolved cited work.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:28.703271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:28.703271Z digest=sha256:a7410d7c45fbb724b3af406ea6780eeb75aaf7a6c3aca4d551b511e050892b86

Observation ee628d70-cb9f-4f04-8bbf-bb9e9d5be3f5 · outbound

This paper cites Modeling uncertainty with hedged instance embeddings,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Modeling uncertainty with hedged instance embeddings,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:28.794358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:28.794358Z digest=sha256:4c8874f60870b741cd386bded5481c0b066f39ce0673003a0e38aee35a09a500

Observation f782ef05-09e4-4362-ad43-99079a6f13f7 · outbound

This paper cites Probabilistic face embeddings,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Probabilistic face embeddings,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:28.894178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:28.894178Z digest=sha256:f0826e9d9c081203a5b43ec3d9c57e1ff6314849a252972b4e985c0668ae94f0

Observation c6bab12d-53b6-4083-9e3a-1ab732218cdd · outbound

This paper cites Data uncertainty learning in face recognition,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Data uncertainty learning in face recognition,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:28.981464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:28.981464Z digest=sha256:4e5d49d6fbb9bb4750c99bd1315963ac49f87eb31101807e2ffd0f7b334dc1cf

Observation bb41864a-658f-4ba8-8766-69d9eb7510e5 · outbound

This paper cites View- invariant probabilistic embedding for human pose,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval View- invariant probabilistic embedding for human pose,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:29.110889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:29.110889Z digest=sha256:93f547f5fc403d4517b08a6bcf50fc3faf8cf683eeca51034e31c40f59732e56

Observation 873e0b6f-5d08-406b-bf91-ed57d0c74aef · outbound

This paper cites Probabilistic embeddings for cross-modal retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Probabilistic embeddings for cross-modal retrieval,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:29.293239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:29.293239Z digest=sha256:7d8756b70271f6de89cb8ed6c10f895e5bcf73cab2044d7bf08932aeb4909c0c

Observation fa5eb562-b6de-4f19-8e39-63fa6dc94db5 · outbound

This paper cites UATVR: uncertainty-adaptive text-video retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval UATVR: uncertainty-adaptive text-video retrieval,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:29.457803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:29.457803Z digest=sha256:cc61eb90fde4997381986b0f1439bc27bbd0fff3682b45968de6f01180f58d48

Observation 1b20c5da-334b-4528-b4f4-4aed9c1dc539 · outbound

This paper cites FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval FAME-ViL: Multi-Tasking Vision-Language Model for Heterogeneous Fashion Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:29.616204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:29.616204Z digest=sha256:b2beae14fb0f5015027058249892353623b95677f22213aa7f06297d49f2f516

Observation e4161197-3ad7-4cc3-96fb-2cad8ccf6042 · outbound

This paper cites Position-guided Text Prompt for Vision-Language Pre-training.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Position-guided Text Prompt for Vision-Language Pre-training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:29.766288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:29.766288Z digest=sha256:f9dfbc36cbd3dfa8166995ed98d9ddaa345d84c36a231a2f057061ee28f38636

Observation 55782fe3-41b3-400b-84f8-223174cbb07d · outbound

This paper cites Multi-Modal Representation Learning with Text-Driven Soft Masks.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Multi-Modal Representation Learning with Text-Driven Soft Masks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:29.897347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:29.897347Z digest=sha256:d5bc5b7961d175245e9ae06e472d8b60f26dd8c1fee898ba35ac07b2bd499e9f

Observation a8fa8345-6409-4780-8603-142a90d5912b · outbound

This paper cites GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval GALIP: Generative Adversarial CLIPs for Text-to-Image Synthesis

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:30.021535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:30.021535Z digest=sha256:31b99df3f4c47fcc6d2d84d0359df1759a27cabb7bd9d10c1c6737da15e092ba

Observation b90c1fea-d017-40a4-8b86-c7722c51034b · outbound

This paper cites MAGE: MAsked Generative Encoder to Unify Representation Learning and Image Synthesis.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval MAGE: MAsked Generative Encoder to Unify Representation Learning and Image Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:30.136040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:30.136040Z digest=sha256:5ea661d816ac328a61d0d8c2206adee5e01742547134eb9c0f81ec9e3d313ec8

Observation afae6587-692c-4e98-8b8d-1a7523cd624b · outbound

This paper cites Towards accurate text-based image captioning with content diversity exploration,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Towards accurate text-based image captioning with content diversity exploration,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:30.270733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:30.270733Z digest=sha256:3d97e95d766cc05680d7a4950a5598e5efc123b4eb5373eae9e309db436e2a65

Observation 7bd18602-762d-4dad-9fc4-9408ba13c2ab · outbound

This paper cites Show and tell: A neural image caption generator,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Show and tell: A neural image caption generator,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:30.415898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:30.415898Z digest=sha256:ea90b47b9b48d48de9eb84738e1cf3b86c32d8d84d69a5bdf8ef1de9b34cbccf

Observation 18f86d71-9ce3-48b6-b2e6-24ba3afd280c · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Deep visual-semantic alignments for generating image descriptions,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:30.692071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:30.692071Z digest=sha256:1faafd90f901151a7318647e795b2ce02bd2b15c9ade9a6833ca96a2cd640260

Observation f4c571ea-4557-4e20-be37-90c37aa54e39 · outbound

This paper cites Visual semantic reasoning for image-text matching,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Visual semantic reasoning for image-text matching,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:30.775799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:30.775799Z digest=sha256:083c724a3c356e4874f918d49d68297a10cec6c65d4d3c9f43e5404019ca3091

Observation 872d62fc-edf7-4732-b12e-326a1ae016c6 · outbound

This paper cites Conditional prompt learning for vision-language models,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Conditional prompt learning for vision-language models,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:30.892132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:30.892132Z digest=sha256:a8f156f41823f3b42b661a19274bba25008ad94eb9b190e41a79ff6f57564a2e

Observation 7e3137e0-7a29-4ed9-b12a-65283ea13786 · outbound

This paper cites Bridging video-text retrieval with multiple choice questions,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Bridging video-text retrieval with multiple choice questions,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:31.011422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:31.011422Z digest=sha256:40304807cc3882a5f4f6741ad2619b3fe043131d61434d538025f93754c13076

Observation ba364f4b-3f3f-4c41-ae53-b72a7cf8fd04 · outbound

This paper cites Visual abductive reasoning,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Visual abductive reasoning,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:31.132987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:31.132987Z digest=sha256:c7f86d8b17eed5d5ab85422d2a21804a3175cb7e133e75362a970a519ed446cf

Observation a46d115a-6ce9-4df6-97c3-439448401aeb · outbound

This paper cites Membridge: Video-language pre-training with memory- augmented inter-modality bridge,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Membridge: Video-language pre-training with memory- augmented inter-modality bridge,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:31.248812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:31.248812Z digest=sha256:a8ad53eb3e72f813aa8f818eea05f1dc27b4afd3825197c3b76e872fb558a9a2

Observation 923334b6-6923-4ac6-834c-51880dc0bc68 · outbound

This paper cites A straightforward framework for video retrieval using CLIP,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval A straightforward framework for video retrieval using CLIP,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:31.384176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:31.384176Z digest=sha256:261b62113af3e4c05df9041a3e0e277e771ce26ca968310a4289b02b3cea1c88

Observation 23ff43d9-5430-4dac-8342-1a89a74d838c · outbound

This paper cites Hisa: Hierarchically semantic associating for video temporal grounding,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Hisa: Hierarchically semantic associating for video temporal grounding,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:31.477161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:31.477161Z digest=sha256:a0cd0a15fd4a7e2a3689b426ecf2835650e456fcf99e806303ced683d7172aad

Observation 65e664f4-4038-4b80-b2f6-896b07ce172f · outbound

This paper cites Concept-aware video captioning: Describing videos with effective prior information,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Concept-aware video captioning: Describing videos with effective prior information,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:31.592455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:31.592455Z digest=sha256:18863c0f4977651904a04640a8f903103aa2fff253d75f5db3d1c84051ba504d

Observation 002c68af-697b-402a-9c69-ac0050ef0f40 · outbound

This paper cites Hierarchical representation network with auxiliary tasks for video captioning and video question answering,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Hierarchical representation network with auxiliary tasks for video captioning and video question answering,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:31.709425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:31.709425Z digest=sha256:cea6487a7b6910a7de13ead6aa8f347aa119ceb6848de12b990c623003e1aea1

Observation d73c0d86-0663-4861-82f5-e8790bd60c9a · outbound

This paper cites Cross-attentional spatio-temporal semantic graph networks for video question answering,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Cross-attentional spatio-temporal semantic graph networks for video question answering,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:31.812435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:31.812435Z digest=sha256:a073fc9a16d2ad4b7aaaf6714abbbe7062e268676db722cb1c5785b26323362a

Observation 420785a2-87ff-4ebc-9995-d7fa2e27604e · outbound

This paper cites Adaptive spatio- temporal graph enhanced vision-language representation for video QA,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Adaptive spatio- temporal graph enhanced vision-language representation for video QA,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:31.925922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:31.925922Z digest=sha256:966699ab9a858e95ff40e818819724305909da435833eaf84b20531f79d971fa

Observation 5a1d09eb-1f79-4419-a6dc-d1e6056ac403 · outbound

This paper cites Exploring language hierarchy for video grounding,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Exploring language hierarchy for video grounding,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:32.061616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:32.061616Z digest=sha256:9277e32edcd5985548eeda7a6b5ff8a4a7879f8ac4dbd97270e1dcdaca19b36b

Observation d75c092c-1544-43d0-a39e-b17988689ab7 · outbound

This paper cites Layer Normalization.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Layer Normalization

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:32.099327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:32.099327Z digest=sha256:a34cfca1cb8084dbecd970ebe71ab9203bc7f9c16e1f2f42bd7f6f0f3dfe50ea

Observation 66a2ca70-a2d6-4263-b222-c076b00d3a6c · outbound

This paper cites Attention is all you need,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Attention is all you need,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:32.180909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:32.180909Z digest=sha256:2623258a206ead6a2b943112df0e9c4907fc03bbbeb869d1073cd1e47f847739

Observation a18ce8eb-18a6-408f-8d67-a1f5f22e59de · outbound

This paper cites ERA: A Dataset and Deep Learning Benchmark for Event Recognition in Aerial Videos.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval ERA: A Dataset and Deep Learning Benchmark for Event Recognition in Aerial Videos

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:32.357873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:32.357873Z digest=sha256:4d46366e622eb12a8e3b1a823e24c7a73748f5453ff65cc39d7105b3a633f800

Observation 68194454-ded6-4e32-918d-2dfbe69aee6b · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:32.462871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:32.462871Z digest=sha256:873060242b6bf68acfa957baca3a58ad1373cfc91979b753edb6757f28021d59

Observation 470f4a6b-2dfe-4047-8fdb-ddbf8073072c · outbound

This paper cites Decoupled weight decay regularization,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Decoupled weight decay regularization,

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:32.572859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:32.572859Z digest=sha256:7ea15c60b23e69fb1840fce00118c8b2e84b3cbe39528de42251cb8555f31df6

Observation f834c26c-f347-4705-9ba6-bb0f73c85096 · outbound

This paper cites SGDR: stochastic gradient descent with warm restarts,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval SGDR: stochastic gradient descent with warm restarts,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:32.630351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:32.630351Z digest=sha256:483d5cb087e72f177820602452303806f1f0344aa8ef40697ef7871aaaa63469

Observation 75667931-de7d-47aa-a09f-601be6a4bd2d · outbound

This paper cites Disentangled Representation Learning for Text-Video Retrieval.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Disentangled Representation Learning for Text-Video Retrieval

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:32.734611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:32.734611Z digest=sha256:78f4eb232963c20014f76f118d21821d2de08e29e02cd800a158e0f9cbd278b3

Observation aef1d775-333c-4235-b609-b219180a5b23 · outbound

This paper cites Unified coarse-to-fine alignment for video-text retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Unified coarse-to-fine alignment for video-text retrieval,

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:32.792332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:32.792332Z digest=sha256:3b1316cc7fe63e2834327767f68e6686d969b1ec5570585eb1d90bb41ef2e594

Observation 7db79281-f41a-4fcb-9d0b-321f8338fdf3 · outbound

This paper cites Text is MASS: modeling as stochastic embedding for text- video retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Text is MASS: modeling as stochastic embedding for text- video retrieval,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:32.873537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:32.873537Z digest=sha256:a809b33dcedac62893de477d27e8e31b838b23591544a1d0ae844f3281dd62fb

Observation 6e4f03fb-4a4a-41b7-b328-e08aab1cbf7c · outbound

This paper cites DGL: dynamic global-local prompt tuning for text-video retrieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval DGL: dynamic global-local prompt tuning for text-video retrieval,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:32.954982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:32.954982Z digest=sha256:da4adcaa04f8c6776fcb842e6936910aa88f7817bee4f0fc3487825aba13328b

Observation 0777012d-2953-4d31-9fae-33978252ff21 · outbound

This paper cites Text-video retrieval with global-localsemantic consistent learning,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Text-video retrieval with global-localsemantic consistent learning,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:33.087130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:33.087130Z digest=sha256:294dd0de11ab047951c6a0e3cf35772f819ec5fca022c943ba39112423afb15a

Observation e1a50d70-ce7a-42e3-ae7c-798b52a53934 · outbound

This paper cites Tempme: Video temporal token merging for efficient text-video re- trieval,.

MSAM: Multi-Semantic Adaptive Mining for Cross-Modal Drone Video-Text Retrieval Tempme: Video temporal token merging for efficient text-video re- trieval,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T09:27:33.138881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:27:33.138881Z digest=sha256:6e797b4218f5cd4b8ef04ac8ac2cdef10f1c6184d52370a9bac62dd7dc1ad40f

Pith citing papers

No inbound Pith citation observations are available.