Pith. sign in

Paper Citation Record · LEDGER

Visual Semantic Description Generation with MLLMs for Image-Text Matching

As of 19 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 0 inbound Pith citation observations for arXiv:2507.08590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08590 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:19:58.847247Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4cc6cc2-5ebc-46f6-9f6f-bdcaf9a95f63 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Faster r-cnn: Towards real-time object detection with region proposal networks,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.439282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.647447Z digest=sha256:b9800154c8bc8573afb43095825b030c3e4f262bb9ea246ac3b78314f70c0ac6

Observation 820545a0-8f94-4934-b661-03f81fcd31c3 · outbound

This paper cites Stacked cross attention for image-text matching,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Stacked cross attention for image-text matching,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.421214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.652953Z digest=sha256:25f196b3cbc1cc2e94ca25c907259e10a97e966fa88dade1614bc837081e2912

Observation cdfd30ec-1d79-4a7c-be7b-f4ca4a409cfe · outbound

This paper cites Image- text embedding learning via visual and textual semantic reasoning,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Image- text embedding learning via visual and textual semantic reasoning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.397910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.657786Z digest=sha256:d28f2ab9d70df6a9f7ec4cfd1eb206f2d2f83aaf8bfa32830fb969fff14f1920

Observation 514a9348-726f-42c5-b9a6-7cd810702ce2 · outbound

This paper cites At- tentive mask clip,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching At- tentive mask clip,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.367248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.663498Z digest=sha256:3e11b8985da893e441aa9f3c113242d5e1b2856ca37afad449ec83afa505893c

Observation c5fe10b2-5c13-4113-89f0-2502927467bb · outbound

This paper cites Composing object relations and attributes for image-text matching,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Composing object relations and attributes for image-text matching,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.352015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.669323Z digest=sha256:6c0f4e8b988c1eac7a7d07110a2a92ce37547fdd0cd811f04ed3ab2255764eba

Observation 8d94c9a9-67e6-475f-adf3-a00f85950f4e · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning transferable visual models from natural language supervi- sion,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.336427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.675803Z digest=sha256:175ec2536da3a2ebbbe572a4b93fed405cfbff07be0906ff1060ca076b653d85

Observation 8ca9b18e-f722-4b92-b8a7-a6d85806f3d2 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Sigmoid loss for language image pre-training,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.319859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.683871Z digest=sha256:b85e0ccd9d40ccfe018232f6d236e1ad61ba63dc9f72a8a32fe30796a2442fc3

Observation 2f68182e-5c82-4a45-bea4-6ba588a05ed2 · outbound

This paper cites Regionclip: Region-based language-image pretraining,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Regionclip: Region-based language-image pretraining,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.292337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.693360Z digest=sha256:efa6c765cf9d9b5a4b0a6bdb6446f70f1a422b5104c96ecbf21741fa43ca055c

Observation 30654223-7950-4359-9836-f79e4b7a9526 · outbound

This paper cites Sclip: Rethinking self-attention for dense vision-language inference,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Sclip: Rethinking self-attention for dense vision-language inference,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.267034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.703590Z digest=sha256:ddc91fb55e5ad350cd5cdec10ef3c3b792ff295284958d5728441c899dd510ea

Observation 9bccf390-e917-45ae-b837-fc8e5ac865aa · outbound

This paper cites Improving clip training with language rewrites,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Improving clip training with language rewrites,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.250030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.715883Z digest=sha256:cc7bb3d4b5e0d75b39221eff4e1d27aff6ae75d712d461e7922ca4f7a99446c7

Observation f1acdda5-029f-4367-ad21-f8a58209a01a · outbound

This paper cites Improving multimodal datasets with image captioning,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Improving multimodal datasets with image captioning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.234196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.728908Z digest=sha256:2776920eedddc30094afef0d9a9bd6b4dd299b19efd5f34352845c1d0cb6183b

Observation cbb654d8-1fa7-4e50-b034-6b2ba87c5c8a · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.218150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.735122Z digest=sha256:77acdc484f191af341188b9a452cf75d9e09b76eb363369b1a85c8a279498569

Observation 0ee958d2-bfaf-42c2-be02-df1ebbc15013 · outbound

This paper cites MLLMs-Augmented Visual-Language Representation Learning.

Visual Semantic Description Generation with MLLMs for Image-Text Matching MLLMs-Augmented Visual-Language Representation Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.740424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.740424Z digest=sha256:17d0615aa51ffd671cf3f7bdf39bb97dc8bb275b88fff681e0cb049c2d94a6cf

Observation 6af04eb6-4472-4714-bb0f-7cc3062c3059 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Visual Semantic Description Generation with MLLMs for Image-Text Matching MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.748432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.748432Z digest=sha256:7601a7ee99cfbaa9397420555ef3f90f2f3cc9c852aca27994dbf4ebc05b0047

Observation 0033427b-733e-4acb-b098-13d0f736b924 · outbound

This paper cites C-Pack: Packed Resources For General Chinese Embeddings.

Visual Semantic Description Generation with MLLMs for Image-Text Matching C-Pack: Packed Resources For General Chinese Embeddings

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.753670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.753670Z digest=sha256:d2a2de5f85aabf5755b37fa6f3b7c6d9e4a01953766781154fbb24199ebbeb30

Observation 08f96dca-8f4c-404a-be87-e9c1a9a24f75 · outbound

This paper cites Fine-grained image- text matching by cross-modal hard aligning network,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Fine-grained image- text matching by cross-modal hard aligning network,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.201309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.761302Z digest=sha256:8aa71961943479e22c4cbc0b214c9f638f7a8713287fd101c33af6a228c82d9e

Observation 50e81f65-4615-40ad-b95f-ed89de7a1fc4 · outbound

This paper cites Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.183805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.768032Z digest=sha256:0d76bb2dabb1805df7845f2b17c28a904b9bfe73f8d1a0736e91b2cafe8dcf03

Observation fcd0442a-4ab7-452f-a98d-8230c52abf65 · outbound

This paper cites Learning the best pooling strategy for visual semantic embedding,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning the best pooling strategy for visual semantic embedding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.166182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.772689Z digest=sha256:44b1a8f21365ad2994fb809dfd92f13ce8dedea6694267969b62c32843d61bc6

Observation e001793f-3add-4bdf-9b6a-1e4239ba249c · outbound

This paper cites Learning semantic relationship among instances for image-text matching,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Learning semantic relationship among instances for image-text matching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.145425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.777092Z digest=sha256:a29bb8a50546b76e7b9fbdc1eb924bd9eeb36a32c5bc05d3be3482b6a3beaea7

Observation 26a67493-2443-427f-b1e0-30aa4136f551 · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Sinkhorn distances: Lightspeed computation of optimal transport,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.124932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.797389Z digest=sha256:75b4a5476485ce8c84808e8eda5e8171fde72f902074f4a35ae3160779602286

Observation f2d0d17a-2ff4-4404-9b24-cbaf2c7c7a03 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Deep visual-semantic alignments for generating image descriptions,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.105418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.805931Z digest=sha256:a8408da0a0376f54b55e80da388af89f6401eb7efb0c9b2db792e1f6859a3774

Observation 593e71a4-bd87-4236-a740-ca1389172458 · outbound

This paper cites N24News: A New Dataset for Multimodal News Classification.

Visual Semantic Description Generation with MLLMs for Image-Text Matching N24News: A New Dataset for Multimodal News Classification

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:19:58.811615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:19:58.811615Z digest=sha256:9b8304660444ed2f12497cef98f1cadd48ffcc6493ee67d0538c99434e84fe05

Observation e83deb29-8974-4017-8f07-bfd8cee1eefa · outbound

This paper cites Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Exploring a fine-grained multiscale method for cross-modal remote sensing image retrieval,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.084295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.817342Z digest=sha256:9b93bfbcc6d5932efa41ba0ccb96bab3962d84c8d33fd411ea12b74154047c98

Observation b76fc5ee-b355-4008-8d7e-2ea5b2803cc3 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Florence-2: Advancing a unified representation for a variety of vision tasks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.065552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.825844Z digest=sha256:1758d5b6758e0f98615d242e2a3c623423ca0dc6c4fec2c22e7443141c05d7dc

Observation 337190bb-dc06-4ccb-9f17-3e968b6363a9 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distillation,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Align before fuse: Vision and language representation learning with momentum distillation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.045119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.832266Z digest=sha256:b1a4a4e746f172e7a2dd3f313e3dd72713c7a96cb33da9e393804f9af7161d22

Observation 8ab328d0-2d7d-4a6b-bbe6-4b9ed6e2c7ca · outbound

This paper cites Remote sensing cross-modal text-image retrieval based on global and local information,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Remote sensing cross-modal text-image retrieval based on global and local information,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.019781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.842725Z digest=sha256:faddff9332e5a412d2aa724abd669cb4892c445498bdeca225ec67eb3aea96bc

Observation 5d951393-c6d2-46c2-b4a3-189abee9d957 · outbound

This paper cites Hypersphere-based remote sensing cross-modal text-image retrieval via curriculum learning,.

Visual Semantic Description Generation with MLLMs for Image-Text Matching Hypersphere-based remote sensing cross-modal text-image retrieval via curriculum learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:59.000703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:19:58.847247Z digest=sha256:b8c3d270688ea0a20ea59a496e5007409964bc0138516275e2e6a9e05160e9dc

Pith citing papers

No inbound Pith citation observations are available.