Pith. sign in

Paper Citation Record · LEDGER

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text

As of 14 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 1 inbound Pith citation observation for arXiv:2507.10095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10095 v2

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:47:58.308176Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:48:58.248509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:48:59.054054Z

Reference resolution

94 of 94 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved63
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85b82364-f8a0-4410-8018-94e29995b356 · outbound

This paper cites ComAlign: Compositional Alignment in Vision-Language Models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ComAlign: Compositional Alignment in Vision-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:47:58.985842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:50.844567Z digest=sha256:692627541029f5ab9d50162a3a13f7a02b541b0182577e28d8660060ab37ce31

Observation 5eb06bef-2fc3-4f30-9434-d7eee1608165 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:50.931195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:50.931195Z digest=sha256:39182202e168b42608c664b11e76f89363de810bdc886b64eca17e10e842d610

Observation afd3135c-ea32-4aeb-bf47-f167dbe7863a · outbound

This paper cites InternLM2 Technical Report.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.023553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.023553Z digest=sha256:7bf722969d390ccdbd7ae2d29b7544615c20402e074479d8672a45d1be046fc7

Observation a0393b45-96bd-4029-b95a-7c94ffa108bf · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.123031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.123031Z digest=sha256:61e55f94c26194334b7bf971bca0a024b556b35709f80bd20e05c0ece630b4ff

Observation 70c2c9f4-f23b-4a93-8bef-33bfa0a088ba · outbound

This paper cites PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.190600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.190600Z digest=sha256:506bca25df65589fd37cee3fd09c0ca47ce35daab01e3d65427991f8672839a1

Observation ed9adfa0-7708-43dd-8730-9ff4da8f548a · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.300206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.300206Z digest=sha256:c922bc990c4a8fd6f20f91320d185d204ba29da779f92b88ef3ef8a24f99ebbc

Observation 28858f04-a530-40a7-a1b3-995cb6c67809 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.375599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.375599Z digest=sha256:82b4a231040375a0b2d38abf534cc14f80d87e33f1d41c60da0b860ee4b974c4

Observation 938540eb-a592-4be7-8f5a-1fcd2866819a · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.456033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.456033Z digest=sha256:29b8f477351e3fe1b22d0afea027ef8e8ec0b719a185e773a06fc3f62708a42c

Observation 7d540b81-e646-4792-a26d-645d458057f3 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Reproducible scal- ing laws for contrastive language-image learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.529823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.529823Z digest=sha256:a8f61a9d29762396c7dd052aab6074723d3e0daa994de49b6252175994770b22

Observation ade1a9c4-2ce5-4298-96df-6fdf3230edad · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.616122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.616122Z digest=sha256:90a1556153d7c8722736ab70589ed6fed03872725ce8e04e22134100a2b549c8

Observation 4faabd50-9abd-4a31-b138-b46392c1996c · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.691092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.691092Z digest=sha256:5999b083986207440d8c092fc86b4bb2a2dc80d244a937ccb2fcfaefb7990e9d

Observation f8e15959-695a-42b4-bd32-9668e0d48eb0 · outbound

This paper cites Maskclip: Masked self-distillation advances contrastive language-image pretraining.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Maskclip: Masked self-distillation advances contrastive language-image pretraining

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.757234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.757234Z digest=sha256:58ae54b14cbde88d21290b3a3129b4285b21f07fc6d21be48fc1543176c8d17f

Observation 899646ea-d9f0-42c5-b499-564324a4c96f · outbound

This paper cites Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.836489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.836489Z digest=sha256:9c8992e8b309a18e38dc0cba2a17bfe549963f3b5cbc561b144437a4e70c9d42

Observation a088e4c0-c9d3-4ddb-ac43-352e9f520b67 · outbound

This paper cites Improving clip training with language rewrites.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Improving clip training with language rewrites

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.895009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.895009Z digest=sha256:2c24a9a491de57ab9923b7309c5bad2dbf02000bc3d4dcdbda6b608ac2c77cbd

Observation dea49971-4fc1-4e00-81e8-0d2aecc88b25 · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.976190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.976190Z digest=sha256:3984c0b97e55890eec9da1f5a273133b7e006ef7dd3aa0771f299010e8e12b57

Observation c0fb7b0c-7aae-4444-9522-df10b1fbf046 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.096635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.096635Z digest=sha256:649c2674067cecca01fbb5aec40df1080682c7f369bb48471f23bb658b11cb56

Observation f37346fa-c5ac-44bc-a306-54ab2c5f6b8a · outbound

This paper cites SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.167649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.167649Z digest=sha256:16c78a1111aa39da19261958ebb66d92c1a1e60bf4c483e333f8c69f8d21330b

Observation 56d02a32-66c5-4526-996d-8626deded757 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Momentum contrast for unsupervised visual rep- resentation learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.241373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.241373Z digest=sha256:f3f6a646f55466707290d821d1ca858c1feaeec5f4f60314825edc00619b4290

Observation 3bd88bcf-110d-4fe7-a18b-43e7ba3f1b94 · outbound

This paper cites Masked autoencoders are scalable vision learners.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Masked autoencoders are scalable vision learners

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.301323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.301323Z digest=sha256:ce0c9d7b6eae55a912d01025e228d0d60277e575a8fe6d91b75a4b83068548e4

Observation 0993e728-622f-42ee-8eff-7ab51f8dbd05 · outbound

This paper cites Natural adversarial examples.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Natural adversarial examples

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.382258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.382258Z digest=sha256:00a6058db63e3df7cc8351699c677e9e6af0a3af0ea9ce889b65f3bf537775e0

Observation 175916ac-0f86-4e73-a6b0-5db895b865da · outbound

This paper cites Dynamicid: Zero-shot multi-id image personalization with flexible facial editability.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Dynamicid: Zero-shot multi-id image personalization with flexible facial editability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.468856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.468856Z digest=sha256:a7402ba63be291ab8de95dbe16f44ce8b60439f6f823cca3d28f52f219564025

Observation eb2da0f2-6969-471f-a46e-15dc5b7c6095 · outbound

This paper cites Vcoder: Ver- satile vision encoders for multimodal large language models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Vcoder: Ver- satile vision encoders for multimodal large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.522994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.522994Z digest=sha256:7f317de4cc6310fa20f5d223296592d2d17557ab827fb28657311a08ea405201

Observation f184e6ab-a845-4049-ba47-7fba77f79468 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scaling up visual and vision-language representation learning with noisy text supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.606217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.606217Z digest=sha256:f517f337de639601e8d7174a9658a65e048944d5961cefe9ee4270224acb0af9

Observation ea6a9366-3a8d-4047-b5c8-45ebcf0c6d39 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.671434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.671434Z digest=sha256:afb87d7a9b8c0b1ab5081bde9a3edc6fbcc83ba6211fa6179a34ca4d49d83c53

Observation f8bbd309-d12f-4dfc-8d96-8c87a5b26b65 · outbound

This paper cites Krizhevsky and G.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Krizhevsky and G

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.755130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.755130Z digest=sha256:6e9746ecfa9fa916d16c6711ceca6b784f98e110219a4f82528ae5e8a25887fc

Observation d5a3b105-2cab-43fb-9598-5497516bbd4a · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Veclip: Improving clip training via visual-enriched captions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.863818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.863818Z digest=sha256:2adb1c23ff6636ba8d53e314a5fcb06f7c92bdb1cae2b430d5b619c4aafc62b4

Observation 9707efd1-86d7-446f-9c2d-266ad069bae7 · outbound

This paper cites ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.943918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.943918Z digest=sha256:454a73aa6f74512a2bd85bfad08059e60888cc6de6757c9e3ab5a66ad1a6eacd

Observation 1bf5e39d-976d-4250-895c-595d092070c6 · outbound

This paper cites Language-driven Semantic Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Language-driven Semantic Segmentation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.066777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.066777Z digest=sha256:30e48ee02de0aa5a47ed3e0d0d8b7cfa1c00810e7bd285fc56c0271fd7685cc4

Observation df3ab08e-13fe-4c2e-9915-ff13b3447768 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.169799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.169799Z digest=sha256:6e74bdc60b3b5def84f99895f6df26d84d4c7169c1d0bbbfa6120a386cc5170d

Observation 38902c7f-4cd6-4b7a-88ad-60d35621164e · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.279551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.279551Z digest=sha256:b12b822855c040cec168c236d6838a761e885c6e99312c25b824433c34aeba01

Observation 689f7079-8927-4709-ae38-99d07585b202 · outbound

This paper cites Grounded language- image pre-training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Grounded language- image pre-training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.318476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:53.358274Z digest=sha256:4d6daad239f6a9fb0a1b1afbcfa7249a30322d1c5697fb85e04f54d54e7c8ae9

Observation 8b5e591e-c7e2-422f-be41-a434ecd1d60a · outbound

This paper cites Scaling language-image pre-training via masking.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scaling language-image pre-training via masking

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.302907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:53.439117Z digest=sha256:ab271d81aebf0deba619d6a56c2b8be755426fee59a22ec421c73528a42f48a8

Observation f75d187f-ca3b-4a50-bb8d-30094aacadc0 · outbound

This paper cites Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.532210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.532210Z digest=sha256:95a319803891fd363cb63f7031d8c9a9f71ff3f6b97038d944e9a9a21bd4c5be

Observation bd69c4e9-a10d-4040-9c09-556be97cd576 · outbound

This paper cites Improved baselines with visual instruction tuning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Improved baselines with visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.632235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.632235Z digest=sha256:2872a4ca9e4cc7639c50278bb18fa59e37a97bb26977d3c7189f736a62259a87

Observation ce4e1047-352d-4c57-baa3-160ebe7909a5 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.276811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:53.688289Z digest=sha256:71987c00c6d5500d0bbe989725c88e53c79de0af1cb3772334334c0fc4f1b384

Observation 87c2de30-0b79-43e9-b44a-1e6ae5637304 · outbound

This paper cites Decoupled Weight Decay Regularization.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Decoupled Weight Decay Regularization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.754810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.754810Z digest=sha256:7ab1c35299d9ef56eb0017ca2e99ed3eacc6ab08fe56df24ccc33855ef546036

Observation dfe63ba5-7b7c-4215-a763-0e7e4f8ec317 · outbound

This paper cites Open vocabulary semantic segmentation with patch aligned con- trastive learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Open vocabulary semantic segmentation with patch aligned con- trastive learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.261207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:53.837321Z digest=sha256:a8d885a7720c56381e2e788dece8ecd72081eeb640fba4ca82ee2cf10af8f76c

Observation c1f8caa5-333f-45ba-b831-33fdf3fa9b2c · outbound

This paper cites TULIP: Token-length Upgraded CLIP.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text TULIP: Token-length Upgraded CLIP

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.923146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.923146Z digest=sha256:95be1801e813e51b4151af21f973209426665de730bb8759fbfe474695a40d27

Observation 029cca05-7d20-4c51-87fd-21dbe3b0d59a · outbound

This paper cites Wonderturbo: Generating interac- tive 3d world in 0.72 seconds, 2025.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Wonderturbo: Generating interac- tive 3d world in 0.72 seconds, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.246042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:54.028385Z digest=sha256:d4a4911cd820cc2ef7ddf286c52d9c2a1b66802e5bd5a398a3279118d632daf5

Observation dc291d73-dc63-47b4-ade4-79c070d4be35 · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Im2text: Describing images using 1 million captioned photographs

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.230216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:54.123427Z digest=sha256:d21c9f4b030d18768bf5a021d06217e93848ad6d7ee822bb35354eaeb10b81e7

Observation 6f0d734b-6310-4ce9-90ef-036f7181759c · outbound

This paper cites Scalable diffusion models with transformers.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:54.242148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:54.242148Z digest=sha256:53acefbbe4b2ab8bdf6e0b8c93dcba5dde22a87ecf1c75094d81048bbba96165

Observation cbb2f623-2e56-4221-8d1b-a3ef970c5bf9 · outbound

This paper cites Bizgen: Advancing article-level visual text rendering for info- graphics generation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Bizgen: Advancing article-level visual text rendering for info- graphics generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.200175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:54.384756Z digest=sha256:8585db2000c05e9c35e493083ab45aeaa02d4a61e4882338b070a4265b613b1d

Observation c7914707-cbcc-420d-9889-39d67cf425e5 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.185478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:54.522718Z digest=sha256:d6a95e20978c2306f206f808e2ac4100c1a5fd754dfbf181155ca68b799e9cf2

Observation 28f7ae4a-c9b4-4fb9-8716-61328c3dfd68 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Learning transferable visual models from natural language supervi- sion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.171011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:54.645406Z digest=sha256:2ec57269cc7faedbe96e86c733e313b4b695f40aeb10cbabdbbe58e840cfd1fc

Observation bab91242-e00e-4b59-90fb-3b242f9cda8a · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.155853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:54.728047Z digest=sha256:6cc8c0692477eeffdae4936e411d733504e1ae2601307332274645c0b49a09d9

Observation 727ad07b-659c-488b-9d1e-ec414cfc7994 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text High-resolution image syn- thesis with latent diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.141043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:54.847562Z digest=sha256:68d8c9ff2a037bb7150d5b477e8b41ec2fe43031337665f3a4df757a91ea5730

Observation 0ff8bf5a-e4fa-4a0c-a26d-8c423df35567 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.124884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:54.951984Z digest=sha256:e93a479d85a6607d414b0b4e0eb94eb50b350e664c9fed606006cb0b05ebae3d

Observation 245d499d-fe53-48d0-bb61-499da5580eb4 · outbound

This paper cites Umg-clip: A unified multi-granularity vision generalist for open-world understanding.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Umg-clip: A unified multi-granularity vision generalist for open-world understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.110017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:55.094107Z digest=sha256:b39126179b325a79908ce5c68a89dec913d8ff0d1d8572deb69e1976ab28fe18

Observation 3d6d531e-fd2d-4f98-a773-afb55277d8dd · outbound

This paper cites Localizing Objects with Self-Supervised Transformers and no Labels.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Localizing Objects with Self-Supervised Transformers and no Labels

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.183779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.183779Z digest=sha256:79ecdda3e934b001d047431bca048b93a03e6884955cf6cddee4b707a5aa0574

Observation bc494d9b-2f1e-44c9-b2e0-76e80a0af7f5 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.324952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.324952Z digest=sha256:fe889b59af902d7285b0dee3198191e165d50842ec9cb70ef373c24306a2a60f

Observation 4e448dc6-6881-4b5d-a333-17c921c4b40e · outbound

This paper cites Yfcc100m: The new data in multimedia research.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Yfcc100m: The new data in multimedia research

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.095838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:55.486925Z digest=sha256:c2a49fe4f57307cf97468df58e725d56307b133bab5c8b55af0cd8dd41f74340

Observation 65efc709-fe4d-4d36-9117-1bc24eb43b2e · outbound

This paper cites A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.080064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:55.616713Z digest=sha256:587a0e5939b570e2fa9d13de40973e42ab92725d48797987e58a25523cbbb462

Observation b3c68827-f803-439e-90b3-2966eeb41bdd · outbound

This paper cites Position-guided text prompt for vision-language pre- training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Position-guided text prompt for vision-language pre- training

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.061108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:55.737623Z digest=sha256:e3f58b07b2ef07aec926fb3f7f8afc63ff472a71e22e8555b67703e1485c4b19

Observation 7e086497-050a-4e97-acf6-b36067dc6e10 · outbound

This paper cites Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.044173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:55.855959Z digest=sha256:dbb59980026dd8fa2b10b7cde168f07b5694af73852026822183c08a124c5806

Observation 0a439027-8353-4b1c-981a-0ec2b4a4eb96 · outbound

This paper cites LoTLIP: Improving Language-Image Pre-training for Long Text Understanding.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text LoTLIP: Improving Language-Image Pre-training for Long Text Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.987654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.987654Z digest=sha256:5b60749daf2770e495702ffcce0a887a426985eabc18f8fb783bd7fd0231def6

Observation e3c8eee6-d1a6-4c8e-8fae-a6925cb17173 · outbound

This paper cites FLAIR: VLM with Fine-grained Language-informed Image Representations.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FLAIR: VLM with Fine-grained Language-informed Image Representations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.051883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.051883Z digest=sha256:0ed11d80574ab3054cc3d7b9c81c3d4f61025599c8472ee212c1230d15731b2b

Observation ab5e76e4-5cb3-4201-8efd-22c5c53a1fb7 · outbound

This paper cites Demystifying CLIP Data.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Demystifying CLIP Data

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.114498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.114498Z digest=sha256:cefc171f3e5908b4087a79dc4bf5e8a13452724f4fc104e8e119abb09c5354c4

Observation bc37e39e-0277-4cf0-8044-0384f944813d · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.156463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.156463Z digest=sha256:55b9c9f2cf1cd2f699c65b5d227c54955f2e6a6436bf454684133bbe1ce26178

Observation f1bee241-eed6-4c4f-a81d-9147de762f80 · outbound

This paper cites Fb-diff: Fourier basis-guided diffusion for temporal interpolation of 4d medical imaging,.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Fb-diff: Fourier basis-guided diffusion for temporal interpolation of 4d medical imaging,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.027925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:56.218338Z digest=sha256:b91fcdd53c8864fb93c6d6b7d80f7017d913f7408d25ca612e6d3c9c3bdebb9f

Observation 5ddd13ad-b507-4f8b-a92e-e35421b82a72 · outbound

This paper cites Temporal differential fields for 4d motion modeling via image-to-video synthesis, 2025.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Temporal differential fields for 4d motion modeling via image-to-video synthesis, 2025

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.009948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:56.315764Z digest=sha256:320a674a073653aa9caa925e2f306748cded88bf22b606386ad37de097825205

Observation ca00cd65-efd2-4abd-a442-23a2eee6796e · outbound

This paper cites Capsfu- sion: Rethinking image-text data at scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Capsfu- sion: Rethinking image-text data at scale

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.992184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:56.396155Z digest=sha256:3a715ddca52506659477fb7289bae745236ba44d77b8bc4257441e864ff09d9e

Observation b7b5ceb2-96e7-4128-9bb4-28c5d0c7d0ec · outbound

This paper cites Sigmoid loss for language image pre-training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Sigmoid loss for language image pre-training

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.972522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:56.438155Z digest=sha256:b73a81e49eac1578c5eb8b93856eba0d8e8ea446902aa40e42fb8324f7480042

Observation d91d90a1-acd1-4d8e-ad0a-ea4de23d1568 · outbound

This paper cites Long-CLIP: Unlocking the Long-Text Capability of CLIP.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.501373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.501373Z digest=sha256:d989d0f57357785327f5731774b21f6864427bb76dcdb28f149ffb39c512b1c6

Observation 9735214e-b1a4-4d5f-ab04-97a552d3d2be · outbound

This paper cites Vision-language models for vision tasks: A survey.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Vision-language models for vision tasks: A survey

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.618386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.618386Z digest=sha256:bd0607296e9a53943cbc713ba7678e09c9d65f4611ec9696991d1eb6c5984519

Observation 21800da7-651c-4c08-bc12-c727d5b80f70 · outbound

This paper cites Exploring regional clues in clip for zero-shot semantic seg- mentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Exploring regional clues in clip for zero-shot semantic seg- mentation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.936071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:56.706784Z digest=sha256:0839c00e26426cf2261cf4ab07398ed083de821f33b283c2ef476fc12f97771c

Observation 5421ac72-bac2-442c-a643-c9411cda6979 · outbound

This paper cites Dreamlip: Language- image pre-training with long captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Dreamlip: Language- image pre-training with long captions

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.914531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:56.791346Z digest=sha256:d0e97c53a223953717cda2f69a4b6bff9f25f4b841a6b4f1e1bffae208cfba20

Observation 462a7847-6e62-4390-99db-6e3b41da792b · outbound

This paper cites Zegclip: Towards adapting clip for zero-shot semantic segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Zegclip: Towards adapting clip for zero-shot semantic segmentation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.899086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:56.860708Z digest=sha256:ea78443c425b5fea2eccffe3f80760c9223162c06628f2d6ec96653f243debf4

Observation 3f7bd7fd-85a6-43f1-882a-8fb7ae5d882a · outbound

This paper cites During the re-caption process, samples are randomly taken from the following 20 prompts.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text During the re-caption process, samples are randomly taken from the following 20 prompts

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.880236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:56.940555Z digest=sha256:20e56a3c6716ec90b1b22f977e5d1590480514fa9b6b245e7cc11496fa4a0312

Observation b27e8572-25d7-417c-a600-bee0d405dda4 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.864988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.004886Z digest=sha256:e541b6be381486944ad1cb18019cfd148f277c1463933a52ea3b5aa25de1afbc

Observation 2979e6db-0f6c-4255-b2db-473a56218dc8 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.849510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.072569Z digest=sha256:a49f08489d5a9dff26fc49c01437877dfd7ce8d817392035ffd52c4d1e93f73d

Observation 630a3f3a-739a-479c-b434-737331f99d67 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.834782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.135153Z digest=sha256:521b3c59749437a993448fa8d3c4653f0e577739fcbeeebbb83268da5679d737

Observation 292b5fe0-9da0-450f-a811-30b2b2ae16f9 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.815510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.256816Z digest=sha256:fe2e77f40ee1165042dfdef6cb0cd3dc536c07dc0a4e1e86b3489b4867274b47

Observation 4399dc0f-e3cb-4c85-9f93-31825f001afe · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.798569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.327849Z digest=sha256:e32afafc29c35345d0a476eb110f60a19295e9078767b5a8012cfe46de7b1126

Observation 5fb6db72-521d-4eea-ba5d-8d15d5f7b6f5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.779688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.390660Z digest=sha256:36b41f1023492c6628e8f17c174b1f02a14c0b4d545ff251d3b6e4ef249110ff

Observation 16ced39f-51fd-4ba1-bc08-1f1ec6ad927e · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.763152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.496477Z digest=sha256:2b50ba5c1da54d2e8ea471684065af85887fb6bb293c2bc28c6a8aca3d28d900

Observation a5479ea5-f9f6-4c66-8327-1e0cbdf7ad6b · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.736635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.556587Z digest=sha256:2440409f2705b1dec50c0649b90a7a908c87377fdc9d86d8192c028fcb708f68

Observation 334b0f3c-161e-4fa7-9fd9-51f853f72bb5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.712612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.621546Z digest=sha256:cdebd0d9fcc131d226c672febc5ef69ac4a33ca95d60a39ca8d274394af99191

Observation 2abaedfc-af62-469a-8c19-7697cfaef980 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.691201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.723436Z digest=sha256:7415ff62b2954fec5d53d02d8222cd5d8631c44c62c4e0c3cb39a584e2ebef4c

Observation 8510c5d8-f4c5-495d-bbdf-be6ca6436edb · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.673643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.764626Z digest=sha256:5ba0deb3d0527bb3a0ec478c059a87db403cd63a6820ee1ab7573827abc3fbd6

Observation fd9f8879-87e2-4a54-84b5-59f3da917bf5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.652297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.828019Z digest=sha256:3088187cdd9a2fb5f3dcb388e77b36abb882d082459329ae3f73157c13781956

Observation 22ced992-f3c1-4445-9f28-1447279b278e · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.629924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.895073Z digest=sha256:a603b0b6efa8aa17c6039758ee04a66c7906c24dd99740180d7eedc0a6d2e066

Observation 9818ae67-df68-4042-a7d5-562a803eab9d · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.600287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:57.960361Z digest=sha256:409a826849edeccac733806dd71623d60a07d37c70f4c3a025e74104defb206f

Observation ac6ce2f1-36ab-44f4-808a-f4d137033c05 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.578432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.024087Z digest=sha256:a80a14666f15517a2b2d95ccb296fb4b97db9df19ae3f311f62595f5bc2c9652

Observation d9984249-4631-4ce7-813a-f1ef90a0a32f · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.556838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.065219Z digest=sha256:7c366451f669be524b69cc47b2dc0d8ba2905b455d1532b6b4a2e1c9cefd3530

Observation 9a450b0f-a173-4834-a6fb-61b3e8b55c1d · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.536723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.176510Z digest=sha256:2d98aaf4c75b7dc0f87a384fe18eb62c008573645dcff2315c918f4021a34e9a

Observation dd41036f-1848-4dae-97e2-edb3d32b46d5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.521128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.244879Z digest=sha256:f5813245e03e60797e2e2c8e53060271d8b2325aba5cb423f2aca7f7d4b9a310

Observation cc697d6f-033e-4ba5-b1f3-f299809ead8c · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.505202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.274042Z digest=sha256:5f7c9155ac0605d51804c6afb6772473c4fc1758731b9cb796368ba26792dfc9

Observation c6a95066-2bf9-4323-96c3-827cc939b0a7 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.487803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.279069Z digest=sha256:2925f9e00ca9ff0e9d1320aba8466086a93b9b1db58606f52fe9c97985f880ea

Observation 7318b5a6-9db1-4156-a524-9e236f79ca00 · outbound

This paper cites We apply a simple filtering method on captions to reduce repeated words, meaningless sentences, and short results.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text We apply a simple filtering method on captions to reduce repeated words, meaningless sentences, and short results

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.469236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.284468Z digest=sha256:d178ecc95d3dcb2d11017d516ff310016fe1303a2c254ec8ce23acf81ac49361

Observation dd714861-6b4d-4d9c-b50a-12f83254acce · outbound

This paper cites 5th of October, there’s a significant event highlighted in blue - the launch.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text 5th of October, there’s a significant event highlighted in blue - the launch

Reference 90

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T17:48:00.250320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.289503Z digest=sha256:575a94e66dfd24d277dab68a385f1456fb94b2077ef8f002ec63c394d6614943

Observation f0a4420a-820d-4fcf-864c-64efaf7c91f2 · outbound

This paper cites Shared Prompts.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Shared Prompts

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.972954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.294865Z digest=sha256:f42aa80dd85dabdfee6bfde06f765314b8b0fa4552503e1f23111cc9d2a3bd9e

Observation a91ca14b-6d26-4954-9d66-065ab404ec9f · outbound

This paper cites 8, the regional prompts obtain stronger responses in the corresponding local patches.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text 8, the regional prompts obtain stronger responses in the corresponding local patches

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.679889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.298832Z digest=sha256:72c6b9a2eb408726ee33dba7fc6453d38a658512100a2a3a2132bc5ecc717ddb

Observation b917b074-82b1-4f5e-9acf-087729fe5ba7 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:47:59.460818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.302953Z digest=sha256:bd6a652ceed9d16638d1407e04468fbd4fa839426883c9aa829c4bc9ffdc939e

Observation a663cbe5-11fc-4b8c-be9f-c887bd864bbe · outbound

This paper cites We replace the original text encoder in the stable-diffusion model with that in Long- CLIP [63] or ours.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text We replace the original text encoder in the stable-diffusion model with that in Long- CLIP [63] or ours

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.238484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:47:58.308176Z digest=sha256:8fbaabe47cf712002bfc23fd06d0240190ae6a6063f92cc72ec7b8558e685e09

Pith citing papers

Observation 660bdf26-2245-4bdf-8e7e-b8d0ad692313 · inbound

Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis cites this paper.

Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:48:59.127431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T05:48:58.248509Z digest=sha256:221618b3f6d350a8984e0084c3847015397aa45d8d1856ddc365d3a91d3f934f