Pith. sign in

Paper Citation Record · LEDGER

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text

As of 8 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 1 inbound Pith citation observation for arXiv:2507.10095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10095 v2

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:47:58.308176Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:48:58.248509Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:48:59.054054Z

Reference resolution

94 of 94 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved63
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85b82364-f8a0-4410-8018-94e29995b356 · outbound

This paper cites ComAlign: Compositional Alignment in Vision-Language Models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ComAlign: Compositional Alignment in Vision-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:47:58.985842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:50.844567Z digest=sha256:f86e00ada3595867ac4d6c8f9021572e8203f8621d29b2b108b0d10d60c831e6

Observation 5eb06bef-2fc3-4f30-9434-d7eee1608165 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:50.931195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:50.931195Z digest=sha256:9ba1e9942d1fb44c1dff4f72b9d89817b58587d4fec626bacc269fb45541e50c

Observation afd3135c-ea32-4aeb-bf47-f167dbe7863a · outbound

This paper cites InternLM2 Technical Report.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.023553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.023553Z digest=sha256:cf92bb00d15e25e9bfc0554423c47c13e93593092bc145a6a238cff590900108

Observation a0393b45-96bd-4029-b95a-7c94ffa108bf · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.123031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.123031Z digest=sha256:749bcde078d8acf25f235d9ece8f3b1794d1ab2b9b68e6218cf9347d5704a436

Observation 70c2c9f4-f23b-4a93-8bef-33bfa0a088ba · outbound

This paper cites PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text PrismLayers: Open Data for High-Quality Multi-Layer Transparent Image Generative Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.190600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.190600Z digest=sha256:0edaba478f6f357b80b28af9375a923af3cba8c845c6a25c639b25716b1c19bf

Observation ed9adfa0-7708-43dd-8730-9ff4da8f548a · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.300206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.300206Z digest=sha256:981af4861a7d37552875687dd9d6ad22134c227a5faa9e4cf879cc2856acdbe3

Observation 28858f04-a530-40a7-a1b3-995cb6c67809 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.375599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.375599Z digest=sha256:9e582ba3b41ec314b020c014875974209daa25663be72734e4b124f43f7b3f5c

Observation 938540eb-a592-4be7-8f5a-1fcd2866819a · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.456033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.456033Z digest=sha256:5dcfac4970a255acdae55a206cca7fc353ee9caebe0dc8bc473f1a5daefdacdc

Observation 7d540b81-e646-4792-a26d-645d458057f3 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Reproducible scal- ing laws for contrastive language-image learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.529823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.529823Z digest=sha256:08ba60a525d1e4e6b1c6fbf8ec5ebd274d9ca1e24b37909cb1bb06949eb52482

Observation ade1a9c4-2ce5-4298-96df-6fdf3230edad · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Instructblip: Towards general-purpose vision- language models with instruction tuning, 2023

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.616122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.616122Z digest=sha256:5d1ff100c5bb5e393a0d0d3f6d357ca5a455e53c27142ea556ed0102cd7b20e2

Observation 4faabd50-9abd-4a31-b138-b46392c1996c · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Imagenet: A large-scale hierarchical image database

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.691092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.691092Z digest=sha256:d6b4d7078bf9ad69425c0566fc2e62298cb84e79e4db4ceb3f4ccef8534140e0

Observation f8e15959-695a-42b4-bd32-9668e0d48eb0 · outbound

This paper cites Maskclip: Masked self-distillation advances contrastive language-image pretraining.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Maskclip: Masked self-distillation advances contrastive language-image pretraining

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.757234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.757234Z digest=sha256:9a2a8ff77c44514f745aba43a00fcf3a056b8de5cbc5ca3338f74fecfbbef0f8

Observation 899646ea-d9f0-42c5-b499-564324a4c96f · outbound

This paper cites Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Go to Zero: Towards Zero-shot Motion Generation with Million-scale Data

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.836489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.836489Z digest=sha256:de0ee3aa203c53ea314f952388ad5905279b1a275a8d96f15a280f75d86d454f

Observation a088e4c0-c9d3-4ddb-ac43-352e9f520b67 · outbound

This paper cites Improving clip training with language rewrites.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Improving clip training with language rewrites

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.895009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.895009Z digest=sha256:19f4ed3a2956414399d1f95e50eae06c2dd25ff014ff4fb1c89c3fc6dfe7fd2f

Observation dea49971-4fc1-4e00-81e8-0d2aecc88b25 · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:51.976190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:51.976190Z digest=sha256:de19f8404a3057ec5edf05889086d494d57d9d9393091ca1a82298df68847e15

Observation c0fb7b0c-7aae-4444-9522-df10b1fbf046 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.096635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.096635Z digest=sha256:47dcfdf7482b6e8545305d48e62f9a9b138cd128fa93fbe4bb2d39b036ce8657

Observation f37346fa-c5ac-44bc-a306-54ab2c5f6b8a · outbound

This paper cites SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.167649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.167649Z digest=sha256:a79ad4eba0db8e552b0621fadad3880044fca61b8a0d009d86656705820955d1

Observation 56d02a32-66c5-4526-996d-8626deded757 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Momentum contrast for unsupervised visual rep- resentation learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.241373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.241373Z digest=sha256:4144b3ba8019215bc63693c1d1aec2003a6da4e7ad7be38f46e9982382363184

Observation 3bd88bcf-110d-4fe7-a18b-43e7ba3f1b94 · outbound

This paper cites Masked autoencoders are scalable vision learners.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Masked autoencoders are scalable vision learners

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.301323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.301323Z digest=sha256:6227a262d1ea7ebd9dd6a78bb818fd3f00ec06e3bd1cdb408a09f35ae44d7667

Observation 0993e728-622f-42ee-8eff-7ab51f8dbd05 · outbound

This paper cites Natural adversarial examples.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Natural adversarial examples

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.382258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.382258Z digest=sha256:f498cc468bed2c3309cecf39eb87e3ef845f22089956169ea27b9ceb96728a73

Observation 175916ac-0f86-4e73-a6b0-5db895b865da · outbound

This paper cites Dynamicid: Zero-shot multi-id image personalization with flexible facial editability.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Dynamicid: Zero-shot multi-id image personalization with flexible facial editability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.468856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.468856Z digest=sha256:599e132d5957b3e56dadcb9a8f92ceb9dd0e248c43fe804f6ad27c8147e71c00

Observation eb2da0f2-6969-471f-a46e-15dc5b7c6095 · outbound

This paper cites Vcoder: Ver- satile vision encoders for multimodal large language models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Vcoder: Ver- satile vision encoders for multimodal large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.522994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.522994Z digest=sha256:e5d16e4f1df61df5ba383db16075d9ba37f0cb8e551eb55e0e1b85722132dcc1

Observation f184e6ab-a845-4049-ba47-7fba77f79468 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scaling up visual and vision-language representation learning with noisy text supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.606217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.606217Z digest=sha256:01d943a9991dcc5be6b338f7015f81aca5fad71e064f7045a2d8ac5466ba85a4

Observation ea6a9366-3a8d-4047-b5c8-45ebcf0c6d39 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.671434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.671434Z digest=sha256:69acb040d45d755fd74fc19eac56817bd21c823bb4871aee007acc597d734c37

Observation f8bbd309-d12f-4dfc-8d96-8c87a5b26b65 · outbound

This paper cites Krizhevsky and G.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Krizhevsky and G

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.755130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.755130Z digest=sha256:6e7442446c8761889e55c73482ebabe1d83fbca681bf56fe67834212ab1f3b6f

Observation d5a3b105-2cab-43fb-9598-5497516bbd4a · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Veclip: Improving clip training via visual-enriched captions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.863818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.863818Z digest=sha256:382a5069f63689292c4f50a26fee1f64798a984d01bf59e61262c0abc9134c4e

Observation 9707efd1-86d7-446f-9c2d-266ad069bae7 · outbound

This paper cites ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text ProxyCLIP: Proxy Attention Improves CLIP for Open-Vocabulary Segmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:52.943918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:52.943918Z digest=sha256:c7a7f1725c03edcf5a8303e5aa687f6d5bf4b43c8a4c78a4941a5a0a71ece853

Observation 1bf5e39d-976d-4250-895c-595d092070c6 · outbound

This paper cites Language-driven Semantic Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Language-driven Semantic Segmentation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.066777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.066777Z digest=sha256:dc93e29b21cf4f3258a2281cb94b5926aca40807f2d13007f721259d86c61450

Observation df3ab08e-13fe-4c2e-9915-ff13b3447768 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.169799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.169799Z digest=sha256:18d1d836308b0bbab9c927a47452875e59236e3aa0f867cd9d29caaf3adac566

Observation 38902c7f-4cd6-4b7a-88ad-60d35621164e · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.279551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.279551Z digest=sha256:d36ce683c4bb445a35d03354b8f9643e50319d3b28308f020835ef77656dac87

Observation 689f7079-8927-4709-ae38-99d07585b202 · outbound

This paper cites Grounded language- image pre-training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Grounded language- image pre-training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.318476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:53.358274Z digest=sha256:df0fef93cd76ec18533592d4f3c7e58215ccb99cabcd911a127f6638cdafc348

Observation 8b5e591e-c7e2-422f-be41-a434ecd1d60a · outbound

This paper cites Scaling language-image pre-training via masking.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scaling language-image pre-training via masking

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.302907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:53.439117Z digest=sha256:697f0406172052ff2b090c402e4be5199bcc1f5507cb9bf86d8dc90ecdd62e88

Observation f75d187f-ca3b-4a50-bb8d-30094aacadc0 · outbound

This paper cites Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.532210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.532210Z digest=sha256:e39b1e31d3907f3c1b2bac6a5d0b237ee2f959826868993ab882edb6a77f2ac9

Observation bd69c4e9-a10d-4040-9c09-556be97cd576 · outbound

This paper cites Improved baselines with visual instruction tuning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Improved baselines with visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.632235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.632235Z digest=sha256:6e8bef6586ce4f4e500296da69f53c539f02059f2517c2f0d35989da56d8507c

Observation ce4e1047-352d-4c57-baa3-160ebe7909a5 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.276811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:53.688289Z digest=sha256:4f1a7b10d6e177abff4206880e4dfdc431b1847e09add27de0efce29a4d9d821

Observation 87c2de30-0b79-43e9-b44a-1e6ae5637304 · outbound

This paper cites Decoupled Weight Decay Regularization.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Decoupled Weight Decay Regularization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.754810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.754810Z digest=sha256:9ad6157221f785fe579195974613b8b27ccb2226e4d4513ac6c1d69a1e6bb520

Observation dfe63ba5-7b7c-4215-a763-0e7e4f8ec317 · outbound

This paper cites Open vocabulary semantic segmentation with patch aligned con- trastive learning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Open vocabulary semantic segmentation with patch aligned con- trastive learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.261207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:53.837321Z digest=sha256:68501439f5439cb30a1632cf8ba0d1708342832486f2d9cac4e0caf45f3c2279

Observation c1f8caa5-333f-45ba-b831-33fdf3fa9b2c · outbound

This paper cites TULIP: Token-length Upgraded CLIP.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text TULIP: Token-length Upgraded CLIP

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:53.923146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:53.923146Z digest=sha256:0921c28445372a7cc807ee3fa095d859142113c49c4b485a6101634104a6b3f8

Observation 029cca05-7d20-4c51-87fd-21dbe3b0d59a · outbound

This paper cites Wonderturbo: Generating interac- tive 3d world in 0.72 seconds, 2025.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Wonderturbo: Generating interac- tive 3d world in 0.72 seconds, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.246042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:54.028385Z digest=sha256:5c6cf9f29ac6e2597d50e1b04dd33fc174639cc68064a2e8fa803b44994d554a

Observation dc291d73-dc63-47b4-ade4-79c070d4be35 · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Im2text: Describing images using 1 million captioned photographs

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.230216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:54.123427Z digest=sha256:35f1596db207bf8d2bbd3de18d7a371b9384e8b402b3c1007d0a5b1171142502

Observation 6f0d734b-6310-4ce9-90ef-036f7181759c · outbound

This paper cites Scalable diffusion models with transformers.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:54.242148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:54.242148Z digest=sha256:b94da56801d995b2cdae331225b26b497a74fdac61e8340a3a0f4475677afbd6

Observation cbb2f623-2e56-4221-8d1b-a3ef970c5bf9 · outbound

This paper cites Bizgen: Advancing article-level visual text rendering for info- graphics generation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Bizgen: Advancing article-level visual text rendering for info- graphics generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.200175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:54.384756Z digest=sha256:d558efcf849e6f17cb9a5e99cc39e4d477791060f1cee1a06c5987aeca726be2

Observation c7914707-cbcc-420d-9889-39d67cf425e5 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Flickr30k entities: Collecting region-to-phrase correspon- dences for richer image-to-sentence models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.185478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:54.522718Z digest=sha256:90fc4faa9fe657a6a1edf377fbdc2054619042a719dca7cc91db8dafb0fe1239

Observation 28f7ae4a-c9b4-4fb9-8716-61328c3dfd68 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Learning transferable visual models from natural language supervi- sion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.171011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:54.645406Z digest=sha256:5e1b4e065c1cc5a4f94e522bc9aa912163bb5c4787d4a12d3772a5da8ad1710e

Observation bab91242-e00e-4b59-90fb-3b242f9cda8a · outbound

This paper cites Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Do imagenet classifiers generalize to im- agenet? In International conference on machine learning , pages 5389–5400

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.155853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:54.728047Z digest=sha256:50671b1b12bd59e7b24e6dcf40b391d3a92d4b9f908423c61f7db573e3be49fb

Observation 727ad07b-659c-488b-9d1e-ec414cfc7994 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text High-resolution image syn- thesis with latent diffusion models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.141043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:54.847562Z digest=sha256:de6bebb898ee9a4b19799100fc1318546cdad519a8d4214eb8f3a692a292b5ea

Observation 0ff8bf5a-e4fa-4a0c-a26d-8c423df35567 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.124884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:54.951984Z digest=sha256:aacc8934a2a45c22f8e1086dbdca37477d7e26aeb4a6326af07aad70513715f0

Observation 245d499d-fe53-48d0-bb61-499da5580eb4 · outbound

This paper cites Umg-clip: A unified multi-granularity vision generalist for open-world understanding.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Umg-clip: A unified multi-granularity vision generalist for open-world understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.110017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:55.094107Z digest=sha256:72c7024bac18e87c7461748c2e064e203ee03e2bacb946920799da78ce026b80

Observation 3d6d531e-fd2d-4f98-a773-afb55277d8dd · outbound

This paper cites Localizing Objects with Self-Supervised Transformers and no Labels.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Localizing Objects with Self-Supervised Transformers and no Labels

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.183779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.183779Z digest=sha256:65a05e7c7878b747b9ba8c0d7d26babe7e86b6f7fce004ba299de96a78fd0b4b

Observation bc494d9b-2f1e-44c9-b2e0-76e80a0af7f5 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.324952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.324952Z digest=sha256:44c6aca039eddebbf80a98003a9f531adf02a9c4be3b4aabb7a08e459a5255ef

Observation 4e448dc6-6881-4b5d-a333-17c921c4b40e · outbound

This paper cites Yfcc100m: The new data in multimedia research.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Yfcc100m: The new data in multimedia research

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.095838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:55.486925Z digest=sha256:6650bf8c38d286e5bc996453f2ba40631d71fbca7472339eb7036d168787a5d3

Observation 65efc709-fe4d-4d36-9117-1bc24eb43b2e · outbound

This paper cites A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.080064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:55.616713Z digest=sha256:65ef91a41846384066a8cb9edeb18817239231ea3b36c1b693006ea6c4f8651d

Observation b3c68827-f803-439e-90b3-2966eeb41bdd · outbound

This paper cites Position-guided text prompt for vision-language pre- training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Position-guided text prompt for vision-language pre- training

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.061108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:55.737623Z digest=sha256:0a83b375304a7ef0f4f7a3aacadd2c82280fb9c87677a48b22a7c5fc9fca40a4

Observation 7e086497-050a-4e97-acf6-b36067dc6e10 · outbound

This paper cites Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Image as a foreign lan- guage: Beit pretraining for vision and vision-language tasks

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.044173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:55.855959Z digest=sha256:d67edf6a94ef3269951bb2397dd599126a1efc9bac740453a45a5ae3386918d0

Observation 0a439027-8353-4b1c-981a-0ec2b4a4eb96 · outbound

This paper cites LoTLIP: Improving Language-Image Pre-training for Long Text Understanding.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text LoTLIP: Improving Language-Image Pre-training for Long Text Understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:55.987654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:55.987654Z digest=sha256:dbead63e8f14caa719c49fa64bba3375acc6abeb07106508c4da7b43dcb20500

Observation e3c8eee6-d1a6-4c8e-8fae-a6925cb17173 · outbound

This paper cites FLAIR: VLM with Fine-grained Language-informed Image Representations.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FLAIR: VLM with Fine-grained Language-informed Image Representations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.051883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.051883Z digest=sha256:6c3c9c19e4820790b3c041d8df36b3843e630e02e82520adbf75a7095a1995c3

Observation ab5e76e4-5cb3-4201-8efd-22c5c53a1fb7 · outbound

This paper cites Demystifying CLIP Data.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Demystifying CLIP Data

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.114498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.114498Z digest=sha256:bfbde29f40a3703f7486a663e45261d0746cc10616607b81e91f03d66bc49fa6

Observation bc37e39e-0277-4cf0-8044-0384f944813d · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.156463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.156463Z digest=sha256:702f10124ea4a0a0ea140ad520591f9fb88c818ce4b5c7e5850ceef31dffa0c0

Observation f1bee241-eed6-4c4f-a81d-9147de762f80 · outbound

This paper cites Fb-diff: Fourier basis-guided diffusion for temporal interpolation of 4d medical imaging,.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Fb-diff: Fourier basis-guided diffusion for temporal interpolation of 4d medical imaging,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.027925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:56.218338Z digest=sha256:352e7e69fa5ab081b250b06f218dcda31531a5bdbbefd60cf0d85c6aaa6e4192

Observation 5ddd13ad-b507-4f8b-a92e-e35421b82a72 · outbound

This paper cites Temporal differential fields for 4d motion modeling via image-to-video synthesis, 2025.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Temporal differential fields for 4d motion modeling via image-to-video synthesis, 2025

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:01.009948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:56.315764Z digest=sha256:938f54123f60092b089366393ef0967b2c02dc05a3888731ec920d4b9517b910

Observation ca00cd65-efd2-4abd-a442-23a2eee6796e · outbound

This paper cites Capsfu- sion: Rethinking image-text data at scale.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Capsfu- sion: Rethinking image-text data at scale

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.992184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:56.396155Z digest=sha256:3f57c4cd3cc0577590f37cfb2e6f328e6d387a6e6c86b43479c2bd6e8c19844a

Observation b7b5ceb2-96e7-4128-9bb4-28c5d0c7d0ec · outbound

This paper cites Sigmoid loss for language image pre-training.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Sigmoid loss for language image pre-training

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.972522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:56.438155Z digest=sha256:f8cb37036d5b2fb3566bf0c5c2593b03363a165decb96e20ef28dee1a463fdb7

Observation d91d90a1-acd1-4d8e-ad0a-ea4de23d1568 · outbound

This paper cites Long-CLIP: Unlocking the Long-Text Capability of CLIP.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Long-CLIP: Unlocking the Long-Text Capability of CLIP

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.501373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.501373Z digest=sha256:3a64916703f0511d92c713bb532cc0e3678c1161e296e9c8b05fe01bde96c622

Observation 9735214e-b1a4-4d5f-ab04-97a552d3d2be · outbound

This paper cites Vision-language models for vision tasks: A survey.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Vision-language models for vision tasks: A survey

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:47:56.618386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:47:56.618386Z digest=sha256:0d9d0d99d8379ea93e7242d51a504dd19f7c1fd48c6dc38a925c2de2d0e954d8

Observation 21800da7-651c-4c08-bc12-c727d5b80f70 · outbound

This paper cites Exploring regional clues in clip for zero-shot semantic seg- mentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Exploring regional clues in clip for zero-shot semantic seg- mentation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.936071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:56.706784Z digest=sha256:10a05673774880e6ea7428852ca236111268611d889e53ebff66d81cedacd9bc

Observation 5421ac72-bac2-442c-a643-c9411cda6979 · outbound

This paper cites Dreamlip: Language- image pre-training with long captions.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Dreamlip: Language- image pre-training with long captions

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.914531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:56.791346Z digest=sha256:084c4b88af0255c7d0782d7698ea601a15d4a928d5851614cfbd5d61e90ba095

Observation 462a7847-6e62-4390-99db-6e3b41da792b · outbound

This paper cites Zegclip: Towards adapting clip for zero-shot semantic segmentation.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Zegclip: Towards adapting clip for zero-shot semantic segmentation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.899086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:56.860708Z digest=sha256:bfb73c74710f572e92be3c4d9d00bac21b7ca3de1aa35b2a3c129f4329af168f

Observation 3f7bd7fd-85a6-43f1-882a-8fb7ae5d882a · outbound

This paper cites During the re-caption process, samples are randomly taken from the following 20 prompts.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text During the re-caption process, samples are randomly taken from the following 20 prompts

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.880236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:56.940555Z digest=sha256:207c280db3498454011a51503232daf3a654d048ca5653d721c7b959a1207b55

Observation b27e8572-25d7-417c-a600-bee0d405dda4 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.864988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.004886Z digest=sha256:8bc3a408cafbea5d90235d9088af1c99f97334cc2e2aa5111fc97352d56bf4be

Observation 2979e6db-0f6c-4255-b2db-473a56218dc8 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.849510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.072569Z digest=sha256:1a003ad964304a9709092787ca490c0f389634863a6cae7f3165f348fcd27cc4

Observation 630a3f3a-739a-479c-b434-737331f99d67 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.834782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.135153Z digest=sha256:d35de736c3bde54f781f81d249659b944df1a40745c8ad57b718e43b6650469e

Observation 292b5fe0-9da0-450f-a811-30b2b2ae16f9 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.815510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.256816Z digest=sha256:1bba3a9730412877605460b8ddb4a0d27dc0a926eec877344b8b771a5b21820d

Observation 4399dc0f-e3cb-4c85-9f93-31825f001afe · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.798569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.327849Z digest=sha256:5baa415015638ebb79e73588cc1f86b55a7c40c5a317cc03af93bf921f3c4081

Observation 5fb6db72-521d-4eea-ba5d-8d15d5f7b6f5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.779688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.390660Z digest=sha256:52a057eba201eca30e4923c6b8be301074fe32d805a059187f5ed99c61b41356

Observation 16ced39f-51fd-4ba1-bc08-1f1ec6ad927e · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.763152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.496477Z digest=sha256:e97c5fe0d443ec2644d37869a9921b1d6ec9608c3c638b3a0427667c782481fd

Observation a5479ea5-f9f6-4c66-8327-1e0cbdf7ad6b · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.736635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.556587Z digest=sha256:84e1eaea1ed49e29a9f6fe2654c172f55338d979f981a6f82fb8ddca123e8e16

Observation 334b0f3c-161e-4fa7-9fd9-51f853f72bb5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.712612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.621546Z digest=sha256:68fb6078c983f47f8e108eceae4ffa235b0d04bc1d19e4d7e8e7e58d2982444a

Observation 2abaedfc-af62-469a-8c19-7697cfaef980 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.691201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.723436Z digest=sha256:334b788c1821bda772387002b7768d3807e4ad42e217e62fbc0579f9520d92fe

Observation 8510c5d8-f4c5-495d-bbdf-be6ca6436edb · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.673643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.764626Z digest=sha256:3de858d62acb1a2c070a42478b7fe43ee3142d00e31258f767d2e6658b9edde1

Observation fd9f8879-87e2-4a54-84b5-59f3da917bf5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.652297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.828019Z digest=sha256:9345b44927c7a6c52cbcc26933df4b4c4ea158c008eb50276a7ef387b3ffc4aa

Observation 22ced992-f3c1-4445-9f28-1447279b278e · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.629924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.895073Z digest=sha256:d7d688d737c5938666e8120f3e4a42c860c918391b73f7b3c34bf42d1ea35663

Observation 9818ae67-df68-4042-a7d5-562a803eab9d · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.600287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:57.960361Z digest=sha256:1a51af57a912be8838789e19aed01223d754d2b859039cf08dd92b6cdd2f6a01

Observation ac6ce2f1-36ab-44f4-808a-f4d137033c05 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.578432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.024087Z digest=sha256:c07acd3e870c9932b81d11bd9f91ed8b147d394d5f2c85134ef9f498e40a2ce2

Observation d9984249-4631-4ce7-813a-f1ef90a0a32f · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.556838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.065219Z digest=sha256:de298b5d620527c2a8a18832767114ac362f0096f0519ea0b6901f577122a779

Observation 9a450b0f-a173-4834-a6fb-61b3e8b55c1d · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.536723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.176510Z digest=sha256:8eb9b050b3bf7949bc7ac0b383b4e3165517efb1c18389c7a9af8055fc251581

Observation dd41036f-1848-4dae-97e2-edb3d32b46d5 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.521128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.244879Z digest=sha256:edf4ed57d550a6e395995d63fb206b93597be3c734307eb99d9d55ca8f7986a5

Observation cc697d6f-033e-4ba5-b1f3-f299809ead8c · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.505202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.274042Z digest=sha256:3f44e7b88c295693a3cc4c4a14c71e085efeb5151ee3d55eeb941b53a22087fa

Observation c6a95066-2bf9-4323-96c3-827cc939b0a7 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:48:00.487803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.279069Z digest=sha256:5ab6ca6afd12e3abfa01ae4175a94954f4fac005c61c142c3d1c4340ec5268e0

Observation 7318b5a6-9db1-4156-a524-9e236f79ca00 · outbound

This paper cites We apply a simple filtering method on captions to reduce repeated words, meaningless sentences, and short results.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text We apply a simple filtering method on captions to reduce repeated words, meaningless sentences, and short results

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:48:00.469236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.284468Z digest=sha256:b972b7fd6206f36221064d0931e31e5dc6791660f6d50f5fa19b12eb6a9897af

Observation dd714861-6b4d-4d9c-b50a-12f83254acce · outbound

This paper cites 5th of October, there’s a significant event highlighted in blue - the launch.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text 5th of October, there’s a significant event highlighted in blue - the launch

Reference 90

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T17:48:00.250320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.289503Z digest=sha256:28c10be0e5a9a7558df2cb0a3f119bbdcfa256d0fc4c5c4c5799b917b84d4ac6

Observation f0a4420a-820d-4fcf-864c-64efaf7c91f2 · outbound

This paper cites Shared Prompts.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Shared Prompts

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.972954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.294865Z digest=sha256:87b87566ec4b50adc3f14eaf77ca1b55dfad89e51a8baecfa3b00f2c5f07a4cb

Observation a91ca14b-6d26-4954-9d66-065ab404ec9f · outbound

This paper cites 8, the regional prompts obtain stronger responses in the corresponding local patches.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text 8, the regional prompts obtain stronger responses in the corresponding local patches

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.679889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.298832Z digest=sha256:cd37a6327a53ef1ec73983e82091d4f8246df59f59f78bb29b2d5aeea7fc86a8

Observation b917b074-82b1-4f5e-9acf-087729fe5ba7 · outbound

This paper cites an unresolved cited work.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:47:59.460818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.302953Z digest=sha256:7608e1d9043afb1c3d9b370a1880f2b4f0f21216967a8026be72d28c5c0cf03c

Observation a663cbe5-11fc-4b8c-be9f-c887bd864bbe · outbound

This paper cites We replace the original text encoder in the stable-diffusion model with that in Long- CLIP [63] or ours.

FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text We replace the original text encoder in the stable-diffusion model with that in Long- CLIP [63] or ours

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:47:59.238484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:47:58.308176Z digest=sha256:ed0779db6bfcdab697707f1374f3baee8c618492ed5ab3b3e22b31cbe44a16d3

Pith citing papers

Observation 660bdf26-2245-4bdf-8e7e-b8d0ad692313 · inbound

Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis cites this paper.

Eigen Neural Network: Unlocking Generalizable Vision with Eigenbasis FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:48:59.127431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T05:48:58.248509Z digest=sha256:2aec2024ed1fc2c73e9d8d0f5e20bf781ac2e1d4f8c27863ff6b8e03d95d32b1