Pith. sign in

Paper Citation Record · LEDGER

Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2211.01335.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2211.01335 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:54:43.528354Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.901557Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 454367fc-f584-4108-83db-4f056ff99b3c · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.193007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:f9f3a2cee286f6c91121d45b9314e666a43f8c5b4ab7ba8bfe768464d84e29c6

Observation 3a5c388e-90f1-4a6b-8046-aee4026ec6ed · inbound

Transmission Line Defect Detection Based on UAV Patrol Images and Vision-language Pretraining cites this paper.

Transmission Line Defect Detection Based on UAV Patrol Images and Vision-language Pretraining Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T18:39:38.888167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:39:38.888167Z digest=sha256:c2c718a55b4eace852202dfadb250bbe107b85e4adbf02eda9a4c703b899ba48

Observation 5ba97f96-d30f-4d48-8a22-a1c68accd075 · inbound

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training cites this paper.

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T18:38:19.021241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:38:19.021241Z digest=sha256:c8aeaecdaafbecf23ff4dce3f52ebc1e6b7d69e5456dd276f4e76642934cad59

Observation 4e6af795-fd8d-4a68-98da-5fce212198a5 · inbound

Text-Video Multi-Grained Integration for Video Moment Montage cites this paper.

Text-Video Multi-Grained Integration for Video Moment Montage Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T17:10:14.553825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:10:14.553825Z digest=sha256:c4e68210d0d6596e3ab57daf99d0825a8686327b04c9275dbbc57254844173ae

Observation b656fe77-a02f-4db8-9d53-252efc70b473 · inbound

From 2D CAD Drawings to 3D Parametric Models: A Vision-Language Approach cites this paper.

From 2D CAD Drawings to 3D Parametric Models: A Vision-Language Approach Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T14:34:00.771229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:34:00.771229Z digest=sha256:43994dd9582221843d64da6036d60951f9e4fb1ceed9ecdd69526d684983a56c

Observation 7e4a0350-2dad-43fd-aa68-10c8dfa1f7db · inbound

Controllable Satellite-to-Street-View Synthesis with Precise Pose Alignment and Zero-Shot Environmental Control cites this paper.

Controllable Satellite-to-Street-View Synthesis with Precise Pose Alignment and Zero-Shot Environmental Control Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T10:18:03.703357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T10:18:03.703357Z digest=sha256:7ccffabf38358b9def09e2caf39e63b9ee5d8bd1ba591a95b01bb52865d957c3

Observation af9a374b-3e77-49e9-8fbb-38a8765a354c · inbound

HarmonyCut: Supporting Creative Chinese Paper-cutting Design with Form and Connotation Harmony cites this paper.

HarmonyCut: Supporting Creative Chinese Paper-cutting Design with Form and Connotation Harmony Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-08T12:10:43.441199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:10:43.441199Z digest=sha256:029a3fbc2063568b46f41e9306f9f582e0dd4f72e90744462dcf2939f7d99edc

Observation 49c3e5ec-8fcb-444f-808a-65b31d513bdd · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 279

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:23.863149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:7b5d35e92625aabd328120cc7a7fc0a81a53d47e91d4d654f5e654bfe008d25b

Observation ea8302fd-67df-4bee-8e33-90af132e764c · inbound

Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark Approach cites this paper.

Multimodal Coreference Resolution for Chinese Social Media Dialogues: Dataset and Benchmark Approach Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T11:54:43.528354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:54:43.528354Z digest=sha256:6a86218c5105a58a21d7d767f27a9a1290f25f5f58b4a1666410a178fb6a65ac

Observation 9013c745-1044-4219-b1c2-e7133c872756 · inbound

VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform cites this paper.

VLM as Policy: Common-Law Content Moderation Framework for Short Video Platform Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-16T11:41:25.763712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:41:25.763712Z digest=sha256:df1d25309c49c6537ff218f2e764ff8cf3c56a3c3029fe459ef2ac219a327c66

Observation 24299011-d4e0-4667-84dc-87a2483404eb · inbound

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding cites this paper.

SeriesBench: A Benchmark for Narrative-Driven Drama Series Understanding Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:09:15.208095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:09:15.208095Z digest=sha256:66a180d0ef18ec65ed8630d0cae939f92606f3dc50c22c1b82261b8c09f708e7

Observation 810a1e1e-88ba-404a-8d59-e2005a2c0637 · inbound

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution cites this paper.

Towards Cross-modal Retrieval in Chinese Cultural Heritage Documents: Dataset and Solution Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:04:46.265412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:04:46.265412Z digest=sha256:d1d47968e00a75d3a244788302a6bc7b8091c9735d24f27c5ccffc8de41cc37a

Observation f1967491-8899-4f2d-85a9-d16e43bfea0e · inbound

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts cites this paper.

Beyond Cropped Regions: New Benchmark and Corresponding Baseline for Chinese Scene Text Retrieval in Diverse Layouts Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:34:01.936414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:34:01.936414Z digest=sha256:39b48a5b8fffd47b99915a5948ee6bc6e048c4a0b71262b5b0aa1e518612ccbc

Observation a473fc8a-db33-45b0-9760-cc662ae9c057 · inbound

FORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets cites this paper.

FORGE: Forming Semantic Identifiers for Generative Retrieval in Industrial Datasets Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T15:17:07.637024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:17:07.637024Z digest=sha256:73677fce2b447c04aadc89b2624fd089c54921aa06021b2328a0b401982c5705

Observation 1be1cad4-722c-4c1d-acf1-a614c9c91e91 · inbound

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model cites this paper.

FG-CLIP 2: A Bilingual Fine-grained Vision-Language Alignment Model Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:18:44.449566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:18:44.449566Z digest=sha256:221048674b0bd8de6bf9062c05586b14d802bbea14b0b39d0b2c288dc5cba688

Observation d2f3e6d3-f9ab-4560-bdf4-bc1cbad7a8e4 · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.245140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T14:08:36.801359Z digest=sha256:56830badf1e7b5af1730df3d7d0653175c05d4c1432d9f4068f1670b1429b993

Observation 9313398e-7389-40e6-9162-02bc5618da1d · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.923852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.923852Z digest=sha256:17227e1638a0dbded71257fd9ec0c145f5916907c14765fe22a20d9bca15f719

Observation aa561e56-aabf-4b11-825c-c256c14093d0 · inbound

Disentangling Fact from Sentiment: A Dynamic Conflict-Consensus Framework for Multimodal Fake News Detection cites this paper.

Disentangling Fact from Sentiment: A Dynamic Conflict-Consensus Framework for Multimodal Fake News Detection Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:01:16.140672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T21:00:09.342056Z digest=sha256:28ad71f0cd2de092ea617ca5eb688f67d9e23d60d4a36e3fe0d4bbe160413301

Observation 36fda9db-7c75-4e95-a259-dc1a30d28b6f · inbound

JARVIS: An Evidence-Grounded Retrieval System for Interpretable Deceptive Reviews Adjudication cites this paper.

JARVIS: An Evidence-Grounded Retrieval System for Interpretable Deceptive Reviews Adjudication Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:31:43.750186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-15T22:31:09.746554Z digest=sha256:83a434fd33886d12c23e095093483641834fd6c4a0506e15d2d21f19c6cd5cde

Observation 14ec2737-114e-4706-9913-3eb083145973 · inbound

Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data cites this paper.

Maximizing Mutual Information Between Prompt and Response Improves LLM Performance With No Additional Data Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T23:48:38.098404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:48:38.098404Z digest=sha256:cacea587b416ba0dd4192c6fff73d4ea59f53520e3690493a732d05fdf15139f

Observation 6002349a-9720-4901-b60c-02f1d9a9b57f · inbound

DRG-Font: Dynamic Reference-Guided Few-shot Font Generation via Contrastive Style-Content Disentanglement cites this paper.

DRG-Font: Dynamic Reference-Guided Few-shot Font Generation via Contrastive Style-Content Disentanglement Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:20:26.321381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T13:13:30.596448Z digest=sha256:0745a5e72139d6d11e5266a58464c594b542a9923752f17f6cecfdd6807bda0a

Observation aaf75cb8-8dac-4480-ae00-d935c580e1fb · inbound

Text-Guided Visual Representation Learning for Robust Multimodal E-Commerce Recommendation cites this paper.

Text-Guided Visual Representation Learning for Robust Multimodal E-Commerce Recommendation Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:12:51.744658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T23:08:33.010725Z digest=sha256:8bf7dbfe7ac0f02448e0acf2102aa6cf08cd470507770e3d576159969cb0d8c7

Observation 92ed66d7-e17b-4c95-90da-fd86a351e428 · inbound

TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval cites this paper.

TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:57:53.206645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T23:54:32.646978Z digest=sha256:06c3c04b5b35ada1f2df283d1ac03b09a1594f2b68e714eba578daea41434ed3

Observation f6f026c4-33d4-4b6c-984c-94d026951ae6 · inbound

MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding cites this paper.

MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:54:45.382242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T14:52:47.828027Z digest=sha256:5913ec0de7693ec39be9fd2ab9ff1611eab46fa6cf57cdf02fcda81f0ecd4035

Observation 556cfd5d-0b4a-4bf6-92c8-72a3b1e2a459 · inbound

UniNote: A Unified Embedding Model for Multimodal Representation and Ranking cites this paper.

UniNote: A Unified Embedding Model for Multimodal Representation and Ranking Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T06:03:08.836179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T05:55:04.367186Z digest=sha256:1834b4f2ea07f253b410eb22c40a4fd60106b702ba4d74a45fa8eea4f6e4d200

Observation b05e1a17-c547-46a2-9f51-c82999cb66b6 · inbound

MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding cites this paper.

MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T22:42:46.338904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T22:40:49.525450Z digest=sha256:69d1cdc36d70d41d4783ef295a4439403f07e65a1545bda779dfb62f67c99c43

Observation 2f3030fa-5072-48f1-a2d1-6265c12946d6 · inbound

Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues cites this paper.

Fine-grained Fragment Retrieval in Multi-modal Long-form Dialogues Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T08:26:48.009848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T06:04:28.939248Z digest=sha256:26c0d3c8dc8a286efc89cbcaf86e96c88390188834baca184aca85246411e7a6

Observation 558e0d6c-e320-4ea9-8f42-4e557a03d566 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 143

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.902902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:fc05986d1c018900e510909fc7b8e727623d162443936c5c29669c3c993faec4

Observation 2888c23e-cc9c-4cee-a36c-e1716aece263 · inbound

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators cites this paper.

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:48.760240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T01:02:20.024793Z digest=sha256:236b427428b5d499f797b2a09afc0ae9a8f620f61fa8f2c49ab7863614bf7abc

Observation 8e187e19-ec37-44f5-819f-2758a2a1621a · inbound

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators cites this paper.

JuZhou 1.0 Technical Report: The First Edge-Native Text-to-Image Foundation Model Trained Entirely on China-Developed AI Accelerators Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T11:46:50.520083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:46:50.520083Z digest=sha256:ac6fbf5170340617e0e458053b3cd5bacab3f3053cd13ce48cbecb455fa7308a

Observation 2e24fe46-9617-4181-b2e5-3317bb8beb13 · inbound

BamiBERT: A New BERT-based Language Model for Vietnamese cites this paper.

BamiBERT: A New BERT-based Language Model for Vietnamese Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.770089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-03T14:29:14.923528Z digest=sha256:a91c09cbb480611904edbf957fd51875a20cda19f0ceb4d8992d00f7ff7120cd

Observation 101c3f87-480f-4fd5-bad1-0d7c46ca2163 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 158

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:20496e2a7d4e28e345e1a3f53a34708d772ab0677eb2141adda6f470686e44cb

Observation 98c1a34a-e216-4b65-a4c7-475005d12a3a · inbound

RecGPT-V3 Technical Report cites this paper.

RecGPT-V3 Technical Report Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T22:57:25.745539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:57:25.745539Z digest=sha256:157fbd45a1138a963723ad8c8a0a1447de283dba8b731c4bcbf77982d806a630

Observation cc146426-2b52-4402-b2f1-41c5f4ee38dc · inbound

CHaystack: Benchmarking Chinese Document Retrieval and VQA cites this paper.

CHaystack: Benchmarking Chinese Document Retrieval and VQA Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T12:43:35.458231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:43:35.458231Z digest=sha256:85c5009eab739c996505f2c4ee4f7f794fc56013e5a458e0caa313161d4c3cdb

Observation 01b51685-1e22-4bc2-9f44-da41d03d2333 · inbound

GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System cites this paper.

GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T11:34:16.740780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T11:34:16.740780Z digest=sha256:2d000efbbe63b6b896912cef21cc29fe9cbb7143d2f2bf1e5d66e66702cda126

Observation 87e9efb2-e9f0-4622-8961-558d1fdb1567 · inbound

Illuminating Visual Identity in Universal Multimodal Embeddings cites this paper.

Illuminating Visual Identity in Universal Multimodal Embeddings Chinese CLIP: Contrastive Vision-Language Pretraining in Chinese

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-04T20:43:17.339889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:43:17.339889Z digest=sha256:45459f4a41e0aa1e495343e26849057e190881784f93f0c7bccdfa3e79c637b4