Pith. sign in

Paper Citation Record · LEDGER

Douyin Multimodal Embedding Model Technical Report

As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2608.02148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02148 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:45:31.800765Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2e5e9e8-61c0-43a6-98d1-7cc14b3bdca6 · outbound

This paper cites MM-BRIGHT: A multi-task multi- modal benchmark for reasoning-intensive retrieval.

Douyin Multimodal Embedding Model Technical Report MM-BRIGHT: A multi-task multi- modal benchmark for reasoning-intensive retrieval

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:26.929476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:26.929476Z digest=sha256:760761a5101ce67b1f8542a42c7bbfb4fff39b7584523edd79d608922f48ab62

Observation fc3fb8a4-21ac-412c-a3cf-6d4da7f5ac17 · outbound

This paper cites Qwen3-VL Technical Report.

Douyin Multimodal Embedding Model Technical Report Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:26.985437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:26.985437Z digest=sha256:28c82c0a087974bb32a93a2a3af56b46773184954539be3f3240323cbacd6d62

Observation 96df572d-21e8-47e7-bc54-94212b28b887 · outbound

This paper cites Qwen2.5-VL Technical Report.

Douyin Multimodal Embedding Model Technical Report Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:27.072738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:27.072738Z digest=sha256:24b64ad80b3ea6d87c34a2be9cab818a39d1e85f12ba85045a093cac2149a2f4

Observation 0b4c2db2-9c79-4ee4-9940-d95b6c443b77 · outbound

This paper cites MS MARCO: A Human Generated MAchine Reading COmprehension Dataset.

Douyin Multimodal Embedding Model Technical Report MS MARCO: A Human Generated MAchine Reading COmprehension Dataset

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:27.157304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:27.157304Z digest=sha256:9454db9ab9c8f0355fea2425c287c87a2e79010bb411cf6abba78c5c4d786c15

Observation 85b14d83-a5c7-42a3-9a20-f78bf9a558f8 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Douyin Multimodal Embedding Model Technical Report Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:27.239176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:27.239176Z digest=sha256:48296c5c75dfa285e8fc64f2400a986ea391502d3bb696eaddba59b53426797b

Observation 11f79045-c8b6-4206-abe6-85a6fec169ce · outbound

This paper cites mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data.

Douyin Multimodal Embedding Model Technical Report mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:27.401932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:27.401932Z digest=sha256:e0ac900a20fdb475f2e8b391812147ee41a01a9a11623fdc32f21921630ae578

Observation 0af506fd-4c66-4879-8ffe-f2479e6095b2 · outbound

This paper cites Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning.

Douyin Multimodal Embedding Model Technical Report Reconstructing Content with Collaborative Attention for Universal Multimodal Representation Learning

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-04T13:48:46.448917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T13:45:27.532154Z digest=sha256:b9bba542a6c7ec1b857eccd6b31f96b576cbf684435119cd69588b07a854f154

Observation a86328f1-2546-442f-a1ca-7ba0a6cb5891 · outbound

This paper cites Pailitao-vl: Unified embedding and reranker for real-time multi-modal industrial search.

Douyin Multimodal Embedding Model Technical Report Pailitao-vl: Unified embedding and reranker for real-time multi-modal industrial search

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:27.666930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:27.666930Z digest=sha256:1d8d80e512626534b6eb497e410c5d228de72029cf2de251ec08cde3ef4df3b1

Observation 849f28f3-18a2-41fb-8b6c-0cb0122d1ba1 · outbound

This paper cites Think then embed: Generative context improves multimodal embedding.

Douyin Multimodal Embedding Model Technical Report Think then embed: Generative context improves multimodal embedding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:27.785164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:27.785164Z digest=sha256:e87609fbfdebdd83b57a8be14a9d9647c8d741beca30c463e39f04df76032c47

Observation f2cb43c9-2299-45cb-85f2-0f2d38b52e32 · outbound

This paper cites Reason to contrast: A cascaded multimodal retrieval frame- work.

Douyin Multimodal Embedding Model Technical Report Reason to contrast: A cascaded multimodal retrieval frame- work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:27.908771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:27.908771Z digest=sha256:d1de71476c8142e97780634c2b34311c729b57d7c824a8d38e80ffa62c84d659

Observation e36ef377-1a59-40d3-9b63-21bd4ca3a351 · outbound

This paper cites DeepSeek-V3 Technical Report.

Douyin Multimodal Embedding Model Technical Report DeepSeek-V3 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:28.054257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:28.054257Z digest=sha256:95a7c69d7410ec32c2a5d4f369de2bb1f9c3693b9dadd18db8559e5e25659d6e

Observation 5e1436e8-7f6f-4dc9-95a0-5ec1c7967912 · outbound

This paper cites BERT: pre-training of deep bidirectional transformers for language understanding.

Douyin Multimodal Embedding Model Technical Report BERT: pre-training of deep bidirectional transformers for language understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:28.166730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:28.166730Z digest=sha256:0f9b301d0a86707d11ccd9bf91808b9d5128c52d13fc674b94521cf8ba152032

Observation 438feb2a-a326-4c9a-bc53-e245f79a0747 · outbound

This paper cites Colpali: Ef- ficient document retrieval with vision language models.

Douyin Multimodal Embedding Model Technical Report Colpali: Ef- ficient document retrieval with vision language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:28.326972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:28.326972Z digest=sha256:bc22700e0b7b0536fc67d03eb8a626d00408ce5b6e760d564bc199ff7adecde4

Observation d03f02ab-4723-40e2-9a9d-c9f032bd0d67 · outbound

This paper cites Moon embedding: Multimodal representation learning for e-commerce search advertising.

Douyin Multimodal Embedding Model Technical Report Moon embedding: Multimodal representation learning for e-commerce search advertising

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:28.380623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:28.380623Z digest=sha256:940c8575db41e6d93a45b5e403c120a7322a34a7104856b45ecda93d8b62e8e5

Observation 9ad0eda1-7b2d-4bbb-a92c-55ac06f81fe2 · outbound

This paper cites Better & faster large lan- guage models via multi-token prediction.

Douyin Multimodal Embedding Model Technical Report Better & faster large lan- guage models via multi-token prediction

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:28.495807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:28.495807Z digest=sha256:d157a291444747a1d500eb3d0ed028a46652e800fc9aa602d6009d065b9bfef2

Observation dab86b88-26af-422d-8886-fb2e9e26e58a · outbound

This paper cites Breaking the modality barrier: Universal embedding learning with multimodal llms.

Douyin Multimodal Embedding Model Technical Report Breaking the modality barrier: Universal embedding learning with multimodal llms

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:28.607678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:28.607678Z digest=sha256:1293898a6dc451d7bae17ab3939d2eaf2c22afca9bd5c7b3abcfb829600def07

Observation a829538a-267a-4c6e-b0b1-6ec0ae92535b · outbound

This paper cites TRACE: task-adaptive reasoning and representation learning for universal multimodal retrieval.

Douyin Multimodal Embedding Model Technical Report TRACE: task-adaptive reasoning and representation learning for universal multimodal retrieval

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-08-04T13:48:45.555058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-04T13:45:28.732826Z digest=sha256:4cb382e8bcb9027feef1ddabe7f3087cba1238276f4a3d628b89bf5e160dc226

Observation 1521658b-b409-4f48-a946-67df0f620d25 · outbound

This paper cites Girshick.

Douyin Multimodal Embedding Model Technical Report Girshick

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:28.881315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:28.881315Z digest=sha256:745cfb38917b5c06606f0d3f86f8287570704aff2e9f3e93bc4cfca48a14cf5f

Observation 3327ab27-7382-49e3-8de4-8e88857010e2 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Douyin Multimodal Embedding Model Technical Report Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:28.935163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:28.935163Z digest=sha256:0b73749403d9659705a824bedeb81e07ac36f2830604ab8d4495f5f1bfcde124

Observation 184c8b13-5186-4910-8241-39e4f4bbe325 · outbound

This paper cites Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig.

Douyin Multimodal Embedding Model Technical Report Le, Yun-Hsuan Sung, Zhen Li, and Tom Duerig

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.001164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.001164Z digest=sha256:b46333557f60b909666af3cd0470258b8a3812d68ce836dc5518acd7fe680876

Observation 33170702-71d0-4368-8c51-87e3a6f11aa5 · outbound

This paper cites Rzenembed: Towards compre- hensive multimodal retrieval.

Douyin Multimodal Embedding Model Technical Report Rzenembed: Towards compre- hensive multimodal retrieval

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.044376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.044376Z digest=sha256:d3c82aee5b73608087cf99fc0699fa6da9fcb8c76fdd82cee0c6e2c19b141ad1

Observation 230d272f-3ada-4480-8aff-8c0fb027a15b · outbound

This paper cites Embed-rl: Rein- forcement learning for reasoning-driven multimodal embeddings.

Douyin Multimodal Embedding Model Technical Report Embed-rl: Rein- forcement learning for reasoning-driven multimodal embeddings

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.110343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.110343Z digest=sha256:0d251ccd42362b4f4105aebc42c1d62633687871b130579e29bf62dc7847b147

Observation 4531ad9b-0111-416c-8436-b85193d4f8ce · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Douyin Multimodal Embedding Model Technical Report E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.154500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.154500Z digest=sha256:1a992154e7b6d759440c3a6a5a16d7ea79343bf0e3ddded9eff5417e4f0bd0c3

Observation 58fd9dc8-9f37-4cfb-9583-827b019baf34 · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

Douyin Multimodal Embedding Model Technical Report VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.224833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.224833Z digest=sha256:286be530f4c01e5075288015378f73ea44a2a366c95cc54b77938132515520c6

Observation 1ac0dddb-5260-4b3a-8beb-15ad190fdeef · outbound

This paper cites Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension.

Douyin Multimodal Embedding Model Technical Report Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.298541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.298541Z digest=sha256:00277dfadab36ac869ddc8f596ebbf94f41df75a9a4e0381b859b82ac327a520

Observation f701cd97-d583-4253-acd9-155400097266 · outbound

This paper cites Natural questions: a benchmark for question answering research.

Douyin Multimodal Embedding Model Technical Report Natural questions: a benchmark for question answering research

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.346453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.346453Z digest=sha256:869101d8c38236c7be34fde0634f45ce73d1dc2e6bf5bd6278d27b13b7933319

Observation 0f6d230e-b47d-4e54-8f55-ff737c995766 · outbound

This paper cites UME-R1: exploring reasoning-driven generative multi- modal embeddings.

Douyin Multimodal Embedding Model Technical Report UME-R1: exploring reasoning-driven generative multi- modal embeddings

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.419912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.419912Z digest=sha256:f4a0dd35d145a666868a5f2f7a15bade70dac274b600c3b968c55ed8ecef6380

Observation c94bdf20-d072-43e0-a85f-9b3628f01f0d · outbound

This paper cites an unresolved cited work.

Douyin Multimodal Embedding Model Technical Report Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.485770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.485770Z digest=sha256:901bb841cff9cb99c7c49a56609b099f1d0a89f81591ac23b0b529b5b0703f56

Observation 7e09e685-1e6a-417f-945c-069531781e74 · outbound

This paper cites an unresolved cited work.

Douyin Multimodal Embedding Model Technical Report Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.563152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.563152Z digest=sha256:d18d7b7cc37436243f10de8b12bcb8876a909b63896ee1430370ac671214bec8

Observation 08329eb3-bc75-4bf2-a7bf-4420ed2763a3 · outbound

This paper cites Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking.

Douyin Multimodal Embedding Model Technical Report Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.614269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.614269Z digest=sha256:c7d1e711536bce98fc3c84f49300f4b9a8cab62fca06c139ba85a4a6cad55604

Observation c6a8e831-3d0a-477b-8325-341ecf6e4dbf · outbound

This paper cites U-marvel: Unveiling key factors for universal multimodal retrieval via embedding learning with mllms.

Douyin Multimodal Embedding Model Technical Report U-marvel: Unveiling key factors for universal multimodal retrieval via embedding learning with mllms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.666382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.666382Z digest=sha256:e8018d3ffdaee67833df4b2f1bcdc9f1f5b3c2b7d41913a4dca482f0a82cd05e

Observation c0a083fd-8d3d-45be-9891-9fdeedfb063f · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

Douyin Multimodal Embedding Model Technical Report Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.714998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.714998Z digest=sha256:7b4583e012934d181829916f3cae46cbe9be9edca3726fbc3c88ded184d88931

Observation 4b86fb3e-564f-451b-9d28-6d192f3723bd · outbound

This paper cites Sail-embedding technical report: Omni-modal embedding foundation model.

Douyin Multimodal Embedding Model Technical Report Sail-embedding technical report: Omni-modal embedding foundation model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.863104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.863104Z digest=sha256:0dc65020a8015f8fe029b3a761e93cac86cd71dd532afd0a6bf1eb8cfe11a21e

Observation b0b60388-57cf-409f-84a1-f9d23ec05f85 · outbound

This paper cites CREM: compression-driven representation enhancement for multimodal retrieval and comprehension.

Douyin Multimodal Embedding Model Technical Report CREM: compression-driven representation enhancement for multimodal retrieval and comprehension

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:29.981168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:29.981168Z digest=sha256:d08a6ffa1fef50ba102ff9d653b8cadfda0bd6486513be052192652ca61ff58e

Observation c0714be6-1d81-4a6d-88f4-2dfcc2f2ead3 · outbound

This paper cites Vlm2vec-v2: Advancing multimodal embedding for videos, images, and visual documents.

Douyin Multimodal Embedding Model Technical Report Vlm2vec-v2: Advancing multimodal embedding for videos, images, and visual documents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:30.137934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:30.137934Z digest=sha256:ef19b448d891eb3808897c5ecc98fa645d36e02df4c15c69c08d88e615f5197f

Observation 3b7e632d-d6fb-40af-af68-989b1e1b3b9a · outbound

This paper cites an unresolved cited work.

Douyin Multimodal Embedding Model Technical Report Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:30.289633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:30.289633Z digest=sha256:675f6159c1c5314991a33f3bbe400d766d43e824243cef002e947dac52b3e4d4

Observation 1b506bc4-fbfa-4fc5-ab42-53ee3669a50c · outbound

This paper cites Through the lens of contrast: Self-improving visual reasoning in vlms.

Douyin Multimodal Embedding Model Technical Report Through the lens of contrast: Self-improving visual reasoning in vlms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:30.439578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:30.439578Z digest=sha256:45bcf0b82c3217528cae045f273e9d8448339c4a2618e6b27f39fe655fc0409f

Observation 84e18d21-0104-443a-bddc-8c3c1f9284a1 · outbound

This paper cites Qwen3.5: Towards native multimodal agents.https://qwen.ai/blog?id=qwen3.5, February 2026.

Douyin Multimodal Embedding Model Technical Report Qwen3.5: Towards native multimodal agents.https://qwen.ai/blog?id=qwen3.5, February 2026

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:30.560590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:30.560590Z digest=sha256:43eb6a4a702865da6641cf6822d87b610d76928d9df42b8dbe9d6c1e86ff27d5

Observation 1937773d-68b3-4527-8bb1-3d6f07b7ddee · outbound

This paper cites Learning transferable visual models from natural language supervision.

Douyin Multimodal Embedding Model Technical Report Learning transferable visual models from natural language supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:30.660850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:30.660850Z digest=sha256:76b403737e6dd0f3d5277a47bb44bfaea4130427387f44b77d49babf501e74e6

Observation 2e43dc25-79ce-4c9d-9d2f-343232dc8528 · outbound

This paper cites Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity.

Douyin Multimodal Embedding Model Technical Report Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:30.777455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:30.777455Z digest=sha256:05fa2d7bc325301e6f323e38b5c21334b27a529dc38896a00520e49c767a1be8

Observation fed15723-cdaf-49d6-91fd-907de12a101c · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Douyin Multimodal Embedding Model Technical Report SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:30.932936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:30.932936Z digest=sha256:6242a513efe32901097907f126c62c1258cf1798cdcae8199a7dd15ed3fbf484

Observation f448a3e9-6c5b-49af-affc-d7a02cd6e33b · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Douyin Multimodal Embedding Model Technical Report Representation Learning with Contrastive Predictive Coding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.012140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.012140Z digest=sha256:1fcc2d425484c4c190ef1e507622511685374bfd2957f34b097e055a670323d3

Observation 963540b2-0ba0-4b71-b623-cc0377cfd685 · outbound

This paper cites Text Embeddings by Weakly-Supervised Contrastive Pre-training.

Douyin Multimodal Embedding Model Technical Report Text Embeddings by Weakly-Supervised Contrastive Pre-training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.103945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.103945Z digest=sha256:62eccbfc4263b6c9290d5b91e6929d6f66d2f858d56ea665e5848fb202c05b4e

Observation c152b24a-261f-45f8-b905-2756a741643a · outbound

This paper cites Explore more, learn better: Parallel mllm embeddings under mutual information minimization.

Douyin Multimodal Embedding Model Technical Report Explore more, learn better: Parallel mllm embeddings under mutual information minimization

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.148989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.148989Z digest=sha256:9b63e34de06084d75a53993662aa9e2fe18be8f3d89f04443a57359415cbd9e5

Observation f8f32691-2a40-43a6-bb53-b74ab612635b · outbound

This paper cites an unresolved cited work.

Douyin Multimodal Embedding Model Technical Report Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.208255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.208255Z digest=sha256:bdb9369570478f4f9b9dceaad92b139be7763605ba90be32cf859dd4645bb654

Observation a2125f33-3919-44e5-8857-86b294cdbd46 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence.

Douyin Multimodal Embedding Model Technical Report Deepseek-v4: Towards highly efficient million-token context intelligence

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.266855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.266855Z digest=sha256:ca145b56265ab8425176df847b5f4e11a92a31e263174eb5db05a1963ffd5e12

Observation 96b413af-49b4-4b85-ab88-c365fe429402 · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step.

Douyin Multimodal Embedding Model Technical Report Llava-cot: Let vision language models reason step-by-step

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.302725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.302725Z digest=sha256:973b10afece83331fb6eb7f29f509f9199a029f15be0b2cb430261691985ea08

Observation 33e5a926-eed1-4f46-a222-5e131b9a9fa5 · outbound

This paper cites R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization.

Douyin Multimodal Embedding Model Technical Report R1-onevision: Advancing generalized multimodal reasoning through cross-modal formalization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.365064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.365064Z digest=sha256:8a7407339fede0e99d73a9b52b273c57c5ac924f5b543116a10a9cc8a6917b3b

Observation a8a00674-792e-41a9-b08c-9da1eb6e3369 · outbound

This paper cites Hotpotqa: A dataset for diverse, explainable multi-hop question answering.

Douyin Multimodal Embedding Model Technical Report Hotpotqa: A dataset for diverse, explainable multi-hop question answering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.409949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.409949Z digest=sha256:4063a5dc0bc5a6d403da5dc4dcd5be13490b11a865a32b4936b174555ab94987

Observation fc80b92a-2336-4626-a085-3f4dacb94b2f · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

Douyin Multimodal Embedding Model Technical Report Coca: Contrastive captioners are image-text foundation models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.450171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.450171Z digest=sha256:59de3015f0525d08b08f3253600ba390eaa474de672c77aa19636aa1fbf6b968

Observation f55123d4-c725-46d2-9c3a-173ca6bcd270 · outbound

This paper cites Visrag: Vision-based retrieval-augmented generation on multi-modality documents.

Douyin Multimodal Embedding Model Technical Report Visrag: Vision-based retrieval-augmented generation on multi-modality documents

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.528172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.528172Z digest=sha256:adf45fc0054ba2d74075a7c31e5b8cab57cd4a9c2d26226bd64add658c835fac

Observation a65f3556-1356-4d0c-9846-a126b8de6b5b · outbound

This paper cites Sigmoid loss for language image pre-training.

Douyin Multimodal Embedding Model Technical Report Sigmoid loss for language image pre-training

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.581468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.581468Z digest=sha256:2e46b5116db0362d304ccf378c110822f492de3a88cfc9e7a5c6bcaef7a49839

Observation f3549d50-53dd-4cf6-80ae-23e4a365d82b · outbound

This paper cites Gme: Improving universal multimodal retrieval by multimodal llms, 2024.

Douyin Multimodal Embedding Model Technical Report Gme: Improving universal multimodal retrieval by multimodal llms, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.640825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.640825Z digest=sha256:e62785fb525743bbe4c4b366e4b79f872c5dbed21c781d144e047058baad727f

Observation 15a46726-0c66-4a5d-9228-7a9374ac55da · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Douyin Multimodal Embedding Model Technical Report Multimodal Chain-of-Thought Reasoning in Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.706719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.706719Z digest=sha256:9f8501447758f346bb7f7b470e92a7eb7127e13bd9d712d6f1ff6838f97421ba

Observation 2d361da6-ed3b-4036-95a7-cc220fbcaec2 · outbound

This paper cites MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval.

Douyin Multimodal Embedding Model Technical Report MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:31.800765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:31.800765Z digest=sha256:f0d3df5fa45725a1054366484fd61a66148ada19e8d11dcd1fbe9275acf600dc

Pith citing papers

No inbound Pith citation observations are available.