Pith. sign in

Paper Citation Record · LEDGER

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

As of 10 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 2 inbound Pith citation observations for arXiv:2511.12449.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.12449 v3

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T22:05:18.226948Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T09:43:58.385814Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T17:05:51.109495Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44bc2dc4-c9f7-42ad-a251-f5b0a39ffd79 · outbound

This paper cites GPT-4 Technical Report.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:14.995841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:14.995841Z digest=sha256:cd887a74a59a2c38bf251ab466bd0ae500104863e0502f005bf99d233de68522

Observation ecf46e18-1751-4409-8ac9-32b2f6b5b370 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.056733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.056733Z digest=sha256:2d6d83c96ed913509e957de8c70b8e1925a5d1afef6a08d83cd4c0c4ecab94bc

Observation f8a362f0-f42c-4e09-81f9-ff34f561065b · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.154011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.154011Z digest=sha256:2d6663d4f63663ff4c88c77663b3be1e878eb6b5600471bd14687105c4071cca

Observation c5af68f4-43e9-4efb-a582-b9ec21bd6692 · outbound

This paper cites Qwen2.5-VL Technical Report.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.204752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.204752Z digest=sha256:c18028b99d399ad1869d1883ad9b63fd3e27f87b0eecc8f5aba005e56283505c

Observation 1f656f4f-c477-4c25-bac1-1f816f1a2001 · outbound

This paper cites Product2vec: Leveraging representation learning to model consumer product choice in large assortments.NYU Stern School of Business, 2022.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Product2vec: Leveraging representation learning to model consumer product choice in large assortments.NYU Stern School of Business, 2022

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.266357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.266357Z digest=sha256:dfb9f8e50e2f030735bfc9a31303b7e85fd480bae2270aef828a1cf2a7904ab0

Observation 8b19d077-880a-4caa-b7d2-4a9a4d3efd52 · outbound

This paper cites On scaling up a multilingual vision and language model.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding On scaling up a multilingual vision and language model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.288369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.288369Z digest=sha256:51269ae191a24743907b877e01d904db47d88761fee7987bc90503ea9d229cf2

Observation 48231c95-b661-424c-8608-901e2d6745b0 · outbound

This paper cites Contrastive language and vi- sion learning of general fashion concepts.Scientific Reports, 12(1):18958, 2022.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Contrastive language and vi- sion learning of general fashion concepts.Scientific Reports, 12(1):18958, 2022

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.348036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.348036Z digest=sha256:fc94c8116d290c72a6aa2d2e2fe3c5f580bc38c8f6306b5e36aaf204f4176b50

Observation ab5b734d-329b-41ab-a5c8-c81e43452649 · outbound

This paper cites Unified generative and discriminative training for multi-modal large language models.Advances in Neural Information Processing Systems, 37:23155–23190, 2024.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Unified generative and discriminative training for multi-modal large language models.Advances in Neural Information Processing Systems, 37:23155–23190, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.398261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.398261Z digest=sha256:3faa57e8fbf74cb97abb678c8ec1878f87a74881c624146a424b3836cc645556

Observation a2e06d79-121f-497d-85f3-a5de17b9f4af · outbound

This paper cites Uniembedding: Learning universal multi- modal multi-domain item embeddings via user-view con- trastive learning.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Uniembedding: Learning universal multi- modal multi-domain item embeddings via user-view con- trastive learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.455903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.455903Z digest=sha256:bf6f692948418937dd2632849634d57aeaa22dcb9721cbf6fa34ec0418b50072

Observation 2cc89181-cee1-48ac-bbc6-6155603e95a5 · outbound

This paper cites M5product: Self- harmonized contrastive learning for e-commercial multi- modal pretraining.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding M5product: Self- harmonized contrastive learning for e-commercial multi- modal pretraining

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.510395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.510395Z digest=sha256:9525add1f0873a1ec368969ae744a8a901fbcc650537dfdd88655a2c20fe27e5

Observation 0e3a5b04-9f49-45ee-bdea-59a3d82b2ffb · outbound

This paper cites Pmr: Prototypical modal rebalance for multi- modal learning.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Pmr: Prototypical modal rebalance for multi- modal learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.550600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.550600Z digest=sha256:3013e19ab522767190268fea1a570c05fe6c0f50481b981131ff0cd1e2348f14

Observation b2df3855-9b10-49ff-80cc-28944bb1ef05 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with sim- ple and efficient sparsity.Journal of Machine Learning Re- search, 23(120):1–39, 2022.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Switch transformers: Scaling to trillion parameter models with sim- ple and efficient sparsity.Journal of Machine Learning Re- search, 23(120):1–39, 2022

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.614901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.614901Z digest=sha256:95f5972ec7409d1b355c29b4612129aada4a59524e6c71285491ed35549ca7ce

Observation 7f30d795-c27b-44a5-baec-85091d690d87 · outbound

This paper cites Moon embedding: Multi- modal representation learning for e-commerce search adver- tising.arXiv preprint, 2025.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Moon embedding: Multi- modal representation learning for e-commerce search adver- tising.arXiv preprint, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.662400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.662400Z digest=sha256:1a829f07e5574e9fe19aaa14944cfae13dc5715150ac493c28a2f74e5b15d442

Observation 3b8160c3-82df-436c-855c-b751003d5745 · outbound

This paper cites Fashionbert: Text and im- age matching with adaptive loss for cross-modal retrieval.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Fashionbert: Text and im- age matching with adaptive loss for cross-modal retrieval

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.726776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.726776Z digest=sha256:b84d187bb247f9e72a778d52e0faa04ffba5cd40e68ca4fffb86108d714dd329

Observation 6b8a2bef-e14d-4263-a187-9d107f4b48ce · outbound

This paper cites CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding CompoDiff: Versatile Composed Image Retrieval With Latent Diffusion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.793078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.793078Z digest=sha256:2de55def4224667625724fd956453017d411ae04db22ffdcaaf7d17dcac54184

Observation 100c8f0d-78bb-4486-a9ac-1597c7aa4419 · outbound

This paper cites Multi-modal preference modeling for product search.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Multi-modal preference modeling for product search

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.900812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.900812Z digest=sha256:5114a3abd75592b0ffe033bbb81a199bca6debb25853eba3456ed9a2851f9f32

Observation 9ba20362-b828-4fd1-ac96-1bdd1b65a4d4 · outbound

This paper cites Au- tomatic spatially-aware fashion concept discovery.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Au- tomatic spatially-aware fashion concept discovery

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:15.957474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:15.957474Z digest=sha256:31fe33526a1cd4795f6a09c56c04a42272ff9c043cf57424226170ff852636f5

Observation c5cb3b9d-b876-4e97-8835-fe5ee40a06bd · outbound

This paper cites Multimodal retrieval in e-commerce: From categories to images, text, and back.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Multimodal retrieval in e-commerce: From categories to images, text, and back

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.003909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.003909Z digest=sha256:dc822be064952e937c734cf24dd9a24daa40c23f7ef58e32b8c78d7c3a4f0735

Observation 597fb384-51fc-483e-b694-c6d6a17b7ab6 · outbound

This paper cites A Multimodal Recommender System for Large-scale Assortment Generation in E-commerce.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding A Multimodal Recommender System for Large-scale Assortment Generation in E-commerce

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.069327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.069327Z digest=sha256:6a6d9bd8fe02010ba3d825a747808d685afe9f20206e6c8941173304e19dcaa3

Observation 7e984a8a-c35f-4888-8ad0-9a22a6d073f1 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.127974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.127974Z digest=sha256:769974d849e3454daf2e4c90a2fbf118231ba97d873ea157efd0c65b8c0c8e2a

Observation 0cd83f5e-0083-4b44-bc5f-ca5e5df421dc · outbound

This paper cites MRSE: An Efficient Multi-modality Retrieval System for Large Scale E-commerce.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding MRSE: An Efficient Multi-modality Retrieval System for Large Scale E-commerce

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.206059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.206059Z digest=sha256:b4600f305eace91c8e47960ae09e6ce7e91038ddf53714f76b4e240017c83d8d

Observation d027f8d6-f9fd-478e-a2dd-95eed8408139 · outbound

This paper cites VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.260894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.260894Z digest=sha256:d085a77eeee30d978ae7b211c1e962294452761a9a4ef9810e0b685ea8c2d129

Observation b93b2c5d-0f84-41cd-a39e-6c316e2a0e4f · outbound

This paper cites Learn- ing instance-level representation for large-scale multi-modal pretraining in e-commerce.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Learn- ing instance-level representation for large-scale multi-modal pretraining in e-commerce

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.371108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.371108Z digest=sha256:575669003fa21bbfaa850edbd8a7fdbf59fe23f99d8c36ef518731929829aec6

Observation b54e816c-2572-4492-be4c-60f68ca312e9 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.428946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.428946Z digest=sha256:aeeac1fd33876b1fac8b691f488fb51a2e4d1a392b6054e1c4d5fe9384373114

Observation 370ec113-c91b-4d60-8a32-58d6043556f0 · outbound

This paper cites Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705, 2021.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Align before fuse: Vision and language representation learn- ing with momentum distillation.Advances in neural infor- mation processing systems, 34:9694–9705, 2021

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.497586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.497586Z digest=sha256:acd2e87bfa2ccb004dd2c662402af0cff17245665ba22cd4c119ef5cf23ce5f1

Observation 1584f7af-50de-4577-8626-11dc9d351081 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.507575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.507575Z digest=sha256:e546514b5c5a2b93583cd534a0d92a12d0217d528208977338bcbd311a4fca5a

Observation 1b1c03d3-f589-4ec0-bb3d-6273bafb12ee · outbound

This paper cites Embedding-based product retrieval in taobao search.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Embedding-based product retrieval in taobao search

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.583725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.583725Z digest=sha256:f4dec9a444e2ebeadb1e92fc23d6a494cc6af8578f8192ee120c247052aca251

Observation f627da08-5f6a-4fe3-a6f5-e657083d9352 · outbound

This paper cites Adversarial multimodal representation learning for click-through rate prediction.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Adversarial multimodal representation learning for click-through rate prediction

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.687532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.687532Z digest=sha256:0660f14c410d2825ba5de637c7c53feccd080686688fbb2e307f19a60c13394d

Observation 82057d68-8acd-4b26-9bd2-626f502d8b5e · outbound

This paper cites UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.766061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.766061Z digest=sha256:0376936a89eb4e13e35cf0e3a082c12c58fc6972da42776a908dab3db6ba02c8

Observation 7602a1db-8b2f-45aa-be7e-5ab2a868d988 · outbound

This paper cites MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.854105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.854105Z digest=sha256:9cf733f5b94c617a8997ef3381dc2a7d6dddc203f344fcdb0bc42f1bb51e84a7

Observation cebaba44-1bf5-4f9f-a63d-9180cc49d397 · outbound

This paper cites Captions speak louder than images (caslie): Generalizing foundation models for e-commerce from high-quality multimodal instruction data.arXiv preprint arXiv:2410.17337, 2024.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Captions speak louder than images (caslie): Generalizing foundation models for e-commerce from high-quality multimodal instruction data.arXiv preprint arXiv:2410.17337, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:16.940763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:16.940763Z digest=sha256:6f4b4d300d08254ef258eab48b8bb4b6fe8f8756b5af67c7fe25c748876321b8

Observation 5fd4b171-9198-411f-9abc-867752dc8fe8 · outbound

This paper cites Ecom- mmmu: Strategic utilization of visuals for robust multimodal e-commerce models.arXiv preprint arXiv:2508.15721,.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Ecom- mmmu: Strategic utilization of visuals for robust multimodal e-commerce models.arXiv preprint arXiv:2508.15721,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.030089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.030089Z digest=sha256:5b8bcbd6bd7821067cce58af7a1d183d029e8a0c611849e5501c9aaf8e00acf5

Observation 6f6c0ddf-bfd9-498e-8146-b02e12fe9759 · outbound

This paper cites Multimodal pretraining, adaptation, and generation for recommendation: A survey.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Multimodal pretraining, adaptation, and generation for recommendation: A survey

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.119164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.119164Z digest=sha256:a49a4dd62a9584982015efec182aa0b8021c9a0861297e10410d294a54fb9267

Observation 6c659285-58e5-4c0f-bd44-97286fb1efc8 · outbound

This paper cites Multimodal pre-training with self-distillation for prod- uct understanding in e-commerce.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Multimodal pre-training with self-distillation for prod- uct understanding in e-commerce

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.156959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.156959Z digest=sha256:6c3b4e77f6c62246bcbce91098ce719cfb030963329b286603ace98c4156e213

Observation 092d6bf8-e64c-4167-b24b-df1e758f4588 · outbound

This paper cites Pretraining representations of multi-modal multi-query e- commerce search.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Pretraining representations of multi-modal multi-query e- commerce search

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.191297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.191297Z digest=sha256:137080f34d3fddc1c27741011dda83baa60ec791843d0ed401e81c5888acbe4c

Observation a0cc48cb-4809-45c5-8421-98b3c28d9174 · outbound

This paper cites ecellm: generalizing large language models for e-commerce from large-scale, high-quality instruction data.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding ecellm: generalizing large language models for e-commerce from large-scale, high-quality instruction data

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.246464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.246464Z digest=sha256:8e0e744778740fc8ca5232eb4d2f11b37469bf25ff17883edd7d5a41bd99425c

Observation 6c17d940-6d1b-465d-abe1-d63651156f28 · outbound

This paper cites Kosmos-2: Grounding multimodal large language models to the world, 2023.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Kosmos-2: Grounding multimodal large language models to the world, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.281869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.281869Z digest=sha256:9df8a0de0cc65bded5f86d43e2bfab9d934eb9c32bd53fae9837defed42e59cf

Observation c0dd4b5c-2a5b-44be-9250-1484fa1c70d9 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.300953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.300953Z digest=sha256:0e9257f073033b65759302d4a9489d02ba51352ff44cbe200331a5af9c28c3fa

Observation 4b31a3c2-e34e-48a4-a4c8-8e53567e7035 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.334918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.334918Z digest=sha256:700e9d387befe92971e371552b105a10e3f6620e9d6a60ab9c9a92cffa839997

Observation 139151d2-f2fc-41fd-a3dd-4b7533ffda38 · outbound

This paper cites LLaMA-E: Empowering E-commerce Authoring with Object-Interleaved Instruction Following.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding LLaMA-E: Empowering E-commerce Authoring with Object-Interleaved Instruction Following

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.373059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.373059Z digest=sha256:175aaf4f83b2b7c2b73ebe630a50e062cef6c30db19f5e1a40a858114f279718

Observation 9eefee5f-073d-4ef6-9d68-809e37c640e2 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Gemini: A Family of Highly Capable Multimodal Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.407008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.407008Z digest=sha256:bd2036157823d8f5d4be9f2e5d5da3943c9e6ef9b660395f9e5a958a2216a350

Observation 13a1aea9-9593-44e1-a1a3-311ef52a1b7c · outbound

This paper cites Merlin: Multimodal & multilingual embed- ding for recommendations at large-scale via item associa- tions.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Merlin: Multimodal & multilingual embed- ding for recommendations at large-scale via item associa- tions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.445991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.445991Z digest=sha256:b15286da97fc5ddfa676fbf352c0c1670021ddd1243532b2e8f7e34ff9e4fba4

Observation e40ec159-b335-4811-9c04-420bfd016ecc · outbound

This paper cites Stablerep: Synthetic images from text-to- image models make strong visual representation learners.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Stablerep: Synthetic images from text-to- image models make strong visual representation learners

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.490123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.490123Z digest=sha256:726282b284dc0467a178e86521180b6231604f1b1fda96647b0f3dc720b70a57

Observation 2082b2aa-7945-4c5d-94d1-9157b08985f6 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.608149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.608149Z digest=sha256:88dfe44ceb857d27bfc80216b4eece73c5da48da7e351470dba05c1938bdf1f7

Observation 81ce17bc-44ed-4f45-b5d7-b2e1d008edf1 · outbound

This paper cites Missrec: Pre-training and transfer- ring multi-modal interest-aware sequence representation for recommendation.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Missrec: Pre-training and transfer- ring multi-modal interest-aware sequence representation for recommendation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.698070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.698070Z digest=sha256:47bfb558e503e3acf5c94907dea7e5246698ebe6921a40ae0a1b665081e07358

Observation bcf5f517-de33-482c-8f61-c9b19fbf1b32 · outbound

This paper cites MIM: Multi-modal Content Interest Modeling Paradigm for User Behavior Modeling.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding MIM: Multi-modal Content Interest Modeling Paradigm for User Behavior Modeling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.724142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.724142Z digest=sha256:b7c209c7c66383ab020163f9522a2592a09bbddced209cfd9f94b839ac34b2a1

Observation 8bfecc26-a443-40af-a4df-ed6dadb46e5c · outbound

This paper cites Commercemm: Large-scale commerce multimodal representation learning with omni retrieval.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Commercemm: Large-scale commerce multimodal representation learning with omni retrieval

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.832740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.832740Z digest=sha256:08c41d5de321b0658c882cae1b9541136636d5bc262589f7e93a0903ee48d434

Observation cdfc8ab6-21f0-4117-b4f9-373cb95d0849 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Florence: A New Foundation Model for Computer Vision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.860103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.860103Z digest=sha256:0c19049374d6fb7f463b1b992e00318420e86eb9fe49838cc2fc767b4705de2f

Observation 6f62f263-7380-4be1-a86e-d376852e06ae · outbound

This paper cites Prod- uct1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Prod- uct1m: Towards weakly supervised instance-level product retrieval via cross-modal pretraining

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.923111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.923111Z digest=sha256:b9ebc66aef32b84ae9e5e104122c3cb27084c4c8cf67729ac5957e215ad73e90

Observation cd6e71e7-b856-4e49-81ac-dbf3da804f14 · outbound

This paper cites Moon: Generative mllm-based multimodal representation learning for e-commerce product understand- ing.arXiv preprint arXiv:2508.11999, 2025.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Moon: Generative mllm-based multimodal representation learning for e-commerce product understand- ing.arXiv preprint arXiv:2508.11999, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:17.974836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:17.974836Z digest=sha256:7330d1fcc5b7d1150c82608c4b6595c0046a53e3ccb513c9410100fbc89d2a31

Observation 3a63e761-c31d-4440-b1cb-7340d8fe431b · outbound

This paper cites Sharper and faster mean better: Towards more efficient vision-language model for hour-scale long video understand- ing.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Sharper and faster mean better: Towards more efficient vision-language model for hour-scale long video understand- ing

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:18.024821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:18.024821Z digest=sha256:47ccdb890019176dca14c4a9b4d1464cd69ba83df38563a0b3ec5dcf4c7f7dc7

Observation 600e39cd-c0b4-4bc7-97ce-9207fe377fb6 · outbound

This paper cites Modality-balanced learning for multimedia recom- mendation.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Modality-balanced learning for multimedia recom- mendation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:18.033214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:18.033214Z digest=sha256:0bdfd3f6dcc2a2ee02542e5510f89007d43f7d950b28f6ccd06c5316f89e2e3d

Observation b8594a87-65de-468d-9c12-c0cb0cf06889 · outbound

This paper cites GME: Improving Universal Multimodal Retrieval by Multimodal LLMs.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:18.066051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:18.066051Z digest=sha256:dafaf4bbd33daaa3f746d98fe54526cf78a656ae085e0d69cc12c12dadf1f718

Observation aad18658-13fe-448e-a68c-a2c41118e8fc · outbound

This paper cites Delving into e-commerce product retrieval with vision-language pre-training.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding Delving into e-commerce product retrieval with vision-language pre-training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:18.114220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:18.114220Z digest=sha256:c25ff59ddb688fb7d052c375d19f5f05e4eb37680ebb11e9cb1392b078150476

Observation 2590b86f-d365-48bd-8cb7-0945ba9214fc · outbound

This paper cites MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding MegaPairs: Massive Data Synthesis For Universal Multimodal Retrieval

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:18.174178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:18.174178Z digest=sha256:8694ed7c766152b5d5cd2406f2d1c85ecfc8843509c1c30db94ce6b5a85e0cfe

Observation fcf9fca9-6c78-4a1a-b3fd-fac642c385ca · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T22:05:18.226948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:05:18.226948Z digest=sha256:7ca2661e4b1a798377508b564ff4665ad84a19afc1f78d3a0a38c232422bee10

Pith citing papers

Observation 31f70ad7-a02e-432e-b384-4c3dba80927f · inbound

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications cites this paper.

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-31T01:08:53.304372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T04:11:05.226043Z digest=sha256:58e6297ca17f7c35a990a50fde2fb8dae57759f1452888fca58c691b89a71096

Observation c896babc-7ec7-4c54-b0cf-9fd7798cfd10 · inbound

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications cites this paper.

JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-31T01:08:53.304372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T09:43:58.385814Z digest=sha256:c7caaff916defd89f6f848e2c72e8c56cbaeba7dcd3633928790ff372fb65632