Pith. sign in

Paper Citation Record · LEDGER

MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2411.17762.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17762 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:25:55.957533Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T23:52:16.787762Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c23031d0-36ec-4096-8cf9-467741caeef5 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.957533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.957533Z digest=sha256:6563b53464e34e07ed7cea36591f0d1107bd2dd1ea5bd89690d7d2a20d252884

Observation 91c82944-264e-4de9-a118-6323c34ab188 · inbound

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies cites this paper.

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:52:16.792292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:51:43.934329Z digest=sha256:c5b4e3c391a6820d0d5cbbb3c4d63049337d924830647973797e34d5705259e3

Observation b9c097b8-2027-4c6b-aa58-b6df8e08860f · inbound

Emerging Properties in Unified Multimodal Pretraining cites this paper.

Emerging Properties in Unified Multimodal Pretraining MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:23:42.072208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:23:41.854132Z digest=sha256:37012ca54c42cf77f3622bb7c345cc05c2338695334a450a6820a721607950df

Observation d5249f30-1d36-4c24-8522-d533979582ff · inbound

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM cites this paper.

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-22T02:10:56.214807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T02:06:35.204166Z digest=sha256:8b280328dcfc972918b8b2fb3778feab7c41103afbaea3ff73bb246df8fa872d

Observation 48b86967-00b7-4792-bd88-fd439dc7beef · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.866825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:4dc177fba230f040c916a325eadd887a847d5c335b223fc186fb0f0dae81ad08

Observation eca16e4f-e042-4a5e-9470-7a5a5991c70b · inbound

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation cites this paper.

UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:49.883147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:49.883147Z digest=sha256:21c94066fe8015e3a122f8b17b440c9d20f6a018dae35b7ebaefa84e099ea762

Observation bac26ac1-be08-415c-9dd1-0396a829c3e6 · inbound

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs cites this paper.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.435592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:25ac4146d442c5e6b61c44dbcb0ec4387b4d6837ad7135b9dfd69b03b19fa40f

Observation f2053d36-e18c-4b43-9566-42a2ccfc4c53 · inbound

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation cites this paper.

ChatUMM: Robust Context Tracking for Conversational Interleaved Generation MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T03:57:32.313692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:57:32.313692Z digest=sha256:a6f0d6ccc6dc780a800ace9bad6cbc9ec92a8e94162ef29f47baa4106f7d1f06

Observation 30e90ac8-9442-4990-8e84-20318512ac85 · inbound

Twins: Learn to Predict Unified Representations with Focal Loss cites this paper.

Twins: Learn to Predict Unified Representations with Focal Loss MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T04:29:49.896752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T04:29:49.896752Z digest=sha256:1054b2a1dbbf97004fdd756faa9e06dd8fb2fe3987ca89b9180a381b96309d26