Pith. sign in

Paper Citation Record · LEDGER

Scaling Vision-Language Models with Sparse Mixture of Experts

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2303.07226.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.07226 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:18:15.273713Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T06:39:37.941910Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation db1965d2-75c2-4e25-958b-8ae9f89d4f2d · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.400638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:6ec12c7fd381424ad50dee4dbbaa0eac7f475452627672c8656f3d801a7d6205

Observation 49cef5ba-849a-45b6-b550-858b89f8a9b4 · inbound

SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning cites this paper.

SMoLoRA: Exploring and Defying Dual Catastrophic Forgetting in Continual Visual Instruction Tuning Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:46:29.633610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:46:29.633610Z digest=sha256:8bbb63b71f5cb0f80a49953e9703f06f0006539d2d5efb2a4c1b6cba1adc4408

Observation 716ee41d-c6a3-4d56-8a65-e064d317c051 · inbound

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey cites this paper.

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 179

Resolution
unresolved
no resolver link, observed 2026-08-11T23:54:24.035371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:54:24.035371Z digest=sha256:dd9d446fe1ba65e28eb5a9f26a1404b5679c8e56409ea0acdb54ab738e34b2dc

Observation 89ef3ba2-c81f-4b66-b766-9065da2ad385 · inbound

SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts cites this paper.

SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T20:40:39.792788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:40:39.792788Z digest=sha256:b9803537c47e1c0c93ac275cd07e5ea0d0dd7f17e95967f91f6c03dd926ddeef

Observation f2447408-82ca-41d8-a704-11e2131eea0b · inbound

LMFusion: Adapting Pretrained Language Models for Multimodal Generation cites this paper.

LMFusion: Adapting Pretrained Language Models for Multimodal Generation Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:11.014189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:11.014189Z digest=sha256:71ddc47532a33ad843d2d62760356174d85ef6ea2d7624efd1c92adc76d772f1

Observation 783102d4-6863-417c-8daf-09a3d388a22d · inbound

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts cites this paper.

Eve: Efficient Multimodal Vision Language Models with Elastic Visual Experts Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:12.602415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:12.602415Z digest=sha256:88d6689dd8c7eb5808afef3fc111a9c0ca7959115cde4123702e8ed2ebaf1986

Observation 5b55ca24-de84-4dc2-8584-72e062453465 · inbound

X-Fusion: Introducing New Modality to Frozen Large Language Models cites this paper.

X-Fusion: Introducing New Modality to Frozen Large Language Models Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T05:18:15.273713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:18:15.273713Z digest=sha256:25b416fb6a014161061f284fe467c432d16d1e8e419756fded8ed17f72d3f002

Observation b25f7791-3c29-4b06-a152-e42c32f6c2c6 · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:15.339882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:15.339882Z digest=sha256:94ee5fbd9c4c3f5ad3a0503831c7f842929778d370b5f00474dc10eafcac9e48

Observation a9c08a38-53d3-457e-969e-ede02d5259a4 · inbound

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities cites this paper.

SMAR: Soft Modality-Aware Routing Strategy for MoE-based Multimodal Large Language Models Preserving Language Capabilities Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:07:18.721641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:07:18.721641Z digest=sha256:d662b57bddcb96c18231495a7f0751ecdb37cfaa013798aeb50cbaa58633719d

Observation 3e3f5af5-d0c2-4d93-89e5-8a1fe41310a5 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:59.424967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:59.424967Z digest=sha256:251731f7a5d8d69ddc43d85a3d195de75d3a71e5b85964ae93043a0690a1d377

Observation 2c11a169-0341-49a9-9974-fe211344aedf · inbound

Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection cites this paper.

Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-time Open-Vocabulary Object Detection Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:54:36.682649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:54:36.682649Z digest=sha256:caa232a1e64bc1a8292c80159cb56f1e3e0ccb3ce7a0af6e12247f9bdf2e9c51

Observation f1052675-1485-411c-be54-ae88fcb51360 · inbound

Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts cites this paper.

Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:59:40.985914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T05:56:18.498128Z digest=sha256:3bfb25e28b376fbb50b3053e0c0553a483d8a4cd6f898b2e3e9d2e45df8728b1

Observation b8b3cdb1-b86f-44ed-b6ca-8b81a2c2cd82 · inbound

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models cites this paper.

Hyperbolic and Evidence-Prioritized Experts for Large Vision-Language Models Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:26:00.056323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T22:43:33.929871Z digest=sha256:de642e7bb18e20113411c95049e88efc02af18d6e5fed761dcf7cac22a55cec2

Observation 95bc0453-fb17-4af7-952e-57687f9eeaf9 · inbound

Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models cites this paper.

Behavioral and Representational Evidence of Binomial Ordering Preferences in Large Language Models Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.943595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T14:18:11.215278Z digest=sha256:d0a8c2e803ad8ed9cf9ff98850e6fc2d0c9e196f7c1974163a081945c93732cb

Observation 02970255-9889-4f45-b202-1d8140b64edb · inbound

Mixture of Cognitive Experts in Large Vision-Language Models cites this paper.

Mixture of Cognitive Experts in Large Vision-Language Models Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T09:13:07.507164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:13:07.507164Z digest=sha256:d37f0f23b36fab8ab73f592bb7529842505bc79485f27fef37e2f62e64ccb7c4

Observation 6b1b71d1-d31c-4c84-a3cd-d3000e89bac7 · inbound

Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification cites this paper.

Weight-Space Mixture-of-Experts for Implicit Neural Representation Classification Scaling Vision-Language Models with Sparse Mixture of Experts

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:26:57.323356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:26:57.323356Z digest=sha256:5bb394273715f695ac7209c5e7600b2a5b384a35cb3d2dc14202c157072a7c72