Pith. sign in

Paper Citation Record · LEDGER

MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2410.11829.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.11829 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T01:04:00.618591Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T17:56:24.994767Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8246b6d9-140e-4581-9051-2b84db952a05 · inbound

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models cites this paper.

Instruction-Guided Fusion of Multi-Layer Visual Features in Large Vision-Language Models MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T01:04:00.618591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T01:04:00.618591Z digest=sha256:b88bec15d8691d3800081a5784c33c15545d78536467c693cbae6b46fe64d7b2

Observation 4a45b99a-6465-4660-92f8-72b2677fe431 · inbound

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach cites this paper.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:33.418185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:33.418185Z digest=sha256:a5acd13842a6f682a5b80c4d56863fb94a6fa4f6783254b3ad28241222d040e2

Observation 51204491-2696-4535-b023-231582a16ba9 · inbound

Self-Disentanglement and Re-Composition for Cross-Domain Few-Shot Segmentation cites this paper.

Self-Disentanglement and Re-Composition for Cross-Domain Few-Shot Segmentation MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:24:16.961980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:24:16.961980Z digest=sha256:aeec6798bf500695b6c42be07fa7f7be5133b133df841256742a71787ecef881

Observation 00fdf60d-60db-47f4-903e-d5d4a14d023b · inbound

MUFASA: A Multi-Layer Framework for Slot Attention cites this paper.

MUFASA: A Multi-Layer Framework for Slot Attention MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:36:29.489126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:36:29.489126Z digest=sha256:74c9d21b41396390307c97a574885b72596e720cb76874d662a894342d54bfd3

Observation e9e52549-37d3-4403-a684-0843d904a529 · inbound

Mema: Memory-Augmented Adapter for Enhanced Vision-Language Understanding cites this paper.

Mema: Memory-Augmented Adapter for Enhanced Vision-Language Understanding MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:56:24.999212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T17:53:29.860496Z digest=sha256:d9413dcc4cdd421746ebbafeaf35bed762fa857a28d5c4bad8c95d49c8731359

Observation e70482fa-705c-4615-b695-0a77e168dd1b · inbound

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization cites this paper.

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:31:23.380820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T05:31:06.963637Z digest=sha256:54a5010b6cb6ff998fdb62d73fa395c73f8cc9b0465065a02285551b394c513d

Observation a8520dfe-616d-4d28-b50e-1060c7f899ad · inbound

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization cites this paper.

Beyond the Last Layer: Multi-Layer Representation Fusion for Visual Tokenization MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:42:30.582412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T07:40:16.927031Z digest=sha256:5e903b22e35f2e6895d5004a144aa932c970f6a760d32c7437518c0a738dbc6b

Observation c2fffe71-b3a8-4c11-b98d-61e614c74e39 · inbound

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models cites this paper.

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:48:44.056463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:48:44.056463Z digest=sha256:d79873473ceb26aeb5d051be092ea5726b4ed9ffe65f34b222f47c70ba7f9c1e