Pith. sign in

Paper Citation Record · LEDGER

Matryoshka Query Transformer for Large Vision-Language Models

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2405.19315.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.19315 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:28:37.214098Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T11:46:56.083031Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e8c4dcd8-496a-4d67-bf48-0cea65accaed · inbound

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression cites this paper.

FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression Matryoshka Query Transformer for Large Vision-Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:28:37.214098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:28:37.214098Z digest=sha256:2f4c11f0565a6ff5dc0c5cf1e4e69f6dab58ae0c3f4848b3a4c401a85d0e325f

Observation 01a6a45f-85ff-4c90-be31-7d84dc1e1bdb · inbound

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs cites this paper.

Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs Matryoshka Query Transformer for Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:57:32.429234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:57:32.429234Z digest=sha256:a3e7fa527758a9c098a9ad08df61b14e3ba876c02edf709dab61f4b7616107b5

Observation ab721087-15d3-425a-a89d-1a02cf732f04 · inbound

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay cites this paper.

p-MoD: Building Mixture-of-Depths MLLMs via Progressive Ratio Decay Matryoshka Query Transformer for Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:31:44.388230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:31:44.388230Z digest=sha256:183b790b8cf2d1e9fe95468a7fae8c37abc996b2326c235ad57642d36ac43c8d

Observation 00a7912e-ed78-4a99-a60c-0fcc7c2dd508 · inbound

Masked Generative Nested Transformers with Decode Time Scaling cites this paper.

Masked Generative Nested Transformers with Decode Time Scaling Matryoshka Query Transformer for Large Vision-Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T19:21:56.226965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:21:56.226965Z digest=sha256:ba7ea96879f903ad104d132db2179281560653f75f05d306d9c21d7cdb646a0d

Observation d5b5039e-25b6-42c4-90b9-336ed569aca1 · inbound

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding cites this paper.

DynTok: Dynamic Compression of Visual Tokens for Efficient and Effective Video Understanding Matryoshka Query Transformer for Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:56:05.312460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:56:05.312460Z digest=sha256:420ca9bea9bda6b11ab65cf756fedccaa041f1ac4b91ab3f747179468eb20954

Observation e2b31772-2236-4f85-81e6-f33db26f10f1 · inbound

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models cites this paper.

Fourier Compressor: Frequency-Domain Visual Token Compression for Vision-Language Models Matryoshka Query Transformer for Large Vision-Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:40:43.049199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T22:40:39.892802Z digest=sha256:32f1c334e0cb15d2befad8b54f7768a0fcf5dac09a1b557aaa3e03eab72956d2

Observation 6df659c3-93f0-4e94-a40c-e3081d66684e · inbound

MIPIC: Matryoshka Representation Learning via Self-Distilled Intra-Relational and Progressive Information Chaining cites this paper.

MIPIC: Matryoshka Representation Learning via Self-Distilled Intra-Relational and Progressive Information Chaining Matryoshka Query Transformer for Large Vision-Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:12.450982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-08T03:35:44.435355Z digest=sha256:a37ef4cb015e7ee7873ff6e580943fce810d627fc74ece61244c2f7f92731ade

Observation eeefe21b-20f1-4c81-8f53-dd5bf9f5f7e2 · inbound

Adaptive Tokenisation Via Temporal Redundancy Masking And Latent Inpainting cites this paper.

Adaptive Tokenisation Via Temporal Redundancy Masking And Latent Inpainting Matryoshka Query Transformer for Large Vision-Language Models

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T11:46:56.084662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T02:55:00.586061Z digest=sha256:dfceaa864c7279f26c0b1c62e702cdc3d2d2cd65c15fd2cc45097ad1c14c7e03

Observation 902cb00f-32af-4c09-9427-9dad7aedafb2 · inbound

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models cites this paper.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Matryoshka Query Transformer for Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.400885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.400885Z digest=sha256:fb04452e98b5d9f2a71c53142f0da7bee7c0594c7c537dcfd6fd8ab3f81d79e7