Pith. sign in

Paper Citation Record · LEDGER

Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2409.09086.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.09086 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:47:27.279636Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T23:04:01.890632Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f2b9f6b4-7c98-4a65-96e6-5ee84759417e · inbound

ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression cites this paper.

ClusterKV: Manipulating LLM KV Cache in Semantic Space for Recallable Compression Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T22:42:43.244750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:42:43.244750Z digest=sha256:48a0920370448809bb5e935a425d181e7770e09dec9ccf433ca721bd1fbd14f5

Observation dff5c09c-b35e-4453-a966-7649deab5b34 · inbound

Efficiently Serving Large Multimodal Models Using EPD Disaggregation cites this paper.

Efficiently Serving Large Multimodal Models Using EPD Disaggregation Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T04:29:34.761337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T04:29:34.761337Z digest=sha256:1e5acfb3088702f8e6f02c4e3c69c921555922de548c7d18757a36a7fc532d76

Observation 5bc27a61-04d8-46e9-8af5-bb45ed83bb6a · inbound

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos cites this paper.

TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T10:47:27.279636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:47:27.279636Z digest=sha256:d191b99d225e6e93ada1b1e3c257ee0b2209a32babfe8efc68c0b0805f2cc5c6

Observation e0cb6fcd-ee4f-48a2-a2f3-9e127d8494a8 · inbound

Technology solutions targeting the performance of gen-AI inference in resource constrained platforms cites this paper.

Technology solutions targeting the performance of gen-AI inference in resource constrained platforms Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:57.236905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T16:36:26.337511Z digest=sha256:ad29cedf65584f2d55a2926f0a423edb8bf64ce17d4a3dca07cc516737a3fd5e

Observation 8aca015f-10e1-40b8-b7da-54541297633e · inbound

Convex Optimization for Alignment and Preference Learning on a Single GPU cites this paper.

Convex Optimization for Alignment and Preference Learning on a Single GPU Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:05:23.135688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-25T05:01:31.560963Z digest=sha256:e739b6a5713092ce82d9c67ca9bb5e4781ed453e8063529ec5aed671f13278f0

Observation 73155149-82c1-4d52-ada3-69ad62ab3d0c · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU

Reference 246

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.892166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:63bc61bb871850d6196398a1b05f7190d5865c1ff8ef45324f07def1e21cd7d0

Observation 1e70f0f2-cf34-40c3-a958-22be8f4e8bf0 · inbound

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering cites this paper.

StreamOV: Streaming Omni-Video Understanding via Evidence-Guided Memory and Response Triggering Inf-MLLM: Efficient Streaming Inference of Multimodal Large Language Models on a Single GPU

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:02.112392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T22:27:17.092553Z digest=sha256:c1d806f8c755871f8edd05deea05b742ba713f85c2143d6fec2b54c3766e0c95