Pith. sign in

Paper Citation Record · LEDGER

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

As of 17 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2608.03812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03812 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:00:10.950717Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact4
  • verified fuzzy22
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4ff1c81-251e-4035-8a6d-98e3d0f22594 · outbound

This paper cites Qwen3-VL Technical Report.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.728615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.728615Z digest=sha256:d51298874104a4a4ac674c2184cd8808f0785b2c5ed39352ec98b9bb2f7cbace

Observation f2a2490a-d605-4a04-b8aa-288cb3f7a3f2 · outbound

This paper cites Token Merging: Your ViT But Faster.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Token Merging: Your ViT But Faster

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.733972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.733972Z digest=sha256:ebc27774860d2a1eb0b91e599266a8f6db00c8f24d4f96ca89e83aba43858c0c

Observation 85a71cfc-c61f-427f-bb6c-9328324d3d24 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.225663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.739333Z digest=sha256:584bebdc4080323db65329a803948107a8f0263fa1a594e1651342600a8c9a42

Observation 6b8993aa-07a6-4baf-91de-9cc073e2eb77 · outbound

This paper cites Avocado: An audiovisual video cap- tioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Avocado: An audiovisual video cap- tioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.744559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.744559Z digest=sha256:98eb19681a5c645197eff5a1a5978835dd0e6af1cd7dff427d7808719ef43de7

Observation fc11dee5-e447-40a7-bf99-f86c01531b48 · outbound

This paper cites MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.749076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.749076Z digest=sha256:5c4efab74a9869ca21bd187f3051f18370130e9249abdc1bce3ab7901244c062

Observation adf782b0-604e-4e06-ac09-592868da529c · outbound

This paper cites OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:00:12.324177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.753980Z digest=sha256:7fa8678b86caa5ca39c9886a6b1b49c084797a27f80c850bbb03524441d3c563

Observation d4c6e423-058f-4fc6-a23b-65cb6784fea9 · outbound

This paper cites OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.759361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.759361Z digest=sha256:bf4cd9f999d5340cd5f663363823756db7fda8c738dec17fa9c18b718618c035

Observation b6767fe0-689d-4823-a9eb-22133b8c728e · outbound

This paper cites Unified spatiotemporal token compression for video-llms at ultra-low retention.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Unified spatiotemporal token compression for video-llms at ultra-low retention

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.212702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.763940Z digest=sha256:1bd4117bd797e96651feb71998a64d77a05a88142f86892040582345769215df

Observation 0aa1385b-96f6-4d00-897b-e93558c66bde · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis.Knowledge-Based Systems, 2016.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Study on density peaks clustering based on k-nearest neighbors and principal component analysis.Knowledge-Based Systems, 2016

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.199415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.768085Z digest=sha256:7278530fa153a0714d3ebe9fed45d7bd7c2dd7ea4acfc0413291323ae4f933dd

Observation c7229e11-888c-461e-bb5b-e131dff37ef1 · outbound

This paper cites Flashvid: Efficient video large lan- guage models via training-free tree-based spatiotemporal to- ken merging.arXiv preprint arXiv:2602.08024, 2026.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Flashvid: Efficient video large lan- guage models via training-free tree-based spatiotemporal to- ken merging.arXiv preprint arXiv:2602.08024, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.772364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.772364Z digest=sha256:c27037ede62ded379e867086858745dec39ba9268e5c84537f3d5844aa19e6cc

Observation 1563059e-d21b-4e68-84a5-fffc1a5b873d · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.776575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.776575Z digest=sha256:7f2905549db8dcdf90a0b32319fba745fb9686727a3d52a07ff451d6fe88f6ac

Observation 85f39ece-efb9-47a9-8589-866d897e697f · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.187016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.781252Z digest=sha256:263dbf3ec637d399d959858ed119ebd0ec66229a245a4c9fe46a053c499ff1a4

Observation df338285-baa2-473b-a284-b677b5a98a8e · outbound

This paper cites ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.785575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.785575Z digest=sha256:8b2dfc3c9ffaf353c13166a5fe5011e037eec87c521be4ca1b3e1f8c635a7711

Observation 8724ec8e-e5ea-457b-b42d-53d63a4e1e83 · outbound

This paper cites EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.790061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.790061Z digest=sha256:ac14a8ba4039eff654181863c8b2d2762b61cf89d1165fbfeff7a4f23c2b21e4

Observation e80f3558-050b-468e-920d-7cfe665d1e4b · outbound

This paper cites Echoingpixels: Aliasing-resistant joint token reduction for audio-visual llms.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Echoingpixels: Aliasing-resistant joint token reduction for audio-visual llms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.173519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.794328Z digest=sha256:0904c5028b830b156b432dff02a7e9a920deba7f8de6470c751b6def94b01315

Observation 9e4c7575-aca0-494a-8d10-8be0b01266c7 · outbound

This paper cites Gemini 3.1 pro model card.https: / / deepmind.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Gemini 3.1 pro model card.https: / / deepmind

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.160700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.798302Z digest=sha256:2c0cd087f3848e2b840f75cbbdcdacfd298d8ec1e9f988f251f3276f05844309

Observation f691b94f-fdd3-45d0-82b7-fb384736a7ad · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.802417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.802417Z digest=sha256:d183cfe16fe17c96b6c6dd4b55cfefe5762f65e824439d6e813c1a6859cec1c6

Observation 85fb0f3a-968e-4add-bd8b-17f03ba460b7 · outbound

This paper cites GPT-4o System Card.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.806834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.806834Z digest=sha256:5157d8c0cfeaeb3a8eff9e62fe272567ec3ea3c8fb7693f73178515755115037

Observation 669715a8-0864-4eda-a1c1-70f44e032344 · outbound

This paper cites ContextGuard: Structured Self-Auditing for Context Learning in Language Models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models ContextGuard: Structured Self-Auditing for Context Learning in Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:00:11.978204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.811574Z digest=sha256:2225f4a20fac647be0fde672174fe80ba17d65be07e9f0b43a66e4d9f0558828

Observation bb040714-fcd5-48b0-a191-969f1ff2c6ce · outbound

This paper cites Token pruning in audio trans- formers: Optimizing performance and decoding patch im- portance.arXiv preprint arXiv:2504.01690, 2025.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Token pruning in audio trans- formers: Optimizing performance and decoding patch im- portance.arXiv preprint arXiv:2504.01690, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.815853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.815853Z digest=sha256:05fde16100d63f4de1f84a7028638fb4907081e5678d9609f8bbe55c8207d6d3

Observation 969fe028-3840-4a74-9ff4-ab0d84dfadbc · outbound

This paper cites Omnivideobench: Towards audio-visual understanding evaluation for omni mllms.arXiv preprint arXiv:2510.10689, 2025.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Omnivideobench: Towards audio-visual understanding evaluation for omni mllms.arXiv preprint arXiv:2510.10689, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.820110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.820110Z digest=sha256:3797d62f6d0738b1cba4ab69208e6329c3e24708068999642bf95d4899333c23

Observation 2a535e23-31cd-43b0-b60a-cd5745492821 · outbound

This paper cites OmniGAIA: Towards Native Omni-Modal AI Agents.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models OmniGAIA: Towards Native Omni-Modal AI Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.824327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.824327Z digest=sha256:48d70d2dc972f0312571a5e529da8d75bc11898875f0190ddebcf6b6ebebd7f3

Observation 02549923-0200-415b-82f1-2e61ecf47e87 · outbound

This paper cites Speech- prune: Context-aware token pruning for speech information retrieval.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Speech- prune: Context-aware token pruning for speech information retrieval

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.147517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.829073Z digest=sha256:dcd007535c895b29a111ef4d5c6bacba135f187e75f2b09beb729eff0370561a

Observation ec615e69-17b1-4750-9904-d44b2db9a144 · outbound

This paper cites Video compression commander: Plug-and-play inference ac- celeration for video large language models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Video compression commander: Plug-and-play inference ac- celeration for video large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.134724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.833105Z digest=sha256:5849e732079bad76ae07cb3673d277b811f4a05af17529e8953bb48eb0b5fe08

Observation 1e67b4fa-d291-4650-8def-ea6526ef4f2d · outbound

This paper cites Gpt-5.5 system card.https://openai.com/ index/gpt- 5- 5- system- card/, 2026.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Gpt-5.5 system card.https://openai.com/ index/gpt- 5- 5- system- card/, 2026

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.121235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.837170Z digest=sha256:298f08c373a3a5aa7229a1bd489e72a6986e39a7b3a41e98ab3739b9c4f36771

Observation c0be0490-6173-4d6f-bfd8-bde6e8008139 · outbound

This paper cites OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:00:11.658798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.841260Z digest=sha256:c75e146a748fb52882f487046e4042691736fb1333e8ba1fbfdf8f777b967f86

Observation 01735dad-625d-4a25-ba0b-57e432753fd7 · outbound

This paper cites Clustering by fast search-and-find of density peaks.science, 2014.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Clustering by fast search-and-find of density peaks.science, 2014

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.108069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.845601Z digest=sha256:b17b8257d7cd844d027bd6455bbb94e43f76d541645f776790103a71d5b38ea6

Observation c7e7f61a-fc50-4892-8ba9-396f3c1eacd7 · outbound

This paper cites Fastvid: Dynamic den- sity pruning for fast video large language models.Advances in Neural Information Processing Systems, 2026.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Fastvid: Dynamic den- sity pruning for fast video large language models.Advances in Neural Information Processing Systems, 2026

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.095057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.849604Z digest=sha256:c8987aa7f4972716287fdcfca9189586faea8994fb297187d0787229a96ca67a

Observation a797368c-461c-4461-b661-4f748b721c06 · outbound

This paper cites Mavors: Multi-granularity video representation for multimodal large language model.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Mavors: Multi-granularity video representation for multimodal large language model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.988394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.853895Z digest=sha256:f97c095000d7488bf6bfb1dc747835b3654be97179e42816d2acd9ae1850fda1

Observation 1dad8e06-66ef-42ec-b8ef-8d8146649fb1 · outbound

This paper cites Audio- visual llm for video understanding.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Audio- visual llm for video understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.846338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.857877Z digest=sha256:55f927ae97fed06951fa4cdf57f84807d46d21ac6f67436985e6b45b0737de9b

Observation 98885e3e-5829-4942-9e14-7b991be14341 · outbound

This paper cites TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.862282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.862282Z digest=sha256:cf04791e8fd151bc5c76679e7a0e9f65026fe594bc73a25b180e9b02c0a477a2

Observation 5ca6b901-294a-4feb-ad39-e4f4a9168434 · outbound

This paper cites video- salmonn 2: Caption-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models video- salmonn 2: Caption-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.866665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.866665Z digest=sha256:27309cc21a2f162cf901163ce9143806fd1f8c7ab1b4cde6c122d6d4632f02b3

Observation 58d07e37-2acf-49f2-b954-dc6c6718216e · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast 9 video large language models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Dycoke: Dynamic compression of tokens for fast 9 video large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.663031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.871020Z digest=sha256:0094791a479fd0155b67d5599ee29e73fdcc86f7588c54e8b3dcdb1cf8a928bc

Observation e7911a0d-46d9-46a3-a5ed-b0260e4b4a63 · outbound

This paper cites Omnizip: Audio-guided dynamic token compression for fast omnimodal large language models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Omnizip: Audio-guided dynamic token compression for fast omnimodal large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.509379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.875168Z digest=sha256:629929203b903b99d3663c0c7c30835d3c60fb6db4a3ccb008047fe871142b2a

Observation f599502c-6604-45db-8365-c1e444359355 · outbound

This paper cites Lvomnibench: Pioneering long audio-video un- derstanding evaluation for omnimodal llms.arXiv preprint arXiv:2603.19217, 2026.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Lvomnibench: Pioneering long audio-video un- derstanding evaluation for omnimodal llms.arXiv preprint arXiv:2603.19217, 2026

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.879590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.879590Z digest=sha256:c5bfcf428990c4fae0ecd9bfc480ff71e9a19249e982826027079d499d0d14b3

Observation 4f9388e2-8b93-4c69-b335-62a917ab6e8f · outbound

This paper cites Qwen3.5-Omni Technical Report.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Qwen3.5-Omni Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.883844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.883844Z digest=sha256:3f59161e1f9fb8a76ed3c852192a77ccbe598a10cb119720a4845c12cf5eb920

Observation 4b29d5af-d0d7-4469-8ab1-0f1a0e14ed1c · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.arXiv preprint arXiv:2511.21395, 2025.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Monet: Reasoning in latent visual space beyond images and language.arXiv preprint arXiv:2511.21395, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.888273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.888273Z digest=sha256:d1564ce0c76aedb306a0e77462f54b52929b9e5e299b331f08715f1ac1763d43

Observation 6ac6f9ca-05d6-47f1-8765-54db3d9cdeb0 · outbound

This paper cites Beacon: Knowing when and how to perform agentic visual reasoning, 2026.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Beacon: Knowing when and how to perform agentic visual reasoning, 2026

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.395159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.892377Z digest=sha256:07da04f4e3d6c3b5b73445b7fc276422454968af912e39a10d26f7871f0850b1

Observation 8282b130-6e75-4184-bba7-5950cd3130ce · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.896454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.896454Z digest=sha256:6e2bbf70b35b3f6ab0eb93c2f56e09dccadcb72140d3a043ad31e8a469d604c6

Observation 701a70f1-3fc6-4030-b6a8-db64f3001a8d · outbound

This paper cites Varcmp: Adapting cross-modal pre-training models for video anomaly retrieval.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Varcmp: Adapting cross-modal pre-training models for video anomaly retrieval

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.264426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.900700Z digest=sha256:d8d2c3b8c6a875a6da9085d3c1361cbb6017ffecd6045b83336cc5153cd42c3c

Observation f595ffb2-ac58-4259-9c6b-752dcc24693c · outbound

This paper cites Avadclip: Audio- visual collaboration for robust video anomaly detection.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Avadclip: Audio- visual collaboration for robust video anomaly detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.136783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.904729Z digest=sha256:dee5d6c08c5ed34b8fd564b05d5459c67d83fe37e4010e1f7d4d466962acfb1d

Observation 9aa36018-c1f7-4156-8b9e-ad6fa986daaa · outbound

This paper cites Stage-adaptive Token Selection for Efficient Omni-modal LLMs.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Stage-adaptive Token Selection for Efficient Omni-modal LLMs

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:00:11.190235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.908725Z digest=sha256:dbbf0d2a41945c74d65dbb30ab01a6340c5737bf8f327eee4a7d3333f614445c

Observation e7cfe3c4-3dc9-4fa6-a3b7-e8a13743fa6b · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.913314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.913314Z digest=sha256:a9ec1b71ace93ff97f23bd8b86d811d2c58dc23758de0db1f45ce755612cff2d

Observation 06452d4f-c1a8-44f1-95eb-50b1c332fe3c · outbound

This paper cites Qwen2.5-Omni Technical Report.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Qwen2.5-Omni Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.917458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.917458Z digest=sha256:d0f3a77e31fa03ba8a82e8b49af6cfa60750c190247b943c0342b49c6add9754

Observation f8895f65-c30c-4837-8fe5-590202e4c06b · outbound

This paper cites Qwen3-Omni Technical Report.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Qwen3-Omni Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.922366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.922366Z digest=sha256:85e6fe779427fe5e8875f39c7cf550155cd8e183bbaffacfe1e8253ca7a7d365

Observation 5ff0a44c-6bed-496d-a2cd-b373b08cf750 · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.926459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.926459Z digest=sha256:b12ed2c7af8f7856c41dfbcdf9d3d4b43b1d98e8b612549659344d6926dc235d

Observation cba8c6a1-d9af-4899-9b7e-b884074ff5dc · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Visionzip: Longer is better but not necessary in vision language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:12.971267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.930514Z digest=sha256:d385d9c72cff54b99477a5bcaed65b9fd9289e2b749a7a3e932e6d30ce323131

Observation 9641e354-286a-4a75-9b54-1649e8272a96 · outbound

This paper cites Audio-centric video understanding benchmark without text shortcut.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Audio-centric video understanding benchmark without text shortcut

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:12.826632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.934466Z digest=sha256:8e2a77ac00b27cbf1dd632eacf6276d0fe936986eb43b713f48df5d8d72d858c

Observation efb0d17f-458b-40e1-ab02-c7e005dbd68c · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.938425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.938425Z digest=sha256:7e5e4f89977b07bfd26cc803736d0c8ef216cf073c7200632522649044322272

Observation 6589d0ef-0127-47bb-b8c3-cffbb2fe0e68 · outbound

This paper cites Lmms-eval: Re- ality check on the evaluation of large multimodal models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Lmms-eval: Re- ality check on the evaluation of large multimodal models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:12.680178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.942553Z digest=sha256:26dbc9fd76b57fb01b07c26043219ea35cbda298cd65b066490846d09c4a21a4

Observation ab4978fd-8478-4d6f-b98e-8b08c42cb70e · outbound

This paper cites Debiasing multimodal large language models via penal- ization of language priors.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Debiasing multimodal large language models via penal- ization of language priors

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:12.551861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-05T12:00:10.946729Z digest=sha256:0f7ca85e98e591baf40130d71817315ae837d27f44897cbb4ce225fc80fa8051

Observation 6b95c888-b354-4cab-b365-f2adae77069a · outbound

This paper cites What is this?.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models What is this?

Reference 52

Resolution
malformed identifier
no resolver link, observed 2026-08-05T12:00:10.950717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.950717Z digest=sha256:a902d41ce5de127880b73d086245da0634ec89b280efbc975b1cc3def032d512

Pith citing papers

No inbound Pith citation observations are available.