Pith. sign in

Paper Citation Record · LEDGER

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models

As of 8 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2608.03812.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03812 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:00:10.950717Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact4
  • verified fuzzy22
  • unresolved25
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c4ff1c81-251e-4035-8a6d-98e3d0f22594 · outbound

This paper cites Qwen3-VL Technical Report.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.728615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.728615Z digest=sha256:c23c6e433fb0dbcef4d2cfe31ba4642a479c99daeb1ef0d78b19e9d39ba6bc05

Observation f2a2490a-d605-4a04-b8aa-288cb3f7a3f2 · outbound

This paper cites Token Merging: Your ViT But Faster.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Token Merging: Your ViT But Faster

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.733972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.733972Z digest=sha256:992e52cc134b5f6bf2bb7c5d571a5de991fb003a17155af1ac6b06795bb4f736

Observation 85a71cfc-c61f-427f-bb6c-9328324d3d24 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.225663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.739333Z digest=sha256:176e68cc217ccf716299651df9b0edd7e595f09ce946525b86ba5a12184c582e

Observation 6b8993aa-07a6-4baf-91de-9cc073e2eb77 · outbound

This paper cites Avocado: An audiovisual video cap- tioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Avocado: An audiovisual video cap- tioner driven by temporal orchestration.arXiv preprint arXiv:2510.10395, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.744559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.744559Z digest=sha256:e4608b217e58d2a387858be452648764810426e47a9fbc18624a79e5aadf26b3

Observation fc11dee5-e447-40a7-bf99-f86c01531b48 · outbound

This paper cites MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.749076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.749076Z digest=sha256:aa4770d35a033bf66b96b231196d9bb326a43459856fb5f8573957c202721877

Observation adf782b0-604e-4e06-ac09-592868da529c · outbound

This paper cites OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models OmniRefine: Alignment-Aware Cooperative Compression for Efficient Omnimodal Large Language Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:00:12.324177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.753980Z digest=sha256:15a0de828156cff54347e349313db3cf90ea97a394b0baf808642ad9e833f302

Observation d4c6e423-058f-4fc6-a23b-65cb6784fea9 · outbound

This paper cites OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.759361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.759361Z digest=sha256:732decc733016c63a071046f92a1463b8077ea7fa4b11d226e6bff5fcf092917

Observation b6767fe0-689d-4823-a9eb-22133b8c728e · outbound

This paper cites Unified spatiotemporal token compression for video-llms at ultra-low retention.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Unified spatiotemporal token compression for video-llms at ultra-low retention

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.212702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.763940Z digest=sha256:95677c5d9517ad37107640aa9bf525f67925ca327314a5009b20fa00f023ad60

Observation 0aa1385b-96f6-4d00-897b-e93558c66bde · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis.Knowledge-Based Systems, 2016.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Study on density peaks clustering based on k-nearest neighbors and principal component analysis.Knowledge-Based Systems, 2016

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.199415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.768085Z digest=sha256:34085a6a16c47814aebd1fb554b57ef10333d2b5d251a5d3fc7750118c0f103f

Observation c7229e11-888c-461e-bb5b-e131dff37ef1 · outbound

This paper cites Flashvid: Efficient video large lan- guage models via training-free tree-based spatiotemporal to- ken merging.arXiv preprint arXiv:2602.08024, 2026.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Flashvid: Efficient video large lan- guage models via training-free tree-based spatiotemporal to- ken merging.arXiv preprint arXiv:2602.08024, 2026

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.772364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.772364Z digest=sha256:ff6614ee6e78eb57b51c8611a053911a9893930a203c0a67caeb49e6d1a4e446

Observation 1563059e-d21b-4e68-84a5-fffc1a5b873d · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.776575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.776575Z digest=sha256:c4f28517e13e877cfe569da7c537b390926387d80e06a3a506aeca10f28bdfe8

Observation 85f39ece-efb9-47a9-8589-866d897e697f · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.187016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.781252Z digest=sha256:82c927a09ab4ff297c859f2842179d7d51d36fb10ee69e333b39f6b1fadc5023

Observation df338285-baa2-473b-a284-b677b5a98a8e · outbound

This paper cites ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.785575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.785575Z digest=sha256:04bd60a306e58844629bb9fdc56d0c1209877e409c2bb9c45841de65717f62c0

Observation 8724ec8e-e5ea-457b-b42d-53d63a4e1e83 · outbound

This paper cites EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.790061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.790061Z digest=sha256:a89ea52e405b27e262b57c1b3cd3647f663b8d3984e344f272c96176556a91c1

Observation e80f3558-050b-468e-920d-7cfe665d1e4b · outbound

This paper cites Echoingpixels: Aliasing-resistant joint token reduction for audio-visual llms.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Echoingpixels: Aliasing-resistant joint token reduction for audio-visual llms

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.173519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.794328Z digest=sha256:375fe7591eacbf41fe18e574c0a7e7ee0340b0ebd7653287b3da9f240808998e

Observation 9e4c7575-aca0-494a-8d10-8be0b01266c7 · outbound

This paper cites Gemini 3.1 pro model card.https: / / deepmind.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Gemini 3.1 pro model card.https: / / deepmind

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.160700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.798302Z digest=sha256:182f3de8a6e0f5b0c803a68934bbf64f18b60b2e93007766c71491d09b5e566a

Observation f691b94f-fdd3-45d0-82b7-fb384736a7ad · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.802417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.802417Z digest=sha256:5f9330730c1dd2049307b08c7da52a40b790b392b6e244ea5c259d5d804adc5c

Observation 85fb0f3a-968e-4add-bd8b-17f03ba460b7 · outbound

This paper cites GPT-4o System Card.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.806834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.806834Z digest=sha256:fef21edfc8573310cff4313e9e4af64803385304a2876c6c4a50b4490c75440c

Observation 669715a8-0864-4eda-a1c1-70f44e032344 · outbound

This paper cites ContextGuard: Structured Self-Auditing for Context Learning in Language Models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models ContextGuard: Structured Self-Auditing for Context Learning in Language Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:00:11.978204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.811574Z digest=sha256:48145b259dedacc6b2a40658372afd04793e61bd0824aceb117dbe69e79cb648

Observation bb040714-fcd5-48b0-a191-969f1ff2c6ce · outbound

This paper cites Token pruning in audio trans- formers: Optimizing performance and decoding patch im- portance.arXiv preprint arXiv:2504.01690, 2025.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Token pruning in audio trans- formers: Optimizing performance and decoding patch im- portance.arXiv preprint arXiv:2504.01690, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.815853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.815853Z digest=sha256:927b30b8a8ef6e9539d996862760be0c8be93e456ad26ca1b5d1d28e07c20e65

Observation 969fe028-3840-4a74-9ff4-ab0d84dfadbc · outbound

This paper cites Omnivideobench: Towards audio-visual understanding evaluation for omni mllms.arXiv preprint arXiv:2510.10689, 2025.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Omnivideobench: Towards audio-visual understanding evaluation for omni mllms.arXiv preprint arXiv:2510.10689, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.820110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.820110Z digest=sha256:c9b366af9c00d4f1ed52930ae4f9fda9ee8f6da1d2a2c4b1a52c865b44252407

Observation 2a535e23-31cd-43b0-b60a-cd5745492821 · outbound

This paper cites OmniGAIA: Towards Native Omni-Modal AI Agents.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models OmniGAIA: Towards Native Omni-Modal AI Agents

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.824327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.824327Z digest=sha256:ceb2404d459e0ad36bb8641c113ccc3b692c84ba37aa2609aeb86aa7a4d48c76

Observation 02549923-0200-415b-82f1-2e61ecf47e87 · outbound

This paper cites Speech- prune: Context-aware token pruning for speech information retrieval.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Speech- prune: Context-aware token pruning for speech information retrieval

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.147517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.829073Z digest=sha256:9b6b967be44cfa82d4f242be76e52fdee56ba9896b852e85279902dd29cd0a36

Observation ec615e69-17b1-4750-9904-d44b2db9a144 · outbound

This paper cites Video compression commander: Plug-and-play inference ac- celeration for video large language models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Video compression commander: Plug-and-play inference ac- celeration for video large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.134724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.833105Z digest=sha256:bf4c4ac2c9183d46c3ae0063feccdc58cef828afe091cab18398b39bb534c17f

Observation 1e67b4fa-d291-4650-8def-ea6526ef4f2d · outbound

This paper cites Gpt-5.5 system card.https://openai.com/ index/gpt- 5- 5- system- card/, 2026.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Gpt-5.5 system card.https://openai.com/ index/gpt- 5- 5- system- card/, 2026

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.121235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.837170Z digest=sha256:b1b7a76851c4996a67e7cdb27d437b2007ecaf2b9be6b6293e8f47a89508c00b

Observation c0be0490-6173-4d6f-bfd8-bde6e8008139 · outbound

This paper cites OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models OmniDrop: Layer-wise Token Pruning for Omni-modal LLMs via Query-Guidance

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:00:11.658798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.841260Z digest=sha256:2585267c964aa2a6c0df28cdcdb53bcc12a4ffd98df320634b6dd41698407358

Observation 01735dad-625d-4a25-ba0b-57e432753fd7 · outbound

This paper cites Clustering by fast search-and-find of density peaks.science, 2014.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Clustering by fast search-and-find of density peaks.science, 2014

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.108069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.845601Z digest=sha256:60d40d12ff3ddaa4b1cea1aaf2cbf89a8e054392c8567fa37def6ad3a9dbf3e8

Observation c7e7f61a-fc50-4892-8ba9-396f3c1eacd7 · outbound

This paper cites Fastvid: Dynamic den- sity pruning for fast video large language models.Advances in Neural Information Processing Systems, 2026.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Fastvid: Dynamic den- sity pruning for fast video large language models.Advances in Neural Information Processing Systems, 2026

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:14.095057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.849604Z digest=sha256:2948e47a0a45db0705b53643b6cd3b067dd684e3750e978c0c1535dfc1c240c2

Observation a797368c-461c-4461-b661-4f748b721c06 · outbound

This paper cites Mavors: Multi-granularity video representation for multimodal large language model.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Mavors: Multi-granularity video representation for multimodal large language model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.988394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.853895Z digest=sha256:dc7d5906d80a0fe31283501f48bd73c2f71d8071a956d7b8083e4c72eb4e0650

Observation 1dad8e06-66ef-42ec-b8ef-8d8146649fb1 · outbound

This paper cites Audio- visual llm for video understanding.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Audio- visual llm for video understanding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.846338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.857877Z digest=sha256:69385b816bb29fd0a1c7efdd4ae9eddc759bc11b29578d4a4fd20d5cfbd0968f

Observation 98885e3e-5829-4942-9e14-7b991be14341 · outbound

This paper cites TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.862282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.862282Z digest=sha256:47fc53cc0aab99a10f957f848ace64a2426d694970b899f732feb5853817445b

Observation 5ca6b901-294a-4feb-ad39-e4f4a9168434 · outbound

This paper cites video- salmonn 2: Caption-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models video- salmonn 2: Caption-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.866665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.866665Z digest=sha256:2c2945e679cae0f53612376d388f11ec4ba053c268672f90e89820bb469e687b

Observation 58d07e37-2acf-49f2-b954-dc6c6718216e · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast 9 video large language models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Dycoke: Dynamic compression of tokens for fast 9 video large language models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.663031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.871020Z digest=sha256:7b9f0113ac6dd24669ab3e93b7bd156cf6cddafb064473c4cc16edb97c8a4c13

Observation e7911a0d-46d9-46a3-a5ed-b0260e4b4a63 · outbound

This paper cites Omnizip: Audio-guided dynamic token compression for fast omnimodal large language models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Omnizip: Audio-guided dynamic token compression for fast omnimodal large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.509379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.875168Z digest=sha256:39f39aa04fae885f25b95db80e7ef6b2f09ce8b1e421b9c721dd6611e9cd456a

Observation f599502c-6604-45db-8365-c1e444359355 · outbound

This paper cites Lvomnibench: Pioneering long audio-video un- derstanding evaluation for omnimodal llms.arXiv preprint arXiv:2603.19217, 2026.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Lvomnibench: Pioneering long audio-video un- derstanding evaluation for omnimodal llms.arXiv preprint arXiv:2603.19217, 2026

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.879590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.879590Z digest=sha256:134d2bbbb54769f37995bdfde2fa447f9392ba161ba27030d0a5c67f3ae00858

Observation 4f9388e2-8b93-4c69-b335-62a917ab6e8f · outbound

This paper cites Qwen3.5-Omni Technical Report.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Qwen3.5-Omni Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.883844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.883844Z digest=sha256:aa30ed72a889a67318fb054c4cc7e322edd63582e16d32662e6847a6511de683

Observation 4b29d5af-d0d7-4469-8ab1-0f1a0e14ed1c · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.arXiv preprint arXiv:2511.21395, 2025.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Monet: Reasoning in latent visual space beyond images and language.arXiv preprint arXiv:2511.21395, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.888273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.888273Z digest=sha256:a3a25d358ab6fe2c6b20b2a5f5247f4e486e4b89585e72a47b687f6d2e77d292

Observation 6ac6f9ca-05d6-47f1-8765-54db3d9cdeb0 · outbound

This paper cites Beacon: Knowing when and how to perform agentic visual reasoning, 2026.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Beacon: Knowing when and how to perform agentic visual reasoning, 2026

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.395159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.892377Z digest=sha256:ee0f43414d03d0c422f507c81a66fba66580fcdd0c031f8827ccaddb4d0f30b3

Observation 8282b130-6e75-4184-bba7-5950cd3130ce · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.896454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.896454Z digest=sha256:9cafcbff50b4edaeae711bd83fa65beb756e97fd228092f66737b8f53f4a92a5

Observation 701a70f1-3fc6-4030-b6a8-db64f3001a8d · outbound

This paper cites Varcmp: Adapting cross-modal pre-training models for video anomaly retrieval.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Varcmp: Adapting cross-modal pre-training models for video anomaly retrieval

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.264426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.900700Z digest=sha256:a537eb7d23afa070df3463796d82b372331b44791e7950f80977180eb1c76544

Observation f595ffb2-ac58-4259-9c6b-752dcc24693c · outbound

This paper cites Avadclip: Audio- visual collaboration for robust video anomaly detection.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Avadclip: Audio- visual collaboration for robust video anomaly detection

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:13.136783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.904729Z digest=sha256:c848e86012a2ded0fd2bc9f67a53e6b2e1cdadd7c671c29af4abb15a9fa780be

Observation 9aa36018-c1f7-4156-8b9e-ad6fa986daaa · outbound

This paper cites Stage-adaptive Token Selection for Efficient Omni-modal LLMs.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Stage-adaptive Token Selection for Efficient Omni-modal LLMs

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:00:11.190235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.908725Z digest=sha256:8c42d124abc3482c93e3dc0cc504064949cea332ecedd88bb2e912f21914b4c0

Observation e7cfe3c4-3dc9-4fa6-a3b7-e8a13743fa6b · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.913314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.913314Z digest=sha256:76d6dba6e847a25389630ee5f8d67cbbae0dc20dd00ba8c74e15aae7a464a8b8

Observation 06452d4f-c1a8-44f1-95eb-50b1c332fe3c · outbound

This paper cites Qwen2.5-Omni Technical Report.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Qwen2.5-Omni Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.917458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.917458Z digest=sha256:6a16af93fc05bdf68bccd74e266ad33bc7ed7819af4781f26c29fc7cb2f79110

Observation f8895f65-c30c-4837-8fe5-590202e4c06b · outbound

This paper cites Qwen3-Omni Technical Report.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Qwen3-Omni Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.922366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.922366Z digest=sha256:64386b0f48629c02229c6f11d7c6bf3e06d00f70beb02b14369a48b8b773f295

Observation 5ff0a44c-6bed-496d-a2cd-b373b08cf750 · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.926459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.926459Z digest=sha256:b820ae61fb411126fea43cc03f078e9575b460d9cd3515dccb0c7cc99cf09fc1

Observation cba8c6a1-d9af-4899-9b7e-b884074ff5dc · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Visionzip: Longer is better but not necessary in vision language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:12.971267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.930514Z digest=sha256:e1b3077fb7712675339a640019e297c6977c1bc93c39976696231396e7d39f99

Observation 9641e354-286a-4a75-9b54-1649e8272a96 · outbound

This paper cites Audio-centric video understanding benchmark without text shortcut.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Audio-centric video understanding benchmark without text shortcut

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:12.826632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.934466Z digest=sha256:4121fff8aacca2e85c56af299ae7a3969f92567780515a76bc79cc14087a5354

Observation efb0d17f-458b-40e1-ab02-c7e005dbd68c · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:10.938425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.938425Z digest=sha256:99d6a8cf8cd1ac5f3313df04aec71abd0e77011455b547813462fabed2a10395

Observation 6589d0ef-0127-47bb-b8c3-cffbb2fe0e68 · outbound

This paper cites Lmms-eval: Re- ality check on the evaluation of large multimodal models.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Lmms-eval: Re- ality check on the evaluation of large multimodal models

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:12.680178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.942553Z digest=sha256:dca0ac2022935082ae172eaf112f552b25bb4b13ad42dda112cae222e38accad

Observation ab4978fd-8478-4d6f-b98e-8b08c42cb70e · outbound

This paper cites Debiasing multimodal large language models via penal- ization of language priors.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models Debiasing multimodal large language models via penal- ization of language priors

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:00:12.551861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T12:00:10.946729Z digest=sha256:90950525fa48c0f667025263381bf646781907fd1a0d7bee3df010d79cf8d10b

Observation 6b95c888-b354-4cab-b365-f2adae77069a · outbound

This paper cites What is this?.

OmniPack: Unified Token Compression for Efficient Omni-modal Large Language Models What is this?

Reference 52

Resolution
malformed identifier
no resolver link, observed 2026-08-05T12:00:10.950717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:10.950717Z digest=sha256:0aba427148f78492ebc1ba78e4073b1d93493f8d6d1738881a994718c3a24a99

Pith citing papers

No inbound Pith citation observations are available.