Pith. sign in

Paper Citation Record · LEDGER

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models

As of 10 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2607.23193.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23193 v3

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:02:54.529833Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 321e6400-3136-4806-92de-a7d1d23f542e · outbound

This paper cites Token Merging: Your ViT But Faster.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Token Merging: Your ViT But Faster

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.403845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.403845Z digest=sha256:445c76fc20decba40b88d0f80d72d4ab490de6cbfe356da6117431bf8261ee19

Observation e0f66cb5-7bf5-47f5-ba9c-fdd856498ff2 · outbound

This paper cites Animageisworth1/2tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Animageisworth1/2tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.407268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.407268Z digest=sha256:a6c494d24275ae9a528d0f404dedf44bf2738831f4c80f5a4627d4219379a13b

Observation 5d09a710-17d7-4815-991c-d667f60a561b · outbound

This paper cites Howfararewetogpt-4v? closingthegaptocommercialmultimodalmodelswithopen-sourcesuites.ScienceChina Information Sciences, 67(12):220101, 2024.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Howfararewetogpt-4v? closingthegaptocommercialmultimodalmodelswithopen-sourcesuites.ScienceChina Information Sciences, 67(12):220101, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.409945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.409945Z digest=sha256:89846d4c0407616fa3df254688a9997e15956a1e21a6a72da4ffa61b5fe6d4d5

Observation e9de0de9-3c6d-481f-834f-0342e2aa925e · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.Advances in neural information processing systems, 36:49250–49267, 2023.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Instructblip: Towards general-purpose vision-language models with instruction tuning.Advances in neural information processing systems, 36:49250–49267, 2023

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.412497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.412497Z digest=sha256:7f63370c1e2512dd8ce5a1115e0f9896530774b4b64620857e6088e5d5d18af4

Observation 5ff12260-6c83-4bfa-aff2-95f38fdb89ff · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.414925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.414925Z digest=sha256:e967491ffbcf6422513719c4abbcde5d60f4d38fab8433ef1035cc00eda65d77

Observation dbd625f2-1585-45a4-825d-d7a855029a5d · outbound

This paper cites OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.417708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.417708Z digest=sha256:8e4c7f584ff1b35bbdce3e6946a762cc3b0d70ff0f4416ead9794674b769442c

Observation 42b3b869-dc30-467d-8c87-279cf1c3c586 · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis.Knowledge-Based Systems, 99:135–145, 2016.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Study on density peaks clustering based on k-nearest neighbors and principal component analysis.Knowledge-Based Systems, 99:135–145, 2016

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.420658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.420658Z digest=sha256:56fd441f250c5f7eded5e37e764456ffbc21a854b0093e24cba2f0f8a4550075

Observation 4de105d5-329a-4524-a97d-03d201a4c5bc · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.422877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.422877Z digest=sha256:2366ef717c316b12e2de091c29d2696485bbf756423169239640bbec188d4031

Observation 059812d1-8ab1-4999-9197-160ea2f0323d · outbound

This paper cites Video-mme: Thefirst-evercomprehensiveevaluationbenchmarkofmulti-modalllmsinvideoanalysis.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Video-mme: Thefirst-evercomprehensiveevaluationbenchmarkofmulti-modalllmsinvideoanalysis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.425121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.425121Z digest=sha256:30d0498d002ec746aef00396b6fc4dcfae45308179262f3c9de2efa8ab5ac76a

Observation bfdb2d56-db0f-4f24-bc84-22ffd3c0f001 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.427425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.427425Z digest=sha256:be3e3b88eb4c47d95f3999358f6ae401f45c3434fb0abc366b9fcb6d86af5fe8

Observation fdc2347b-b2c4-413b-98e5-97eddb1c6c81 · outbound

This paper cites WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.430169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.430169Z digest=sha256:d144eb21ef4a46d14fd474fff26ff909cfabd6c0a33b1192badeb268cd312710

Observation f5cc7892-3b71-4048-9417-80df2d7bb9ac · outbound

This paper cites Prunevid: Visualtokenpruningforefficientvideolargelanguagemodels.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Prunevid: Visualtokenpruningforefficientvideolargelanguagemodels

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.432700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.432700Z digest=sha256:3ab51c28d6f8b8457917055a146efc859719b94411d0a08b7cf899176dd6eabb

Observation 9109553f-81e9-41e3-b035-1c8b65ce404d · outbound

This paper cites GPT-4o System Card.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models GPT-4o System Card

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.434801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.434801Z digest=sha256:3d130dad8e6b047dbd0d036433a43d02df38001826a644126838a4a7a560def7

Observation 15f6d9a0-23c7-41ca-aebf-0b81797b5571 · outbound

This paper cites Efficient multimodal large language models: A survey.Visual Intelligence, 3(1):27, 2025.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Efficient multimodal large language models: A survey.Visual Intelligence, 3(1):27, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.437064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.437064Z digest=sha256:74c26b8d93ed28f95c887d718d550d98e0ac087c49ef928547a6e852b9d56261

Observation 3f2fadfb-0fbd-453b-b8eb-523d0aa91ac5 · outbound

This paper cites Tokenpruninginaudiotransformers: Optimizingperformanceanddecodingpatchimportance.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Tokenpruninginaudiotransformers: Optimizingperformanceanddecodingpatchimportance

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.439280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.439280Z digest=sha256:514ff8e4b548f8db60181fec448f19d8e03d1578602cb49f014a958bbfc943c0

Observation 9abbf839-5c61-43e4-b63c-45e52ad589eb · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.441373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.441373Z digest=sha256:e8d9e9966a8481cccd558a0aafc199e19a3480f5ac8c831155745f2f4a0bbea8

Observation a1da3b83-dd4a-4444-bab5-dd0b8794c78d · outbound

This paper cites Omnivideobench: Towards audio-visual understanding evaluation for omni mllms.arXiv preprint arXiv:2510.10689, 2025.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Omnivideobench: Towards audio-visual understanding evaluation for omni mllms.arXiv preprint arXiv:2510.10689, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.443570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.443570Z digest=sha256:8c166771786531c67ee671218dcde9de658245ccccd8d7866374ac047cfe926c

Observation 4d4a87ba-6e4e-4cfb-8dd4-4b94e61d93cd · outbound

This paper cites Accelerating Transducers through Adjacent Token Merging.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Accelerating Transducers through Adjacent Token Merging

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.445710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.445710Z digest=sha256:75680adceeaef549a8446919303d2371001083c7b4f39e5cbff53e4dae1b9028

Observation 84198396-a7f8-490a-8c42-3d7f356f78eb · outbound

This paper cites Video-llava: Learningunitedvisualrepresentation by alignment before projection.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Video-llava: Learningunitedvisualrepresentation by alignment before projection

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.448375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.448375Z digest=sha256:aaad7b14b686b905f4f78204b1733ff9bd3f81335d4b0926c757fd58b26bb61a

Observation 3dc2ff65-e3aa-4421-a9f2-7a88e5ea5509 · outbound

This paper cites Speechprune: Context-awaretokenpruningforspeechinformationretrieval.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Speechprune: Context-awaretokenpruningforspeechinformationretrieval

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.450552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.450552Z digest=sha256:5ef94f372a7592010bacaa3a8485613a7c5dc35f2b3592b9328493a3078d9376

Observation 16458f34-6e67-4409-b267-c34321a208ce · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.452730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.452730Z digest=sha256:069cfaab6440b56b2eb2a431eaa05e39a8ac4e4e8afaf80a819a7113aa465b23

Observation 002db1f9-8ec1-4c54-8963-c763ae79e93c · outbound

This paper cites Javisgpt: A unified multi-modal llm for sounding-video comprehension and generation.arXiv preprint arXiv:2512.22905, 2025.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Javisgpt: A unified multi-modal llm for sounding-video comprehension and generation.arXiv preprint arXiv:2512.22905, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.455019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.455019Z digest=sha256:c78c38c6fff16322357c40beaf56e87a602b3dde5f1e7b129ba7fc5371c18c70

Observation 17d987d6-eb0f-4b9c-94fe-1528559f14d7 · outbound

This paper cites Quota: Query-orientedtokenassignmentviacotquerydecoupleforlongvideocomprehension.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Quota: Query-orientedtokenassignmentviacotquerydecoupleforlongvideocomprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.457179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.457179Z digest=sha256:6b723441b3f79a76ec2fa503cb2b8849ded6574f15566f32cede1e534dc793cb

Observation 34f7c4df-4658-4ab2-a502-5c223a79f474 · outbound

This paper cites Llm-pruner: On the structural pruning of large language models.Advances in neural information processing systems, 36:21702–21720, 2023.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Llm-pruner: On the structural pruning of large language models.Advances in neural information processing systems, 36:21702–21720, 2023

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.459406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.459406Z digest=sha256:28463c5f6ff713e5a097632d3d91ebae9bd7a1b99a3a0897faf86e32366b0346

Observation 9fdf5772-7ece-4cb8-9b19-6f222f231419 · outbound

This paper cites Ompq: Orthogonalmixedprecisionquantization.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Ompq: Orthogonalmixedprecisionquantization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.461676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.461676Z digest=sha256:b0529eb8f190776508371d3eaa1f5885c4b8ffe7a49267a2e46532306cdf0665

Observation 18f66ed9-9459-48ff-b97a-2420bd797493 · outbound

This paper cites Affinequant: Affine transformation quantization for large language models.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Affinequant: Affine transformation quantization for large language models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.463908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.463908Z digest=sha256:003724ed1fc94bc78d1ca5caba8df96c22ea86774cca909da7549a7c530dc190

Observation 7dc8079b-e33c-4b4b-b4da-c8c00ba098c9 · outbound

This paper cites Norm of word embedding encodes information gain.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Norm of word embedding encodes information gain

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.466408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.466408Z digest=sha256:9efd0b1fc6b39e709f4e980aa95557c2887477f6c0cd0a89d749c46e2a7ee7c8

Observation c0977b41-472a-41a4-960c-2fa59eb9fb87 · outbound

This paper cites Prentice-Hall, Inc., 1993.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Prentice-Hall, Inc., 1993

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.468638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.468638Z digest=sha256:129add69b7b16239abc4d87bc03ce0cf3a47e66819e2db0963062387d25b0b54

Observation 72f6cacd-4bba-4e32-95a1-0fe2c33d9367 · outbound

This paper cites Learning transferable visual models from natural language supervision.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Learning transferable visual models from natural language supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.470907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.470907Z digest=sha256:e76ec694b7a6dd838374ecfb5c75cab16c910ae53ac9d1dd57cd399e8f3a41c6

Observation 45082798-ef54-40bc-a8c1-fd0ebd7be22d · outbound

This paper cites Llava-prumerge: Adaptive token reduction for efficient largemultimodalmodels.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Llava-prumerge: Adaptive token reduction for efficient largemultimodalmodels

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.473156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.473156Z digest=sha256:4e72e815c386842c73b45850f4ebfa817de6767371e825b24885049f5d5dcba5

Observation 59f4436a-e2e2-4915-9584-7b39520743f4 · outbound

This paper cites Holitom: Holistictokenmergingforfastvideolarge language models.arXiv preprint arXiv:2505.21334, 2025.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Holitom: Holistictokenmergingforfastvideolarge language models.arXiv preprint arXiv:2505.21334, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.475377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.475377Z digest=sha256:db17460ea1e2be26461885f3b5c4460b706bfdbecd6cfce92305e1e7b7d6a189

Observation e13f7504-37f7-471b-9130-247bf434fc4f · outbound

This paper cites Whentokenstalktoomuch: Asurveyofmultimodallong-contexttokencompressionacrossimages,videos,andaudios.arXiv preprint arXiv:2507.20198, 2025.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Whentokenstalktoomuch: Asurveyofmultimodallong-contexttokencompressionacrossimages,videos,andaudios.arXiv preprint arXiv:2507.20198, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.477789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.477789Z digest=sha256:51fcc908229b2f86d530307d181359ded0178b2fa3a64bcd48e4e0ecfcd920d5

Observation 202b797a-e3db-4576-8b2a-c09a7bd4aafe · outbound

This paper cites Less is more: A simple yet effective token reduction method for efficient multi-modal llms, 2024.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Less is more: A simple yet effective token reduction method for efficient multi-modal llms, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.480165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.480165Z digest=sha256:8c98207f9610747c0eb22979060cb7a432548599d3b7c50050945eefa0bd44b1

Observation 9942ec5b-7557-407f-b7d9-2993050d133e · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.482439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.482439Z digest=sha256:182cf0ebd9c00c6748f98d4c7890cb407ed403fad2928a52630faeda1e247be3

Observation 8cb3ad99-2e98-4a33-883d-366a444027cb · outbound

This paper cites Lvpruning: Aneffectiveyet simple language-guided vision token pruning approach for multi-modal large language models.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Lvpruning: Aneffectiveyet simple language-guided vision token pruning approach for multi-modal large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.485107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.485107Z digest=sha256:bb35b8fe5aa0c62c228902ce1fa8ddc2d4f70d1b6b2fd51b776a86d2ad47b489

Observation 389089a4-dba0-4775-9416-aeb0594fe69c · outbound

This paper cites video-salmonn2: Caption-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models video-salmonn2: Caption-enhanced audio-visual large language models.arXiv preprint arXiv:2506.15220, 2025

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.487469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.487469Z digest=sha256:8e4e22a83e22fcae198e8e88742b94aac5b1806dbdded6c5dd05b37d9efe1016

Observation f1abf4e1-5c5e-419c-b352-268351089ee5 · outbound

This paper cites Dycoke: Dynamic compression of tokens for fast video large language models.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Dycoke: Dynamic compression of tokens for fast video large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.489634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.489634Z digest=sha256:79dcff2a58b1ba00e7456706675270b0db003d2641ca9b4a761614bff73962cc

Observation 80f3f3d0-320c-447f-b65e-a8f897e7220e · outbound

This paper cites OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.491885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.491885Z digest=sha256:b9c055b87c6ce374729fdc230c77963f93323ba4a6f4c5c8f81eb19fdeb0f254

Observation d90096e6-6bc9-4058-8128-d9105b6d4946 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.494236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.494236Z digest=sha256:dafb932f2ad581cd280ac3bdaf0b1c5603d739ae4baae223549a34d8dac5c776

Observation d68d8d24-61c0-471f-8d6a-fa0e68d8048d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.496554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.496554Z digest=sha256:cbd24d8e60653daf8dedd1f872ead8241834b610cef83a18f2c16c6a91d963cc

Observation 150dbaaf-1f2a-4d27-9ad0-b56e483c3e4e · outbound

This paper cites Longvlm: Efficientlongvideounderstandingvia large language models.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Longvlm: Efficientlongvideounderstandingvia large language models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.499048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.499048Z digest=sha256:bc00bc6caa84861932733b48f85acca7ee75130a2846fb44887d7aa30aa93918

Observation 2a297e47-a33d-4b46-b9ba-98aefa4dca7b · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.501212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.501212Z digest=sha256:74468d13a1fbd068ec8f884e620c9d641bac1f5cfd4cbffcb60fbc70ebd2f18d

Observation 8f0b1861-64ad-44e0-814a-c39585874b8d · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.503571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.503571Z digest=sha256:9dc27a233ef067b2ab84949bb188ecee93b523442926443a17762edc3c64f05c

Observation e84d5c24-21e8-461e-969a-542e51c30061 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.505966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.505966Z digest=sha256:a8be90ff3bf555ceb29292a721e7fe37f7b89f300a1d926e38f11bee97aa0b80

Observation a881aeeb-654d-41b9-aa74-bd3ab75d9706 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.508465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.508465Z digest=sha256:5a7282d9f8c2c828cfe78ac0c57e71cb6511a6eff0cbe93f747bbf7ec3e17584

Observation 2813c016-831d-41fa-9431-b84d010c32e4 · outbound

This paper cites Qwen2.5-Omni Technical Report.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Qwen2.5-Omni Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.510858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.510858Z digest=sha256:af4781725ffeaf10cf2350bc1f4d7c3268ae9a0717b912c96f3df6320245c361

Observation 41fde0e2-583f-4cd2-8245-5a01599fd579 · outbound

This paper cites Qwen3-Omni Technical Report.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Qwen3-Omni Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.513476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.513476Z digest=sha256:5a879f5ceb170d888b6cb2406b9c2cb67190079f1f62f6136690de9486bad28b

Observation 2bcccdc2-9feb-42fb-acce-f629ddb82464 · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.515743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.515743Z digest=sha256:0f16b3ed46b1248d5a863c9c4c841c2dad2b064a11f6269d12e98bafbf62031b

Observation bdb89aad-2a96-41af-aa93-e67ca05fd351 · outbound

This paper cites Visionzip: Longer is better but not necessary in vision language models.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Visionzip: Longer is better but not necessary in vision language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.518149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.518149Z digest=sha256:c282e6f897962819dd7f86a282772592b4ea390035e17167da2305724a01164d

Observation 6bbef8c0-7495-4973-9d00-73a4a76939eb · outbound

This paper cites Omnivinci: Enhancing architecture and data for omni-modal understanding llm.arXiv preprint arXiv:2510.15870, 2025.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Omnivinci: Enhancing architecture and data for omni-modal understanding llm.arXiv preprint arXiv:2510.15870, 2025

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.520291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.520291Z digest=sha256:9393a3e2d17e51cd5ea1a9bfb4c775afa871b796ec3f77b83802741efbb805ed

Observation dd4ec0f3-2ddf-4830-b903-7a0a7a708e1a · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.522646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.522646Z digest=sha256:e00cc2aa5c3ac1951e909ac7e97c08e0efe793f35545276ac329c8bde066f114

Observation 89bf6c89-a6a6-480f-ade9-e192819a8638 · outbound

This paper cites R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.524979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.524979Z digest=sha256:67ed1473c4a33017a4c517bf786539668288f9835b3d763a7b9c6ac86aecaeb2

Observation f42ee904-aff0-419d-9760-a9309a7dff18 · outbound

This paper cites An Information Theory-inspired Strategy for Automatic Network Pruning.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models An Information Theory-inspired Strategy for Automatic Network Pruning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.527486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.527486Z digest=sha256:2f0bb1e91efa254de5cbf086ca13cef0e720473f74fd3731bf0d6fad4f013758

Observation fa2f2caa-5eee-46f6-b717-7e92b3b2fa4c · outbound

This paper cites Self-Embed.

OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language Models Self-Embed

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T04:02:54.529833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:02:54.529833Z digest=sha256:ad3f77d56ef39808c3de48261147cc905c38d7cd2f0ba3d9d28bb642aeebe58d

Pith citing papers

No inbound Pith citation observations are available.