Pith. sign in

Paper Citation Record · LEDGER

ToSA: Token Merging with Spatial Awareness

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2506.20066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20066 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:29.963011Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:29:38.951111Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:29:40.174642Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e43f906-9452-4009-90bd-a78c62ddc673 · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

ToSA: Token Merging with Spatial Awareness Dinov2: Learning robust visual features without supervision,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.312138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.840335Z digest=sha256:1cd0b806bfe3793189dc6e90967362c3db18bf9020540df9217fef5c4fe480e5

Observation 31346f8b-c886-440a-8a6c-0d45e01fb506 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

ToSA: Token Merging with Spatial Awareness Learning transferable visual models from natural language supervision,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.843839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.843839Z digest=sha256:93e031cbc86efdb9ed0d31e07ffd9aa5be81b7e2418e31e307d00abc9684c26d

Observation 47d741a3-4730-4439-8a42-c0a638032c5f · outbound

This paper cites Sigmoid loss for language image pre-training,.

ToSA: Token Merging with Spatial Awareness Sigmoid loss for language image pre-training,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.296994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.848181Z digest=sha256:c1ca9739bf9baf70e08f61c952c56739401f04471505659d53285d501bb65c5b

Observation a04e200e-2497-4f6e-b485-3f1c3baefbb5 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

ToSA: Token Merging with Spatial Awareness An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.852111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.852111Z digest=sha256:9d16f1fddc85bfc132e60326139863f9d5b1a2d53a8546893b4a0b742766b881

Observation 72db5e3e-9c2b-4405-babd-6f01a8e834c0 · outbound

This paper cites Visual instruction tuning,.

ToSA: Token Merging with Spatial Awareness Visual instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.855346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.855346Z digest=sha256:57d11c0695b2cb15bd8eca9c7c2c931d4ce91d2179116fd288cd46623593e813

Observation a940bb97-d4fd-4cf9-8501-0c57c3225c6c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ToSA: Token Merging with Spatial Awareness LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.858437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.858437Z digest=sha256:142299c15555a0838790b5ffe10a86a8d750a2cde2e15555446ef06158665bcd

Observation e74f5825-bd81-4f70-a135-506546da5152 · outbound

This paper cites Efficientvit: Memory efficient vision transformer with cascaded group attention,.

ToSA: Token Merging with Spatial Awareness Efficientvit: Memory efficient vision transformer with cascaded group attention,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.282019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.861924Z digest=sha256:71eddc4b0a342af9d5c4fe68e7b19b54cdb4ee9bd5c83caab24a801ec941d069

Observation 9bd3f93e-bb79-45c5-a42a-7ae4eccb9a80 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification,.

ToSA: Token Merging with Spatial Awareness Dynamicvit: Efficient vision transformers with dynamic token sparsification,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.272427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.864786Z digest=sha256:6ca3ade9ac81eb79ffbbde9489b85f33d84775be9e8218ca0c908ad4748c6bb2

Observation 04f5c40e-20c1-44a1-bbde-7afa92299075 · outbound

This paper cites A-vit: Adaptive tokens for efficient vision transformer,.

ToSA: Token Merging with Spatial Awareness A-vit: Adaptive tokens for efficient vision transformer,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.262824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.867828Z digest=sha256:eb752f7504d935e7e0546f9b8b0eb264980feaba55214c85f5e9de7217ad125a

Observation 8f6813e0-f0c8-4423-bb16-19a4b436e0ba · outbound

This paper cites TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action.

ToSA: Token Merging with Spatial Awareness TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.871387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.871387Z digest=sha256:968475057849a0d5e7d6f9a116786fadf618f0d73957617588303e132cb2fc7d

Observation b4463c08-a9f5-423f-84a5-02b9c02edb06 · outbound

This paper cites Token pooling in vision transformers for image classification,.

ToSA: Token Merging with Spatial Awareness Token pooling in vision transformers for image classification,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.253657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.874526Z digest=sha256:2a2c5cdb9dc5aead1e7c79d0bc9335d82cb7fb02b8bf80b8c09a502461dee739

Observation e64efe56-fc1a-4d30-bb07-24fd401a13f5 · outbound

This paper cites Zero-shot 3d question answering via voxel-based dynamic token compression,.

ToSA: Token Merging with Spatial Awareness Zero-shot 3d question answering via voxel-based dynamic token compression,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.244426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.877445Z digest=sha256:ccfd8171d3a337f4835d2b192acb36401475d84b067fac44e2fef4ff02dd09cc

Observation 7d9d162f-cc20-4684-bca7-c8051db30524 · outbound

This paper cites Token merging: Your ViT but faster,.

ToSA: Token Merging with Spatial Awareness Token merging: Your ViT but faster,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.235279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.880629Z digest=sha256:4bdd0e41b8c483265418484caacf5205f478ac6628699ec399875ba0cbf79634

Observation 71ebff2d-17a0-4d8a-b8e2-2956bc460053 · outbound

This paper cites What do Vision Transformers Learn? A Visual Exploration.

ToSA: Token Merging with Spatial Awareness What do Vision Transformers Learn? A Visual Exploration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.883551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.883551Z digest=sha256:4818c4958f6a75b3da55f4c5b38bf10342747f2794c2deb30c695a3cc74d8168

Observation ae73add2-f905-443d-b7aa-abbab6e06216 · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

ToSA: Token Merging with Spatial Awareness SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.887022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.887022Z digest=sha256:273c46bb3d1ae0424613e1b65c71f9fdaa642987c5ef95bcfec96ef21f7dfa8d

Observation adb71e1e-120d-46fa-8bf5-a8ff8a7e430e · outbound

This paper cites Vqa: Visual question answering,.

ToSA: Token Merging with Spatial Awareness Vqa: Visual question answering,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.890588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.890588Z digest=sha256:d6964c791cbf90ea0c704a76fb8e6d407e49dda93a58691f1e6feb12b7229a3f

Observation 5817ad76-faa0-4ab9-a8ca-94e0cb8a63ae · outbound

This paper cites Gqa: A new dataset for real- world visual reasoning and compositional question answering,.

ToSA: Token Merging with Spatial Awareness Gqa: A new dataset for real- world visual reasoning and compositional question answering,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.220159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.894119Z digest=sha256:1cabcac7344e200b5e28e9e8ae3e210cbd642f3a6e1853329065ac3033732a71

Observation 6bab8da7-a90f-4414-be8d-514a9a4a2bf9 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models,.

ToSA: Token Merging with Spatial Awareness Openeqa: Embodied question answering in the era of foundation models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.211596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.896953Z digest=sha256:bd930ac008a19a5c8c64097aad69f5591e54db486c5e8873a92e27d71f971d75

Observation aa4213b9-0494-4f27-ba2b-1861cf231b88 · outbound

This paper cites Sp-vit: Learning 2d spatial priors for vision transformers,.

ToSA: Token Merging with Spatial Awareness Sp-vit: Learning 2d spatial priors for vision transformers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.202142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.899813Z digest=sha256:1dc42b19cbce04ffcdb538a757d674d54b3b0be08928467a33dab83b1c97429e

Observation d9f1d06b-b422-4d58-aba9-6c12e1392931 · outbound

This paper cites Evo-vit: Slow-fast token evolution for dynamic vision transformer,.

ToSA: Token Merging with Spatial Awareness Evo-vit: Slow-fast token evolution for dynamic vision transformer,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.193380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.902805Z digest=sha256:7b1621a7b30607011632eec53a9eaaaab910a41f7e722133c1b961c518e68dc4

Observation 3f03106a-f2a8-43c6-a505-631d9b8d6fa9 · outbound

This paper cites Not all patches are what you need: Expediting vision transformers via token reorganizations,.

ToSA: Token Merging with Spatial Awareness Not all patches are what you need: Expediting vision transformers via token reorganizations,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.183813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.905853Z digest=sha256:3e63bf0697f110e2cd9746b189e77c44adc320eb1d66a4196ad4d90f4a4cecd6

Observation 718ac7db-d307-450e-a7d2-bd35898f18f9 · outbound

This paper cites PPT: Token Pruning and Pooling for Efficient Vision Transformers.

ToSA: Token Merging with Spatial Awareness PPT: Token Pruning and Pooling for Efficient Vision Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.908663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.908663Z digest=sha256:0969ea812635ca48f19f5fe19a97944b169f397c8ca52d5c84a8a0d623857918

Observation d7efb4b4-3a43-4e95-a7fc-162937ae0998 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference ac- celeration for large vision-language models,.

ToSA: Token Merging with Spatial Awareness An image is worth 1/2 tokens after layer 2: Plug-and-play inference ac- celeration for large vision-language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.174741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.911840Z digest=sha256:391063989c479065e2930bb340235505d9855cf104a1d7866fe39776f6faf764

Observation e2a3183e-3954-428f-8d4d-1c3037f0678d · outbound

This paper cites Sparsevlm: Visual token sparsification for efficient vision-language model inference,.

ToSA: Token Merging with Spatial Awareness Sparsevlm: Visual token sparsification for efficient vision-language model inference,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.164160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.914750Z digest=sha256:b96dab06cdd2b9fcf25a6803e9ba277864ab766f3e60a5d2ad8814abea28e6f0

Observation 69f882be-bbcd-4864-9259-482ec2173ddb · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,.

ToSA: Token Merging with Spatial Awareness Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.154332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.917634Z digest=sha256:20f23ace70247c2fc4c09d97f751a51b651a1f4604aa1259b52ae4d0dad19535

Observation f2b6a679-237c-449a-ba4d-f6d133ceca6a · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models,.

ToSA: Token Merging with Spatial Awareness Spatialrgpt: Grounded spatial reasoning in vision-language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.145170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.920675Z digest=sha256:c0989792dcd56f0d0b83313d58cb8ed1ecc2f7c97144c78c03e6f4e978295a3d

Observation af21c5cc-af7a-428d-8816-33de85b0785c · outbound

This paper cites Attention is all you need,.

ToSA: Token Merging with Spatial Awareness Attention is all you need,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.923678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.923678Z digest=sha256:e6ca4441a7d89fed470ff793e329717b9be38b1962517193e526cf9302bfe9e1

Observation ea6ef157-f9ed-4beb-8122-dba2c9bdb969 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

ToSA: Token Merging with Spatial Awareness Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.926930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.926930Z digest=sha256:c8513b3238e68ded2e1580f44c5060b30ce75c8f908aa45c0afb73c0695b3783

Observation bd4975ca-7cd5-41c2-a9d6-060bf6e5d0b1 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning,.

ToSA: Token Merging with Spatial Awareness Instructblip: Towards general-purpose vision- language models with instruction tuning,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.929843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.929843Z digest=sha256:319584bf6e986b8e8e1589910ea3965d01c766c233b710e2b227352e085ec0df

Observation 242f7156-e487-47c6-8819-3b72cb02be43 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ToSA: Token Merging with Spatial Awareness Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.932806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.932806Z digest=sha256:19a5831ecfee74a9c441cb47ee5223e87e4dee18913bdec1c20e0a5c5487c9d3

Observation a3e07ea0-fdbd-4e2b-8c53-1f464f6b9e09 · outbound

This paper cites Improved baselines with visual instruction tuning,.

ToSA: Token Merging with Spatial Awareness Improved baselines with visual instruction tuning,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.936217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.936217Z digest=sha256:15f2d9fc0941ee1a1472f32de7e1eab747733c374425bae0d97555afb4ecbbb4

Observation 015df59e-fd2d-49de-b961-fa434bf38c49 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models,.

ToSA: Token Merging with Spatial Awareness Llama-vid: An image is worth 2 tokens in large language models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.939078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.939078Z digest=sha256:6b4e6b5172659675050083c07762b9c8bea6ca2d45eb2a9e0f3adedea08dfc80

Observation 7563aa29-e8bf-48af-b677-46ce31a1056b · outbound

This paper cites Vila: On pre-training for visual language models,.

ToSA: Token Merging with Spatial Awareness Vila: On pre-training for visual language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.109792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.941985Z digest=sha256:01cb7c3104900c167361134408e81b078b0afa3f425660c8497dc102faa6e673

Observation 8f2e361f-155a-49e1-8770-0abe81f21a22 · outbound

This paper cites Depth anything v2,.

ToSA: Token Merging with Spatial Awareness Depth anything v2,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.100457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.944983Z digest=sha256:3a1d60a189ad72ac3019c88e15ae504a60bba28f3353eed3c47e4c6d032290cf

Observation 3bf9bb16-7765-40f5-b839-9587f6f7bfcb · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

ToSA: Token Merging with Spatial Awareness AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.947783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.947783Z digest=sha256:e96635bc2e9caaad844179d5eb468f566e98df0e7903117bb660935ef185be16

Observation bcb28fe4-92e5-4dba-ab9a-101a3e869823 · outbound

This paper cites Longvlm: Efficient long video understanding via large language models,.

ToSA: Token Merging with Spatial Awareness Longvlm: Efficient long video understanding via large language models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.091338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.951218Z digest=sha256:88feb0d9ac70fb0bc5c3a1dbf56de1ef5425d22510c3d3fe235fe03782daec22

Observation 8f28d1d8-cdee-4d34-8e7b-83e03edac955 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models,.

ToSA: Token Merging with Spatial Awareness Video-chatgpt: Towards detailed video understanding via large vision and language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.082347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.954230Z digest=sha256:d2af79a508d58ac2e266ffec2a2bf8ec3124756c57fcc4124330e62447bc8008

Observation 2b74cc36-57e4-4d1d-99bb-022d8b87d674 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding,.

ToSA: Token Merging with Spatial Awareness Video-llama: An instruction-tuned audio-visual language model for video understanding,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.072638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.957084Z digest=sha256:93dfcebc83e06e03812de4e4230668c46f092e31e7b44deef7d25732cdb3fac1

Observation fab74acb-a524-4e80-a5b2-bc01b79f2c1e · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

ToSA: Token Merging with Spatial Awareness VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.959967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.959967Z digest=sha256:515262e9df51682a9b14ddf92600edb9e3769154cadf35f34c5895d68f9da48b

Observation f94ba385-b0bb-4504-a4ab-9870b150d347 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding,.

ToSA: Token Merging with Spatial Awareness Chat-univi: Unified visual representation empowers large language models with image and video understanding,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.063216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T23:01:29.963011Z digest=sha256:fafc59ece2ec38d0cf39b086aa2f3c718d92ff9b6b934e63734dcb1dacb284ed

Pith citing papers

Observation 91eac91c-5b43-4613-918e-ed4b68643fa5 · inbound

Warehouse Spatial Question Answering with LLM Agent cites this paper.

Warehouse Spatial Question Answering with LLM Agent ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:29:40.213090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:38.951111Z digest=sha256:9c0198de303638ac43abe1844bf51d77d7b60d8451636b2c47caca536d2e747d

Observation 1ff61093-2310-4b77-bb4b-06a867db39d9 · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T23:59:16.247451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:59:16.247451Z digest=sha256:c281db95c865b6b44479649a9504c9492ba9b927d3094f477d06893a7e71219f

Observation 974bb5d5-894c-4430-9d4f-4be4e26c6499 · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.518353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.518353Z digest=sha256:7be1ceddbebe8b4961bf12b47bbd0d4e4910a118746531eb27544b94a1f5d1ff