Pith. sign in

Paper Citation Record · LEDGER

ToSA: Token Merging with Spatial Awareness

As of 19 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2506.20066.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.20066 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:01:29.963011Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:29:38.951111Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T17:29:40.174642Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1e43f906-9452-4009-90bd-a78c62ddc673 · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

ToSA: Token Merging with Spatial Awareness Dinov2: Learning robust visual features without supervision,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.312138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.840335Z digest=sha256:595d3eda6dd3ebba2223a5d043b79a490dac1a78b5bf4d0ffd04fd1ddf12e745

Observation 31346f8b-c886-440a-8a6c-0d45e01fb506 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

ToSA: Token Merging with Spatial Awareness Learning transferable visual models from natural language supervision,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.843839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.843839Z digest=sha256:95be775ff4faa6caa8a21da810417655f43f7407e9fcab445ce38fb99d76e2aa

Observation 47d741a3-4730-4439-8a42-c0a638032c5f · outbound

This paper cites Sigmoid loss for language image pre-training,.

ToSA: Token Merging with Spatial Awareness Sigmoid loss for language image pre-training,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.296994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.848181Z digest=sha256:c15dafba11c39cfc64a1d251e68cc52f2ca3ecb5b6b0701f43da6a2a35bea4c2

Observation a04e200e-2497-4f6e-b485-3f1c3baefbb5 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

ToSA: Token Merging with Spatial Awareness An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.852111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.852111Z digest=sha256:a409fd4ebf679cab4fccc6582546f43f92add5cb6ebbad63778ca66306d2df32

Observation 72db5e3e-9c2b-4405-babd-6f01a8e834c0 · outbound

This paper cites Visual instruction tuning,.

ToSA: Token Merging with Spatial Awareness Visual instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.855346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.855346Z digest=sha256:4eed24f44ce2c505166ef6422a1d4141c5ac03fdd421d1b1bfb05371bd770f6b

Observation a940bb97-d4fd-4cf9-8501-0c57c3225c6c · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

ToSA: Token Merging with Spatial Awareness LLaVA-OneVision: Easy Visual Task Transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.858437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.858437Z digest=sha256:aa5fd2f978308cf5c2bc67198d3aa895a6a3158e4912b85fcc3d123c32efe6d2

Observation e74f5825-bd81-4f70-a135-506546da5152 · outbound

This paper cites Efficientvit: Memory efficient vision transformer with cascaded group attention,.

ToSA: Token Merging with Spatial Awareness Efficientvit: Memory efficient vision transformer with cascaded group attention,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.282019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.861924Z digest=sha256:2ce6a8a17f8f53e5381217292d50b57374c98942fe382287d138131e60f7d06d

Observation 9bd3f93e-bb79-45c5-a42a-7ae4eccb9a80 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification,.

ToSA: Token Merging with Spatial Awareness Dynamicvit: Efficient vision transformers with dynamic token sparsification,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.272427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.864786Z digest=sha256:3e40b6b24c76984dae20192c0a1267b2dd9063d7fe5259b5f6bafe58cc0b2986

Observation 04f5c40e-20c1-44a1-bbde-7afa92299075 · outbound

This paper cites A-vit: Adaptive tokens for efficient vision transformer,.

ToSA: Token Merging with Spatial Awareness A-vit: Adaptive tokens for efficient vision transformer,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.262824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.867828Z digest=sha256:0cc19d9e57b8f1f2380f91bb17561a8f4e4a4cb853921e51f7fe08e49051b501

Observation 8f6813e0-f0c8-4423-bb16-19a4b436e0ba · outbound

This paper cites TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action.

ToSA: Token Merging with Spatial Awareness TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.871387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.871387Z digest=sha256:b76ced08b43d7a5b5c0a279df5384983aa7c34b088ad65fcc438b3fc9324a3cf

Observation b4463c08-a9f5-423f-84a5-02b9c02edb06 · outbound

This paper cites Token pooling in vision transformers for image classification,.

ToSA: Token Merging with Spatial Awareness Token pooling in vision transformers for image classification,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.253657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.874526Z digest=sha256:cb74b0c7a56655200bfee9e257788a903c6ebc41c7a7491cd84cec5e5e9ef4fc

Observation e64efe56-fc1a-4d30-bb07-24fd401a13f5 · outbound

This paper cites Zero-shot 3d question answering via voxel-based dynamic token compression,.

ToSA: Token Merging with Spatial Awareness Zero-shot 3d question answering via voxel-based dynamic token compression,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.244426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.877445Z digest=sha256:8c380c73a94eee3a73339d5edc9c11d057132f969d56bf3dd587acff19b8543d

Observation 7d9d162f-cc20-4684-bca7-c8051db30524 · outbound

This paper cites Token merging: Your ViT but faster,.

ToSA: Token Merging with Spatial Awareness Token merging: Your ViT but faster,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.235279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.880629Z digest=sha256:a5e1a09d4387272d6b180e0cd5c5ae14468b64fdbe72d3f9431f1d2f82aeca67

Observation 71ebff2d-17a0-4d8a-b8e2-2956bc460053 · outbound

This paper cites What do Vision Transformers Learn? A Visual Exploration.

ToSA: Token Merging with Spatial Awareness What do Vision Transformers Learn? A Visual Exploration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.883551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.883551Z digest=sha256:80d8de19a64cb97e84434fc4cca608d1612913345ed9470738b8d7256e781821

Observation ae73add2-f905-443d-b7aa-abbab6e06216 · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

ToSA: Token Merging with Spatial Awareness SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.887022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.887022Z digest=sha256:f9af7d3f6d53c82d28dadd45deedbe9eafd0fd6661a21b12f2f62fbeeda2c2df

Observation adb71e1e-120d-46fa-8bf5-a8ff8a7e430e · outbound

This paper cites Vqa: Visual question answering,.

ToSA: Token Merging with Spatial Awareness Vqa: Visual question answering,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.890588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.890588Z digest=sha256:8e7242a29dc45f502204562ca97dea0e64cc8339917a6798d0f746c63608e50b

Observation 5817ad76-faa0-4ab9-a8ca-94e0cb8a63ae · outbound

This paper cites Gqa: A new dataset for real- world visual reasoning and compositional question answering,.

ToSA: Token Merging with Spatial Awareness Gqa: A new dataset for real- world visual reasoning and compositional question answering,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.220159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.894119Z digest=sha256:ff3ff33a09e1d15ac3038356ffda2490dff285802368af4ff222f8623f541a93

Observation 6bab8da7-a90f-4414-be8d-514a9a4a2bf9 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models,.

ToSA: Token Merging with Spatial Awareness Openeqa: Embodied question answering in the era of foundation models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.211596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.896953Z digest=sha256:1cd43bda90ab02ee3f43510794c06db1583d6d96eee11fef616d2e0ba11d1f91

Observation aa4213b9-0494-4f27-ba2b-1861cf231b88 · outbound

This paper cites Sp-vit: Learning 2d spatial priors for vision transformers,.

ToSA: Token Merging with Spatial Awareness Sp-vit: Learning 2d spatial priors for vision transformers,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.202142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.899813Z digest=sha256:cac1dd20e5a832091262b7ff2dfb521a36517cece84341ba6ca4571d354ed43e

Observation d9f1d06b-b422-4d58-aba9-6c12e1392931 · outbound

This paper cites Evo-vit: Slow-fast token evolution for dynamic vision transformer,.

ToSA: Token Merging with Spatial Awareness Evo-vit: Slow-fast token evolution for dynamic vision transformer,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.193380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.902805Z digest=sha256:43d163e0123ea3bab78b09e9e011eb517e4a6febbd348932a5f5ac5d91c316fc

Observation 3f03106a-f2a8-43c6-a505-631d9b8d6fa9 · outbound

This paper cites Not all patches are what you need: Expediting vision transformers via token reorganizations,.

ToSA: Token Merging with Spatial Awareness Not all patches are what you need: Expediting vision transformers via token reorganizations,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.183813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.905853Z digest=sha256:fb27821875c44aadc511a63c02e7fdf9b973af4ecc39e837a5d12d117a3746b1

Observation 718ac7db-d307-450e-a7d2-bd35898f18f9 · outbound

This paper cites PPT: Token Pruning and Pooling for Efficient Vision Transformers.

ToSA: Token Merging with Spatial Awareness PPT: Token Pruning and Pooling for Efficient Vision Transformers

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.908663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.908663Z digest=sha256:4eef13b7f394a1e0e30ef307e29c379522e39af745744e671f6f00e0bd10d1d1

Observation d7efb4b4-3a43-4e95-a7fc-162937ae0998 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference ac- celeration for large vision-language models,.

ToSA: Token Merging with Spatial Awareness An image is worth 1/2 tokens after layer 2: Plug-and-play inference ac- celeration for large vision-language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.174741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.911840Z digest=sha256:de97e35dedbcc107cd49d15bac62d588498e4ff1b35fdbd67719e477c62d9187

Observation e2a3183e-3954-428f-8d4d-1c3037f0678d · outbound

This paper cites Sparsevlm: Visual token sparsification for efficient vision-language model inference,.

ToSA: Token Merging with Spatial Awareness Sparsevlm: Visual token sparsification for efficient vision-language model inference,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.164160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.914750Z digest=sha256:4b2f46f0c91c1f721a8d459afa0445442bbb5fc26767611dda8c40e77fd9f24d

Observation 69f882be-bbcd-4864-9259-482ec2173ddb · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,.

ToSA: Token Merging with Spatial Awareness Spatialvlm: Endowing vision-language models with spatial reasoning capabilities,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.154332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.917634Z digest=sha256:f157068775db3efa346226667786c890a4adaf047696dd9054740d4664cd866d

Observation f2b6a679-237c-449a-ba4d-f6d133ceca6a · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models,.

ToSA: Token Merging with Spatial Awareness Spatialrgpt: Grounded spatial reasoning in vision-language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.145170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.920675Z digest=sha256:4e3f5beabf8332ddee4bd7a998d964be05c26973c1aae47ab951c1b538539b4b

Observation af21c5cc-af7a-428d-8816-33de85b0785c · outbound

This paper cites Attention is all you need,.

ToSA: Token Merging with Spatial Awareness Attention is all you need,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.923678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.923678Z digest=sha256:d99257d6b0c11292610ff4dc42836a3da554001a60a5c0fc6aa2c247426883e4

Observation ea6ef157-f9ed-4beb-8122-dba2c9bdb969 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

ToSA: Token Merging with Spatial Awareness Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.926930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.926930Z digest=sha256:98b955b38c9afb305041a4a3b941a228b6dd17a5237f9d61929cf1ec587ee617

Observation bd4975ca-7cd5-41c2-a9d6-060bf6e5d0b1 · outbound

This paper cites Instructblip: Towards general-purpose vision- language models with instruction tuning,.

ToSA: Token Merging with Spatial Awareness Instructblip: Towards general-purpose vision- language models with instruction tuning,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.929843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.929843Z digest=sha256:81bcf534dca6238908700c9739776127412f05f557c31760b4dab84e19574f73

Observation 242f7156-e487-47c6-8819-3b72cb02be43 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ToSA: Token Merging with Spatial Awareness Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.932806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.932806Z digest=sha256:c172eda8e5175cfc3e8390b6644018ff6879eb9ac1d52b3bfaf7177de5a65f41

Observation a3e07ea0-fdbd-4e2b-8c53-1f464f6b9e09 · outbound

This paper cites Improved baselines with visual instruction tuning,.

ToSA: Token Merging with Spatial Awareness Improved baselines with visual instruction tuning,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.936217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.936217Z digest=sha256:249d78ef2998f93752724efef28c794b4fcd7a620bd2b7e74bf621896b89201d

Observation 015df59e-fd2d-49de-b961-fa434bf38c49 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models,.

ToSA: Token Merging with Spatial Awareness Llama-vid: An image is worth 2 tokens in large language models,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.939078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.939078Z digest=sha256:90c9f4192f1e2c9e6b24b3a04dc7c6c3e8179849504c24da209f02ad147f636f

Observation 7563aa29-e8bf-48af-b677-46ce31a1056b · outbound

This paper cites Vila: On pre-training for visual language models,.

ToSA: Token Merging with Spatial Awareness Vila: On pre-training for visual language models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.109792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.941985Z digest=sha256:20b154193675b59706675f374bc477c13f9db1c1ef07cbff36728ac5fae14b68

Observation 8f2e361f-155a-49e1-8770-0abe81f21a22 · outbound

This paper cites Depth anything v2,.

ToSA: Token Merging with Spatial Awareness Depth anything v2,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.100457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.944983Z digest=sha256:3fd1a952dcd9e0b012f30d20034cd257d62b7fb5b1cdf3400e41f9963f438ff5

Observation 3bf9bb16-7765-40f5-b839-9587f6f7bfcb · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

ToSA: Token Merging with Spatial Awareness AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.947783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.947783Z digest=sha256:2197c5fad16d66afe937ed431220779b772495a2f5d5321b0280a92539b486ef

Observation bcb28fe4-92e5-4dba-ab9a-101a3e869823 · outbound

This paper cites Longvlm: Efficient long video understanding via large language models,.

ToSA: Token Merging with Spatial Awareness Longvlm: Efficient long video understanding via large language models,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.091338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.951218Z digest=sha256:82e11337aac6e0ec8c558e979dcb23f53233144ecfba66bac61faa77a87d6a75

Observation 8f28d1d8-cdee-4d34-8e7b-83e03edac955 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models,.

ToSA: Token Merging with Spatial Awareness Video-chatgpt: Towards detailed video understanding via large vision and language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.082347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.954230Z digest=sha256:50e14ff392a8ca4a104cdd784f9fee36dccdedaee9e69b961497f9d80445e4e2

Observation 2b74cc36-57e4-4d1d-99bb-022d8b87d674 · outbound

This paper cites Video-llama: An instruction-tuned audio-visual language model for video understanding,.

ToSA: Token Merging with Spatial Awareness Video-llama: An instruction-tuned audio-visual language model for video understanding,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.072638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.957084Z digest=sha256:4695de0a0f21978ca6d15e87af1a03290d733b88b1cb35f1245d565f33dcc6e1

Observation fab74acb-a524-4e80-a5b2-bc01b79f2c1e · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

ToSA: Token Merging with Spatial Awareness VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:01:29.959967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:01:29.959967Z digest=sha256:852ac274ef359571b6635df4f621a8db03bb2fc7feb86e686f23811072a2a392

Observation f94ba385-b0bb-4504-a4ab-9870b150d347 · outbound

This paper cites Chat-univi: Unified visual representation empowers large language models with image and video understanding,.

ToSA: Token Merging with Spatial Awareness Chat-univi: Unified visual representation empowers large language models with image and video understanding,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:01:30.063216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T23:01:29.963011Z digest=sha256:4717a56ffdf5ece25c67c6ecf6a3cd5ec0bc303caeea2f9c4154d01071787d5f

Pith citing papers

Observation 91eac91c-5b43-4613-918e-ed4b68643fa5 · inbound

Warehouse Spatial Question Answering with LLM Agent cites this paper.

Warehouse Spatial Question Answering with LLM Agent ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:29:40.213090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T17:29:38.951111Z digest=sha256:ef01a991fe8b0e3d5d46b3442005aab08903b4a68e7ede6802714bffe7a6c132

Observation 1ff61093-2310-4b77-bb4b-06a867db39d9 · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T23:59:16.247451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:59:16.247451Z digest=sha256:0905a963b4c29ae759e2bb7fa920338ad9f88ad95777137f9afed7c803d30621

Observation 974bb5d5-894c-4430-9d4f-4be4e26c6499 · inbound

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation cites this paper.

WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T04:01:48.518353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:01:48.518353Z digest=sha256:e9f5f35990f29f26290300ad1e454b1b83695d8768fc5f738caf5fecf16e4d71