Pith. sign in

Paper Citation Record · LEDGER

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding

As of 16 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2507.07744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07744 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:40:05.200664Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c171e9f0-1219-48da-b4e1-b594575d6f54 · outbound

This paper cites Joint visual and audio learning for video highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Joint visual and audio learning for video highlight detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:12.354245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:00.652223Z digest=sha256:cdc6a9fa916ba3b45d1072e82482dd50cd00cb83b8f3186c7a55e20c86c6a035

Observation 9dcd0f8f-f622-4c8e-b43a-e26cea825954 · outbound

This paper cites FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:00.715128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:00.715128Z digest=sha256:60cb8ecdd41ab5d4baa60a15d55c1faa58fde60fb052523496de9049ba9a1e5a

Observation 8c7eba0d-06a4-4bbc-842e-ddab6282ee95 · outbound

This paper cites End-to- end object detection with transformers.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding End-to- end object detection with transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:12.193724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:00.779071Z digest=sha256:dcb03a70ba254276084a9b7071cb689ba304ccdf7a14f20f5c0dc35224e8c59a

Observation 654734bc-97cd-4a42-8fa6-90ae36976868 · outbound

This paper cites Slowfast networks for video recognition.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Slowfast networks for video recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:00.848713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:00.848713Z digest=sha256:aa2bbc3c1ce6b650951dc644e3913fe98085eb29e5abc6059f2b93ca4af36f46

Observation 75d26b24-e8b2-4486-8b32-79c9a7cc3419 · outbound

This paper cites The use of ranks to avoid the assumption of normality implicit in the analysis of variance.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding The use of ranks to avoid the assumption of normality implicit in the analysis of variance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:12.007086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:00.913841Z digest=sha256:35a3ff188af50b8bb1b8b1a9464179aecf9aecf8e66ebdabaefb07b0a5b225d6

Observation cdf82518-2ae2-4220-849a-0bf54a004368 · outbound

This paper cites Tall: Temporal activity localization via language query.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Tall: Temporal activity localization via language query

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.839389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:00.995731Z digest=sha256:47972ce5c6df9fc31c926e1682af342ca0e058c13e19bb5ce3f57fb7060adc1b

Observation 9e3b8fdc-7227-40c6-b3f7-6591f537a59f · outbound

This paper cites Clip-adapter: Better vision-language models with fea- ture adapters.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Clip-adapter: Better vision-language models with fea- ture adapters

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.672575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:01.076972Z digest=sha256:e6a758bb0040d460745030808631b2abd6cf1c785d11ef171ef07e852bd73cd1

Observation 7058ff4b-37a9-4dfe-a399-c044ba2fbe00 · outbound

This paper cites Saliency-guided detr for mo- ment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Saliency-guided detr for mo- ment retrieval and highlight detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:01.160740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:01.160740Z digest=sha256:fc3a7dd4ff414d3a2e03a06dfe4239371397af07b82f19e3471e7a7e33b1f293

Observation a7a763af-351e-423e-b119-84f00fc9f336 · outbound

This paper cites LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:40:06.018099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:01.243689Z digest=sha256:398999871021121e05753e4e9193f572978200c3589caeca2cc1fd51519dd2be

Observation be569355-1b5e-4e5e-9ee8-d15ddfbaf142 · outbound

This paper cites Unleash the Potential of CLIP for Video Highlight Detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Unleash the Potential of CLIP for Video Highlight Detection

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:40:05.757945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:01.315277Z digest=sha256:52268f155e3cffa9ee7581713bd69335fa8795ee16158bbb662dff8e1bd19e5b

Observation c87d8e20-ff76-4ffa-8e6b-08460fe55dd3 · outbound

This paper cites Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:01.379564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:01.379564Z digest=sha256:9ae7517abc839f594fc33845e1c51998c862d5b450d0eae0aa42d92805661179

Observation a5466ecd-48ef-446e-a5ba-6f557beac4a3 · outbound

This paper cites Mini-net: Multiple instance ranking network for video highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Mini-net: Multiple instance ranking network for video highlight detection

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.490094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:01.461875Z digest=sha256:338f240f16e4bc85836b8863adb5d9ca2d0ad2bf4baec4b5890112aed7119702

Observation c99d1de7-495b-47b5-9e05-6b70969c5966 · outbound

This paper cites Fifth berkeley symposium on mathematical statistics and probability.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Fifth berkeley symposium on mathematical statistics and probability

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.152005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:01.530384Z digest=sha256:01f4d92a53e27598a749fd8ce292e2125124fce0c316184d453acad38e1145f4

Observation 4bb2cb10-9f50-4a4e-b88b-35c86fc67807 · outbound

This paper cites Knowing where to focus: Event-aware transformer for video grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Knowing where to focus: Event-aware transformer for video grounding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.779308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:01.644826Z digest=sha256:c14eb8a97b542b388ba764df895b97e84b9505f3786cb99d770a940dada97d87

Observation d6455af2-90d2-4edc-8066-b87cf2bbe5d1 · outbound

This paper cites Vi- sual prompt tuning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Vi- sual prompt tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.378488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:01.720406Z digest=sha256:68e15849fcb77964ceb7428083a7b35e7a2d2500cb7aac781ab6971dfe1c59ef

Observation 027ddf54-2842-4fb2-af89-78088e6f6da9 · outbound

This paper cites FractalNet: Ultra-Deep Neural Networks without Residuals.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding FractalNet: Ultra-Deep Neural Networks without Residuals

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:01.791278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:01.791278Z digest=sha256:1575b279af61fc86eb0c1ce85b1c6240f1503065bb5e8e0aeb2945457eabda13

Observation 1ade79a0-108e-4edd-aaba-28d11e4eb1c9 · outbound

This paper cites Bam-detr: Boundary- aligned moment detection transformer for temporal sentence grounding in videos.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Bam-detr: Boundary- aligned moment detection transformer for temporal sentence grounding in videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.187552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:01.879168Z digest=sha256:b723c003f95e3bf16e31b82ca11d0260c77203feca72b983e6ce1fa878e5217a

Observation 26a573f4-6ffc-41d4-82b9-e287e889ca43 · outbound

This paper cites Detecting moments and highlights in videos via natural language queries.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Detecting moments and highlights in videos via natural language queries

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.085279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:02.017528Z digest=sha256:02ec4b88102c483283d45fafcd52e32270e4c28c85027d8fb22851d94cdec7d5

Observation ac65712d-020c-41bf-9d4c-13f2eb98229f · outbound

This paper cites Dn-detr: Accelerate detr training by intro- ducing query denoising.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Dn-detr: Accelerate detr training by intro- ducing query denoising

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.967704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:02.079804Z digest=sha256:bf10ed97d1b233d3ee6748010b87785bb6ec4145a65c8c976759954899d18af4

Observation b3a7ea83-9220-4053-984c-0aa808346661 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Univtg: Towards unified video- language temporal grounding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.857558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:02.199351Z digest=sha256:3887d1e90bdbbe2a3b96ee5dc85f7541432a03bc6cfdd8af118fda655dcd3440

Observation beb6e54a-05e5-4ed0-927f-41a94c3afc0e · outbound

This paper cites Focal loss for dense object detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Focal loss for dense object detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.316676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.316676Z digest=sha256:9fda5cc468f736317b6fd56b7c8ba3704a6d8000c61506b6adec50e2dd789514

Observation 9d306071-2c72-4600-a409-34c126147cb0 · outbound

This paper cites Frozen clip models are efficient video learners.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Frozen clip models are efficient video learners

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.737363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:02.409539Z digest=sha256:4daaebef823cc0012b67c50c9f61bdaace8edd3a7cbff7c532438b1ac02a1d99

Observation f82a852d-52f8-4d0b-b24c-e873451843be · outbound

This paper cites DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.500974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.500974Z digest=sha256:4c2d490d2cadc029c7972ffc62f2921bb43571f64c2c5eeab37765c5d992ae63

Observation a4d84933-1429-4601-952d-6e011314d4f8 · outbound

This paper cites End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.576985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:02.567378Z digest=sha256:f883d7c20033713d25fccc6aee49cbe50093bf5b7b65cec65745ba665ba7f104

Observation 962d6db7-09bb-4a60-9218-3d4e5fe0ff3f · outbound

This paper cites $R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding $R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.643521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.643521Z digest=sha256:b17cc7c6e5d2cc1b1a864abf867cb6fcc0f68e37c1066c7b1a3cccbba9b91548

Observation 4bb7e029-ef6d-4f59-977d-318d65752d03 · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.450671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:02.714083Z digest=sha256:8ec4d9f260ef3b8787209441132149c1cb395dc1bc5e56e2863db1bb5b44bc88

Observation 34b3b1a3-7981-425a-a102-7663f633fb81 · outbound

This paper cites LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.771931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.771931Z digest=sha256:d0a5ff448a3cda8866251e980624c178c016d94c200deffe4ad7481281191f57

Observation d8b4e273-c1b5-45ef-b906-8cbe788b8875 · outbound

This paper cites The surprising effectiveness of multimodal large language models for video moment retrieval.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding The surprising effectiveness of multimodal large language models for video moment retrieval

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.837075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.837075Z digest=sha256:c941a4176b20e9fda6a86d764f0c6188a23d91f08682303254bb1501673ef4aa

Observation f40d9eda-caae-4d09-a7d6-9c140d7b66e5 · outbound

This paper cites Correlation-guided query-dependency calibration in video representation learning for temporal grounding.CoRR, 2023.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Correlation-guided query-dependency calibration in video representation learning for temporal grounding.CoRR, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.354168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:02.921358Z digest=sha256:a3c40abf51aceb74fa7177f0d474337a011dc37b28622ee6a08cbb50a2b2bab1

Observation 96ce5b25-054a-4166-a25e-886ed9a96a9b · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.205916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:03.042994Z digest=sha256:0adb6bad1ed80391d58b17f5edefeff336d0ffaea772c835f439bd74e878f79a

Observation dd6a6245-faf9-4110-ad95-d857ae3466e3 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Representation Learning with Contrastive Predictive Coding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:03.121427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:03.121427Z digest=sha256:de712305ca7ac7832c1b1487f1028fecf2ed4c953be96fcc89c0efb1e92f85a5

Observation da0e6ad1-6cb4-4050-8c67-39d70e34a7a4 · outbound

This paper cites St-adapter: Parameter-efficient image-to-video transfer learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding St-adapter: Parameter-efficient image-to-video transfer learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.051372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:03.202428Z digest=sha256:ce0d61f7c5ccfaa5764d693c78bc5faa97d1e7192a0059786b483d8fd7bd768d

Observation 0019f4b3-10f4-4eb8-81ba-ed5f4fb2d95c · outbound

This paper cites Sada: Semantic adversarial unsupervised domain adaptation for temporal action localization.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Sada: Semantic adversarial unsupervised domain adaptation for temporal action localization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.867577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:03.289501Z digest=sha256:62caac7e677f2f88f660a3d791573add3efd487228e5dd42df9a8e0b58824bc0

Observation d16eecaf-28d0-4610-b94b-afdcbbaf92c1 · outbound

This paper cites Disentangling spatial and temporal learning for efficient image-to-video transfer learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Disentangling spatial and temporal learning for efficient image-to-video transfer learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.643095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:03.370771Z digest=sha256:4824ca8bf764c4b399b2594ebba13560c303a6231cf0bce6cc1fe4bc6654a034

Observation ed20371f-0403-4410-b007-1bdccf5441a7 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Learning transferable visual models from natural language supervi- sion

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:03.510044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:03.510044Z digest=sha256:873b95696c49ec8db690d8c18f86165c65829620d28c597672007ca206a958e8

Observation cc3e3217-270b-42e1-b80f-c8f76e0fb56f · outbound

This paper cites Coherent multi-sentence video description with variable level of detail.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Coherent multi-sentence video description with variable level of detail

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.444433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:03.632073Z digest=sha256:5108f7343d89eda27f811ff85ab7d93ce6b13f743ae5dcef60540be028227653

Observation 22c1ed98-f3ce-4cd9-ae74-f64b4bd54f25 · outbound

This paper cites Ranking domain- specific highlights by analyzing edited videos.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Ranking domain- specific highlights by analyzing edited videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.271012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:03.734332Z digest=sha256:f6cac6ad86626dc9c4916f29adf0633a3f1821df6f0c8af12b0ef39d3a17871d

Observation 55b6973d-c485-45ac-9314-207f4d83df5b · outbound

This paper cites Lst: Lad- der side-tuning for parameter and memory efficient transfer learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Lst: Lad- der side-tuning for parameter and memory efficient transfer learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.112200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:03.833480Z digest=sha256:ac4654dfc2eca7a10d6f7b0180dc9e3c9fc995556d0a531cc09d9e70ab8b2a9b

Observation 0afbd36f-a5ef-489d-9838-1bf659250beb · outbound

This paper cites Attention is all you need.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Attention is all you need

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:03.914238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:03.914238Z digest=sha256:84bf63aaae192a585fc00608e9b6d9e4a81879e973d8fbe58a18773e554e8f47

Observation dc814c50-2861-428e-a020-1ced0515292a · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.944574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:04.067540Z digest=sha256:58866fc0be76103d3f947448514214eb8f32cc08e49b20dd0fd6d10168d9f745

Observation dae05960-f941-4840-985e-c37e44273954 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:04.192757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:04.192757Z digest=sha256:a2e6bf38b09e247054eb18a1a90f2b5fd1325b530987b7f22b5efeeef70f60c6

Observation fcaac928-c242-44ff-ba16-b035f1054970 · outbound

This paper cites Vision transformer with deformable attention.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Vision transformer with deformable attention

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.710516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:04.293577Z digest=sha256:b946b7d450a98258e3f33382cc4b7cfded2d60d934a6862d6509d8882e356be6

Observation 6e3c69d9-d6f6-42f5-b21b-21bce7ccd6d1 · outbound

This paper cites Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.511209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:04.371035Z digest=sha256:854a56aa1f1a04015ff0bcbd09c8a4ec84e0489a259b14e58e1a0fb1f6d6fbdd

Observation 542b7c04-9546-4022-bcc1-47839712a2bf · outbound

This paper cites Cross-category video high- light detection via set-based learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Cross-category video high- light detection via set-based learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.318203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:04.445706Z digest=sha256:07da7cf6d0d6c59a765d9739a30b8aa8e9a96400c67295580ed82ca892e3cd65

Observation f6177304-f351-4ff1-9927-97870806d7c9 · outbound

This paper cites Unloc: A unified framework for video localiza- tion tasks.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Unloc: A unified framework for video localiza- tion tasks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.110265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:04.563858Z digest=sha256:58aad3a0e79d0a3a77bd3631462f7b1683c69cd996f6e247258b00623b4de034

Observation 15576df0-3d5b-4402-aa8c-5705814ab4a4 · outbound

This paper cites Understanding negative sampling in graph representation learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Understanding negative sampling in graph representation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:06.893073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:04.698446Z digest=sha256:5666574284edfde4f5f42ab4442b0b2b0ab6bf811c6186e9983e170eda877d01

Observation a89e8865-aa01-4c01-bda5-86a15edc9dd0 · outbound

This paper cites Parameter-efficient is not sufficient: Exploring parame- ter, memory, and time efficient adapter tuning for dense pre- dictions.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Parameter-efficient is not sufficient: Exploring parame- ter, memory, and time efficient adapter tuning for dense pre- dictions

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:06.739159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:04.818146Z digest=sha256:7cd31080ffc413ea91cb761a28f60e7034cbe53d4a60c381d6d63fa0976c63ef

Observation d05ac6e5-583e-49f8-a42f-a901e059b357 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:04.919669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:04.919669Z digest=sha256:1a8452f0e2e72f41aa4de8ea2e4545d174f23662cce1d16458a2be2e4f4613fb

Observation f196f109-36e6-44b2-a4b1-68282aa34209 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Conditional prompt learning for vision-language mod- els

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:06.521597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:05.020736Z digest=sha256:7156734f0019b1b8a9a06d76e10a863953b8dbaa46a2adbebf116e9bf502c1c8

Observation 93b92d5b-bd99-4e09-8c12-71c2c64db6fa · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:05.125881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:05.125881Z digest=sha256:7315160dc8178f60306db76ee3aaa04829914789a910221e961bd87a853e7707

Observation 60b4fea0-4029-4292-816c-9e4963f77bea · outbound

This paper cites Notably, these losses are applied to all the intermediate layers inde- pendently.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Notably, these losses are applied to all the intermediate layers inde- pendently

Reference 51

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:40:06.366128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T18:40:05.200664Z digest=sha256:f38df92aef05d34bcceebf583687314e5c5a685fae3d2327435b2ccdaf8d8c12

Pith citing papers

No inbound Pith citation observations are available.