Pith. sign in

Paper Citation Record · LEDGER

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding

As of 15 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2507.07744.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07744 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:40:05.200664Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c171e9f0-1219-48da-b4e1-b594575d6f54 · outbound

This paper cites Joint visual and audio learning for video highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Joint visual and audio learning for video highlight detection

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:12.354245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:00.652223Z digest=sha256:ecee242da09edd4262b16b6a19de94b59600f58ef776f3e6a0c5bc9f42834a2f

Observation 9dcd0f8f-f622-4c8e-b43a-e26cea825954 · outbound

This paper cites FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding FlashVTG: Feature Layering and Adaptive Score Handling Network for Video Temporal Grounding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:00.715128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:00.715128Z digest=sha256:60cb8ecdd41ab5d4baa60a15d55c1faa58fde60fb052523496de9049ba9a1e5a

Observation 8c7eba0d-06a4-4bbc-842e-ddab6282ee95 · outbound

This paper cites End-to- end object detection with transformers.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding End-to- end object detection with transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:12.193724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:00.779071Z digest=sha256:03c2e1adb0aaf412f215e13cf5b3b71a1a2db91a38d020c776c5b4d15beea582

Observation 654734bc-97cd-4a42-8fa6-90ae36976868 · outbound

This paper cites Slowfast networks for video recognition.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Slowfast networks for video recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:00.848713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:00.848713Z digest=sha256:aa2bbc3c1ce6b650951dc644e3913fe98085eb29e5abc6059f2b93ca4af36f46

Observation 75d26b24-e8b2-4486-8b32-79c9a7cc3419 · outbound

This paper cites The use of ranks to avoid the assumption of normality implicit in the analysis of variance.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding The use of ranks to avoid the assumption of normality implicit in the analysis of variance

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:12.007086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:00.913841Z digest=sha256:153daae129fa8d1514d984d8d51336613592bc573b129c0d976202399b8304bb

Observation cdf82518-2ae2-4220-849a-0bf54a004368 · outbound

This paper cites Tall: Temporal activity localization via language query.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Tall: Temporal activity localization via language query

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.839389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:00.995731Z digest=sha256:0da0822de6766b34b8f24fcf443174952473c5e51ee3cbff1dbf295e94f385ff

Observation 9e3b8fdc-7227-40c6-b3f7-6591f537a59f · outbound

This paper cites Clip-adapter: Better vision-language models with fea- ture adapters.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Clip-adapter: Better vision-language models with fea- ture adapters

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.672575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:01.076972Z digest=sha256:9b75b6423a7c5d4193bd96356a7bb21a4b2eb61bb603cff665358b5fd3b6017d

Observation 7058ff4b-37a9-4dfe-a399-c044ba2fbe00 · outbound

This paper cites Saliency-guided detr for mo- ment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Saliency-guided detr for mo- ment retrieval and highlight detection

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:01.160740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:01.160740Z digest=sha256:fc3a7dd4ff414d3a2e03a06dfe4239371397af07b82f19e3471e7a7e33b1f293

Observation a7a763af-351e-423e-b119-84f00fc9f336 · outbound

This paper cites LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding LoSA: Long-Short-range Adapter for Scaling End-to-End Temporal Action Localization

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:40:06.018099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:01.243689Z digest=sha256:c9905b6abd08276aee41dbbb3d8cb8c1fe0ce2f760b61b1218f6c1b73ce45caf

Observation be569355-1b5e-4e5e-9ee8-d15ddfbaf142 · outbound

This paper cites Unleash the Potential of CLIP for Video Highlight Detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Unleash the Potential of CLIP for Video Highlight Detection

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:40:05.757945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:01.315277Z digest=sha256:fbeb28d564fcd12ac959327f89fbd0c6dfcf5436cb394b656fbf458799d6b838

Observation c87d8e20-ff76-4ffa-8e6b-08460fe55dd3 · outbound

This paper cites Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:01.379564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:01.379564Z digest=sha256:9ae7517abc839f594fc33845e1c51998c862d5b450d0eae0aa42d92805661179

Observation a5466ecd-48ef-446e-a5ba-6f557beac4a3 · outbound

This paper cites Mini-net: Multiple instance ranking network for video highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Mini-net: Multiple instance ranking network for video highlight detection

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.490094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:01.461875Z digest=sha256:2e0df25f34af33316403b0b5564c11ba26503d8f65aedfc49bc2d6832aa8f229

Observation c99d1de7-495b-47b5-9e05-6b70969c5966 · outbound

This paper cites Fifth berkeley symposium on mathematical statistics and probability.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Fifth berkeley symposium on mathematical statistics and probability

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:11.152005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:01.530384Z digest=sha256:33c21bda2f1238628997e422ead6957b8f0c8e51a0480b682b1417bc67af40de

Observation 4bb2cb10-9f50-4a4e-b88b-35c86fc67807 · outbound

This paper cites Knowing where to focus: Event-aware transformer for video grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Knowing where to focus: Event-aware transformer for video grounding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.779308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:01.644826Z digest=sha256:3f217421069b6d4fde519708c8f857761e9605eda39df1a920e07f7beee29da3

Observation d6455af2-90d2-4edc-8066-b87cf2bbe5d1 · outbound

This paper cites Vi- sual prompt tuning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Vi- sual prompt tuning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.378488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:01.720406Z digest=sha256:4286718bf575e6f6bf8795a52e6882a1ad49211a6df0e8eaf6c82dd1ae96ffea

Observation 027ddf54-2842-4fb2-af89-78088e6f6da9 · outbound

This paper cites FractalNet: Ultra-Deep Neural Networks without Residuals.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding FractalNet: Ultra-Deep Neural Networks without Residuals

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:01.791278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:01.791278Z digest=sha256:1575b279af61fc86eb0c1ce85b1c6240f1503065bb5e8e0aeb2945457eabda13

Observation 1ade79a0-108e-4edd-aaba-28d11e4eb1c9 · outbound

This paper cites Bam-detr: Boundary- aligned moment detection transformer for temporal sentence grounding in videos.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Bam-detr: Boundary- aligned moment detection transformer for temporal sentence grounding in videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.187552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:01.879168Z digest=sha256:eeffa8727924f94c7bc2ca9dde016f57b2c69318aa1dbd787f44a94b45b13065

Observation 26a573f4-6ffc-41d4-82b9-e287e889ca43 · outbound

This paper cites Detecting moments and highlights in videos via natural language queries.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Detecting moments and highlights in videos via natural language queries

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:10.085279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:02.017528Z digest=sha256:4bb56e84593bf5f81030ceaded415432e59a4557935e8e9d2903cdb4e909f979

Observation ac65712d-020c-41bf-9d4c-13f2eb98229f · outbound

This paper cites Dn-detr: Accelerate detr training by intro- ducing query denoising.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Dn-detr: Accelerate detr training by intro- ducing query denoising

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.967704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:02.079804Z digest=sha256:6087fbf9fc359f8561b77cb7188148d331ce25b9c77809db0d8926be07dbd369

Observation b3a7ea83-9220-4053-984c-0aa808346661 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Univtg: Towards unified video- language temporal grounding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.857558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:02.199351Z digest=sha256:21bbbc8f93e34848df2c8b5d6dd4e4a7fcc386dbf932e755861bc32cffe00916

Observation beb6e54a-05e5-4ed0-927f-41a94c3afc0e · outbound

This paper cites Focal loss for dense object detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Focal loss for dense object detection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.316676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.316676Z digest=sha256:9fda5cc468f736317b6fd56b7c8ba3704a6d8000c61506b6adec50e2dd789514

Observation 9d306071-2c72-4600-a409-34c126147cb0 · outbound

This paper cites Frozen clip models are efficient video learners.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Frozen clip models are efficient video learners

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.737363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:02.409539Z digest=sha256:def5de3316b98215fa4896ee71ea140d4ecb211dc092d73e9678558c1e074d08

Observation f82a852d-52f8-4d0b-b24c-e873451843be · outbound

This paper cites DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETR

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.500974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.500974Z digest=sha256:03e5cda3e9a1b4db6c12866510e158103332054971f0fc6974c2828e77daa461

Observation a4d84933-1429-4601-952d-6e011314d4f8 · outbound

This paper cites End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding End-to-end temporal ac- tion detection with transformer.IEEE Transactions on Image Processing, 31:5427–5441, 2022

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.576985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:02.567378Z digest=sha256:0e7b005805e22540868d4240c4ff23af79548c6298cd8cdce12c2cf60c64a8af

Observation 962d6db7-09bb-4a60-9218-3d4e5fe0ff3f · outbound

This paper cites $R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding $R^2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.643521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.643521Z digest=sha256:039922655ab62419cffdb5b1c1beef11f72fcbd8ce5ae4b20e16d7a9a5898e32

Observation 4bb7e029-ef6d-4f59-977d-318d65752d03 · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.450671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:02.714083Z digest=sha256:2b76e5a8ccd9bc02e7232f6e882a857b131059f853266dbbab6e6457da63a2cb

Observation 34b3b1a3-7981-425a-a102-7663f633fb81 · outbound

This paper cites LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding LLaVA-MR: Large Language-and-Vision Assistant for Video Moment Retrieval

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.771931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.771931Z digest=sha256:121a18032b847b888222e5baa5818be4375cd77141ef7e17776c638924699cf0

Observation d8b4e273-c1b5-45ef-b906-8cbe788b8875 · outbound

This paper cites The surprising effectiveness of multimodal large language models for video moment retrieval.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding The surprising effectiveness of multimodal large language models for video moment retrieval

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:02.837075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:02.837075Z digest=sha256:c941a4176b20e9fda6a86d764f0c6188a23d91f08682303254bb1501673ef4aa

Observation f40d9eda-caae-4d09-a7d6-9c140d7b66e5 · outbound

This paper cites Correlation-guided query-dependency calibration in video representation learning for temporal grounding.CoRR, 2023.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Correlation-guided query-dependency calibration in video representation learning for temporal grounding.CoRR, 2023

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.354168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:02.921358Z digest=sha256:53405ac1789a42d8012be126099a843c9c57c9ee4a703904a16f53def87c371e

Observation 96ce5b25-054a-4166-a25e-886ed9a96a9b · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.205916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:03.042994Z digest=sha256:598c6265aeadc7087fb06110d23b68b98a89610101e4a8682bc44581bb8493a0

Observation dd6a6245-faf9-4110-ad95-d857ae3466e3 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Representation Learning with Contrastive Predictive Coding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:03.121427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:03.121427Z digest=sha256:de712305ca7ac7832c1b1487f1028fecf2ed4c953be96fcc89c0efb1e92f85a5

Observation da0e6ad1-6cb4-4050-8c67-39d70e34a7a4 · outbound

This paper cites St-adapter: Parameter-efficient image-to-video transfer learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding St-adapter: Parameter-efficient image-to-video transfer learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:09.051372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:03.202428Z digest=sha256:b5fba72b5e0c7337ada97ec6b252a83d20cfd37609f6aaaf0795e113d031ac9f

Observation 0019f4b3-10f4-4eb8-81ba-ed5f4fb2d95c · outbound

This paper cites Sada: Semantic adversarial unsupervised domain adaptation for temporal action localization.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Sada: Semantic adversarial unsupervised domain adaptation for temporal action localization

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.867577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:03.289501Z digest=sha256:fbef8affa97563b045e7c4768212c5177ecaaf745e03f495914e83bd32f5cbae

Observation d16eecaf-28d0-4610-b94b-afdcbbaf92c1 · outbound

This paper cites Disentangling spatial and temporal learning for efficient image-to-video transfer learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Disentangling spatial and temporal learning for efficient image-to-video transfer learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.643095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:03.370771Z digest=sha256:12dbd7b381882f9cb1d937fc69535f2f902a94cda51c3021277a4711a5386495

Observation ed20371f-0403-4410-b007-1bdccf5441a7 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Learning transferable visual models from natural language supervi- sion

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:03.510044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:03.510044Z digest=sha256:873b95696c49ec8db690d8c18f86165c65829620d28c597672007ca206a958e8

Observation cc3e3217-270b-42e1-b80f-c8f76e0fb56f · outbound

This paper cites Coherent multi-sentence video description with variable level of detail.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Coherent multi-sentence video description with variable level of detail

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.444433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:03.632073Z digest=sha256:b24db49526861d6b260312a0c9807c77e46716eff4f6159619e85417db234916

Observation 22c1ed98-f3ce-4cd9-ae74-f64b4bd54f25 · outbound

This paper cites Ranking domain- specific highlights by analyzing edited videos.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Ranking domain- specific highlights by analyzing edited videos

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.271012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:03.734332Z digest=sha256:250a96e4c7b6720c516ab45f8713013a6718339419fd06d626ed5e939717bbdc

Observation 55b6973d-c485-45ac-9314-207f4d83df5b · outbound

This paper cites Lst: Lad- der side-tuning for parameter and memory efficient transfer learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Lst: Lad- der side-tuning for parameter and memory efficient transfer learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:08.112200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:03.833480Z digest=sha256:581d2e29aa0fb8032fd8b7a777b09d364adb8afdc4dd496804def55521a5ddfa

Observation 0afbd36f-a5ef-489d-9838-1bf659250beb · outbound

This paper cites Attention is all you need.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Attention is all you need

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:03.914238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:03.914238Z digest=sha256:84bf63aaae192a585fc00608e9b6d9e4a81879e973d8fbe58a18773e554e8f47

Observation dc814c50-2861-428e-a020-1ced0515292a · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.944574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:04.067540Z digest=sha256:6c9cc9a8945dc742cbf63c31d6c65fd31c8a9a924c18bd07e8189e2acf2872dd

Observation dae05960-f941-4840-985e-c37e44273954 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:04.192757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:04.192757Z digest=sha256:a2e6bf38b09e247054eb18a1a90f2b5fd1325b530987b7f22b5efeeef70f60c6

Observation fcaac928-c242-44ff-ba16-b035f1054970 · outbound

This paper cites Vision transformer with deformable attention.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Vision transformer with deformable attention

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.710516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:04.293577Z digest=sha256:c7f601fec6e76dedd57b4375142d344cf64b7e00d8244e4fe2ee5f3b5355b253

Observation 6e3c69d9-d6f6-42f5-b21b-21bce7ccd6d1 · outbound

This paper cites Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.511209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:04.371035Z digest=sha256:abb8b13ae28577a0bba277fee04867ad254d5493ace37fbe5cb8e8176f5ad18f

Observation 542b7c04-9546-4022-bcc1-47839712a2bf · outbound

This paper cites Cross-category video high- light detection via set-based learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Cross-category video high- light detection via set-based learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.318203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:04.445706Z digest=sha256:d5165cd6d4dab419b56f221372424d1c89e51f06fa19c738c4ba786f6d1fccb5

Observation f6177304-f351-4ff1-9927-97870806d7c9 · outbound

This paper cites Unloc: A unified framework for video localiza- tion tasks.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Unloc: A unified framework for video localiza- tion tasks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:07.110265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:04.563858Z digest=sha256:ad28387b60c4c11391136c0d92958e3b5bf59c848a41f1f28446811b5f577bf9

Observation 15576df0-3d5b-4402-aa8c-5705814ab4a4 · outbound

This paper cites Understanding negative sampling in graph representation learning.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Understanding negative sampling in graph representation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:06.893073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:04.698446Z digest=sha256:f9296412bb64a346a47f895140653a7e42159d164dcf4a0b7fefb863b4ca6119

Observation a89e8865-aa01-4c01-bda5-86a15edc9dd0 · outbound

This paper cites Parameter-efficient is not sufficient: Exploring parame- ter, memory, and time efficient adapter tuning for dense pre- dictions.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Parameter-efficient is not sufficient: Exploring parame- ter, memory, and time efficient adapter tuning for dense pre- dictions

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:06.739159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:04.818146Z digest=sha256:00258cf457f46c921a0cb125674783e5905aeee99a65ac2c4cb2390fa4a20f2d

Observation d05ac6e5-583e-49f8-a42f-a901e059b357 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:04.919669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:04.919669Z digest=sha256:1a8452f0e2e72f41aa4de8ea2e4545d174f23662cce1d16458a2be2e4f4613fb

Observation f196f109-36e6-44b2-a4b1-68282aa34209 · outbound

This paper cites Conditional prompt learning for vision-language mod- els.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Conditional prompt learning for vision-language mod- els

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:40:06.521597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:05.020736Z digest=sha256:048244948068d9793521375788e8724e8d66a797bd9cc95b147b05c4628c372b

Observation 93b92d5b-bd99-4e09-8c12-71c2c64db6fa · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:40:05.125881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:40:05.125881Z digest=sha256:7315160dc8178f60306db76ee3aaa04829914789a910221e961bd87a853e7707

Observation 60b4fea0-4029-4292-816c-9e4963f77bea · outbound

This paper cites Notably, these losses are applied to all the intermediate layers inde- pendently.

Sparse-Dense Side-Tuner for efficient Video Temporal Grounding Notably, these losses are applied to all the intermediate layers inde- pendently

Reference 51

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:40:06.366128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T18:40:05.200664Z digest=sha256:4a6570de3e540b4f6dd0bc0f7a018ce55ec245ad342c25a7806a7251fa7a66c5

Pith citing papers

No inbound Pith citation observations are available.