Pith. sign in

Paper Citation Record · LEDGER

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding

As of 8 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 0 inbound Pith citation observations for arXiv:2607.13421.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13421 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T05:17:47.022013Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

63 of 63 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bce00fdc-7986-4685-885a-3ac341e6a2e9 · outbound

This paper cites GPT-4 Technical Report.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:39.096786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:39.096786Z digest=sha256:0798928603271079116109f8bc17d9a68b18df4f03407565a56838a7f5f88f05

Observation 56b3b335-fe7c-481a-b9b3-4592bcad190e · outbound

This paper cites In: ICCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ICCV

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:39.159518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:39.159518Z digest=sha256:fef8c6471261a54ce6f8a2066f50c5052aff640c5014500ed9d6f3d8ba92bde9

Observation fffcc511-4cbe-4e95-a81d-6f627784c3b2 · outbound

This paper cites In: ICCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ICCV

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:39.307561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:39.307561Z digest=sha256:99aca4b21eded19115fad57f17f84aa2d729f8893a0e448e68a300732f948bc2

Observation 04ac3970-7b4a-4a7d-9a4f-801259c6d0df · outbound

This paper cites Qwen3-VL Technical Report.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding Qwen3-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:39.479468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:39.479468Z digest=sha256:341c1eb693f267913638ab0694075cff523d47918c766ac97d51553325061a0c

Observation 8697d48f-f578-4fb9-a8c3-f6ec62328d30 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding BEiT: BERT Pre-Training of Image Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:39.663567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:39.663567Z digest=sha256:fcf0ee882ad71c31252a00b161202e6b02312a12e26d5cd0b9658fbf66dbb1d9

Observation cb7596d7-befa-411f-b099-1aebc97e5fb3 · outbound

This paper cites In: ICCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ICCV

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:39.834142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:39.834142Z digest=sha256:b2eaa4156c85d27460aadc8c7d75b7bd68db6585854ce359061dee36b701f490

Observation 70f0c01c-5073-4f14-8748-b2c0945a8cdb · outbound

This paper cites In: ECCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ECCV

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:40.021515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:40.021515Z digest=sha256:3fd75d11489eb19b29dc59e8ebd961c74b5d8e37f383eb30121cac779608e74c

Observation ea028b39-321e-4507-abd9-9acd9ae96f91 · outbound

This paper cites In: WACV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: WACV

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:40.191155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:40.191155Z digest=sha256:e2508d229adb8a8eab971c82f191cf38eae280965fd3d4c2c1041b4239dfeaef

Observation f927b9b4-d763-4058-ad52-0615401eb269 · outbound

This paper cites In: ECCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ECCV

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:40.283064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:40.283064Z digest=sha256:2604be3154b92cbc53d494750dcb9319bb0fa66662e897b5385718c2fa323a84

Observation d0f2c486-0e70-4e63-86ad-f3dc94ef2ad2 · outbound

This paper cites Chen et al.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding Chen et al

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:40.405951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:40.405951Z digest=sha256:fb08af2ffa49b2cfd007626c53248ce2e0d41ca5881aeb11dc77c12f75096d81

Observation dff85100-7380-46b0-b12d-bc5a3b20f351 · outbound

This paper cites NeurIPS34, 28442–28453 (2021).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding NeurIPS34, 28442–28453 (2021)

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:40.564173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:40.564173Z digest=sha256:6d74931bc1dadfd9cf1d506ef25996e1569f3587ba83b5a01533e0a276b14ed7

Observation f2ce36e5-bc24-4a4c-a23b-038724aebe86 · outbound

This paper cites TCSVT (2025).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding TCSVT (2025)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:40.714933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:40.714933Z digest=sha256:de78d676ee02733acdb1cb532d5ea87212c5eafd9294efd0d9bd81e597399a3c

Observation 74fad117-945d-42ae-855b-c5f7cbcec15f · outbound

This paper cites TPAMI (2025).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding TPAMI (2025)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:40.859781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:40.859781Z digest=sha256:9352013e92246d7f5bc6a0f265179953cdce96694a00bb8fa61efefce12c820d

Observation fbfc2b56-d228-4b02-9781-cb0f45497a71 · outbound

This paper cites In: ICCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ICCV

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:40.975393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:40.975393Z digest=sha256:945e63821ac526e3ec85612a1f899fe3f54910bbe0986a4c8abe31b607bc4c3a

Observation 86077990-35d4-4812-85bc-e1b74e19382a · outbound

This paper cites NeurIPS37, 121670–121698 (2024).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding NeurIPS37, 121670–121698 (2024)

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:41.087192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:41.087192Z digest=sha256:494d03cb99ec8d3f941f8da1da005d9460cc82ffb509c4033f9d8b2b3e0525c3

Observation 07180e20-0350-4de4-a1a9-00e6c3a3e94c · outbound

This paper cites arXiv preprint arXiv:2510.09274 (2025).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding arXiv preprint arXiv:2510.09274 (2025)

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:41.238768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:41.238768Z digest=sha256:97f326a465528400a84fe5a5abfc82bfe2d7b076814dc25e586b0f9502ff7cc5

Observation 0ffa4f0b-e9c3-4e6f-b946-850a3e492767 · outbound

This paper cites In: ICCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ICCV

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:41.382381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:41.382381Z digest=sha256:8b7b5954314e35c7ba58d0f1e74378af2c34361e07990f52f6238b64f397f1e2

Observation c4c6eadf-ca31-43d6-af6b-5b5b0bf2afaa · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:41.568196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:41.568196Z digest=sha256:15b8f630af6d512581b174df893f1781d0f11ca73890ca08c8b14ee9b7b211ba

Observation 5f80730b-f7e7-4cb2-9f6c-a291497d1aa4 · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:41.672525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:41.672525Z digest=sha256:761c53af4e244ca1d407318093bcbfd2e72d3685bbc2d295a70599f154c05a65

Observation 17ded26d-5018-4dd5-af56-21810e6032ed · outbound

This paper cites Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:41.794200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:41.794200Z digest=sha256:1bf057ecee872d277c4b12e6bb87cbfa492ade058fa0003356e18e18c6483029

Observation fb8ce704-8bf2-421d-933d-c4ffb203fab2 · outbound

This paper cites arXiv preprint arXiv:2511.21375 (2025).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding arXiv preprint arXiv:2511.21375 (2025)

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:41.935783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:41.935783Z digest=sha256:e4475d3a5a74d2524308780fbc8d41d543b017320006b6e77acc7d42d588761f

Observation e7ea3c9f-5bb8-48e2-978e-96ec6f2459a8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:42.055938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:42.055938Z digest=sha256:ff0ad4fac900ed37be5591d33621dfc95bd52c9c6eb2265663fbca8b398e9f6e

Observation 70b87539-3223-4dfd-a4e8-be4a7401e44e · outbound

This paper cites Seed1.5-VL Technical Report.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding Seed1.5-VL Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:42.165867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:42.165867Z digest=sha256:58f37f2faca131da67c705f5ae2b2f5f4e146476917fe349c4590aaafac33d9e

Observation ea60af27-848d-4dc0-aa7c-c1f7fe55f970 · outbound

This paper cites In: ECCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ECCV

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:42.285710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:42.285710Z digest=sha256:9a4e3a47af34c83e9cd4c1c2945f25b4ba75b6f0a333d4c7834f100f0d0edd5e

Observation a511c4cc-8283-4071-96a5-8949f2af4a19 · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:42.437338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:42.437338Z digest=sha256:61ba49ca5a0980340233a800af63f88ab28db1d57710b4cd10374aa120f2e0db

Observation 2092d07f-4e57-4af6-9d6a-a0ea7fe82cd0 · outbound

This paper cites NeurIPS35, 29192–29204 (2022).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding NeurIPS35, 29192–29204 (2022)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:42.527596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:42.527596Z digest=sha256:e4a3ccfa33dbaa79c9ef5144ac6e1729b34dca27aaad3da4724fc37507fd526e

Observation 2a6e2998-b0da-4cae-939e-763d6786a99d · outbound

This paper cites In: ICCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ICCV

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:42.605078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:42.605078Z digest=sha256:0e5c0e9b5953563dd663cb8f717a4efda844e36f9fe3140b320838135750e240

Observation f02887a4-993c-4f0d-9cce-9725f2e468ba · outbound

This paper cites In: ECCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ECCV

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:42.718826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:42.718826Z digest=sha256:9f76c7c7e0d183d94c9307d981fdf53676200ff5e62cacae8e6811f69f8410f8

Observation 3b75e24d-8bfc-402f-ab70-930e72108dd1 · outbound

This paper cites NeurIPS34, 11846–11858 (2021) ScanFocus: A Coarse-to-Fine Framework for STVG 17.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding NeurIPS34, 11846–11858 (2021) ScanFocus: A Coarse-to-Fine Framework for STVG 17

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:42.859875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:42.859875Z digest=sha256:267b864456e93a35debb26db34fca2ae5a728941f2011c961025d64810d1bf7f

Observation 4bfba24c-2ee2-403c-ba0a-2fcc4b01b495 · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:43.037933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:43.037933Z digest=sha256:83e50814e3562ecf96cf5264806559a5a45b745615b3d0216fe02d4a052199c8

Observation 1391d910-287b-489d-acb0-a217d7bd76cc · outbound

This paper cites IEEE transactions on medical imaging43(1), 96–107 (2023).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding IEEE transactions on medical imaging43(1), 96–107 (2023)

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:43.193557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:43.193557Z digest=sha256:0f6eaed754d59b7c09019cdeca4da9eb2802a88375770ec3eea35f96c69eb362

Observation 91efcafc-e575-4ddf-8871-044421af8dde · outbound

This paper cites In: ICCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ICCV

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:43.313101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:43.313101Z digest=sha256:dbf938039b7d4c0a3080ee7e1c89dc09fd4b8e1e54cdb7f35e1f595347a9f1f3

Observation a43a8f1c-e8f6-4fb2-98ba-c9e319a55292 · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:43.436505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:43.436505Z digest=sha256:d9560808809033c4d21b3e7ea689d92538c94ac469e0863143837e64ec3ac4fe

Observation 62aa8eb0-9b51-4253-8b78-b214b2fb319d · outbound

This paper cites In: ECCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ECCV

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:43.581486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:43.581486Z digest=sha256:8b7f12494bdb9b089bea28edd46a8d477e8a638855216294c44725372c7b4815

Observation 633ecaa7-960f-410a-bb1a-b5a1dfa67c29 · outbound

This paper cites Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding Correlation-Guided Query-Dependency Calibration for Video Temporal Grounding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:43.721458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:43.721458Z digest=sha256:18638d980de81519bffa3d02d047087cb7e4899a9a3dd19aac31bbfd607bf4c9

Observation 5d430845-3463-408d-8b86-c14b28184ec0 · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:43.841313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:43.841313Z digest=sha256:3dbd09e925cb10bb665d4ac3362e4f7052c48596ccb7820178d5f681eff560d5

Observation e91696a6-37f8-4e1e-a3fd-34901271d027 · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:43.958043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:43.958043Z digest=sha256:09d16a906bd2f5e28a74d586fe341d8ede0c50253a0ead32bf308a920c24f1ef

Observation 248b5e85-54e8-49e4-8439-d6c13b63b94f · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:44.036849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:44.036849Z digest=sha256:747c826899cfd574cc7b67718d89e5d0ef2bcc0b8bc82c4caf18c0f94fcb4d72

Observation 4f1189d7-dc7b-4e42-9316-7538e753655f · outbound

This paper cites In: ICCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ICCV

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:44.215895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:44.215895Z digest=sha256:948619fa5eb8487cec0ce1a4d922d9a773f8bb64243d5d96aa9dab7a03df631a

Observation b84e8799-72d9-4cae-a1c5-4827fec583b4 · outbound

This paper cites In: ICCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ICCV

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:44.305497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:44.305497Z digest=sha256:4143ab71e377c97973583f81a85439d6776808549ef4da3b7850ceb8c28b3c95

Observation 4b42ed6e-1f69-4ac5-98be-1b9421207f4e · outbound

This paper cites Augmented 2D-TAN: A Two-stage Approach for Human-centric Spatio-Temporal Video Grounding.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding Augmented 2D-TAN: A Two-stage Approach for Human-centric Spatio-Temporal Video Grounding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:44.381877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:44.381877Z digest=sha256:8490d25b3b49124eb0884e9b9b19a55b75d356eb63d21781220e889394f6c6d1

Observation 4f3c1bbe-7f2e-4ad6-830c-ed2a24a2b8ea · outbound

This paper cites TCSVT32(12), 8238–8249 (2021).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding TCSVT32(12), 8238–8249 (2021)

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:44.548749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:44.548749Z digest=sha256:0e2a7b64fa02dd11229a0a4bd31214cb2eec832e203f5854ea31d288f11a9868

Observation 2ef85720-1114-4d63-be36-7eb41c8e011c · outbound

This paper cites NeurIPS35, 10078–10093 (2022).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding NeurIPS35, 10078–10093 (2022)

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:44.650534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:44.650534Z digest=sha256:bcf5fcb5767b693ea1d6f7fc50759a46a91c1169934165888dac0739787f4aed

Observation d402bc82-b25d-4bad-9c9f-1173066318f8 · outbound

This paper cites NeurIPS30(2017).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding NeurIPS30(2017)

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:44.766279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:44.766279Z digest=sha256:6741146887d379f748109eb8e20f8883d61281d04777f80260f4f221306cf1c1

Observation dc37d913-647b-4a08-b324-53b57bf945ee · outbound

This paper cites SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:44.897525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:44.897525Z digest=sha256:634ba524aab4ad40ca0a1e020eecd953588a0f75ed0ba261665adb53b46eb0ee

Observation 8759f66a-83b3-4a91-a558-4c509a1e19ba · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:44.972677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:44.972677Z digest=sha256:b71c490aeb0b42e52f8248cd0ab3f54ee300c6f2bb895a485f4196974462dbc7

Observation d3481a49-e06b-42db-b1af-a13c5b8856e7 · outbound

This paper cites In: ACMMM.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ACMMM

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:45.054602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:45.054602Z digest=sha256:0c164be1aa61e72855ce73c4dafac017212cc403399efad1717d4c647be308a7

Observation af2658b8-168c-4d1e-b088-82f8d523a586 · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:45.112372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:45.112372Z digest=sha256:fbb69d86c061376e33aab0209b55867bacb13ecd57a9cdb8f962313bbca73699

Observation a1663ef7-0790-48de-a90a-d1dfa8c091e1 · outbound

This paper cites In: AAAI.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: AAAI

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:45.203736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:45.203736Z digest=sha256:56a6195087e804722315466ba83a4f562b299f2a5245d7be0d53272f1d422a3f

Observation f47ade2c-c816-4c00-aecc-5090a1676721 · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:45.301585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:45.301585Z digest=sha256:2b639b03bc97f64783a9140d37fd9e3642df7eaf24fe774d9e5525f4d160cbcf

Observation c097ae7e-8a10-4049-ae4f-60ed5d11d1dc · outbound

This paper cites In: ECCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ECCV

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:45.452568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:45.452568Z digest=sha256:4e1c0b349b62445eadbb93141c9aae7c8f20ff22f8416546724b5de5906375cc

Observation e73b660a-ce54-4aa4-b79e-17f59402b661 · outbound

This paper cites In: AAAI.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: AAAI

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:45.538491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:45.538491Z digest=sha256:480b569c795be978578f6511d3f18baf38edc778745b6046d3498eaf5e4912d7

Observation e73249df-0c27-437b-a006-1f356d2854c1 · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:45.677072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:45.677072Z digest=sha256:0bf185d286de6b5c14de166ea93338462c43bd1d101415fa566083ef6d9b0a4c

Observation 533acf56-4a6f-4539-a65f-eef1bc03081a · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:45.777515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:45.777515Z digest=sha256:c3c51b437ed9de185b186d63b866505296477876ee3a99dd01b9e998aa501e9b

Observation 8d38f1f0-76ef-4a0d-8297-c4c54d19845d · outbound

This paper cites 2rd Place Solutions in the HC-STVG track of Person in Context Challenge 2021.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding 2rd Place Solutions in the HC-STVG track of Person in Context Challenge 2021

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:45.903712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:45.903712Z digest=sha256:8f6a9c6471170ff4e137eca1477719f5af438a8a3d1aefb88716eb15fa80df0a

Observation c0885fc8-ae96-4c6f-bb83-3c209acb3f7b · outbound

This paper cites In: AAAI.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: AAAI

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:46.092194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:46.092194Z digest=sha256:a54e2e32693260c41284bf6c7f967fe3595bf3e29327de567c1a9702f5d36b7b

Observation b7ebfe4d-e444-45b3-bdff-a587eca97c71 · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:46.225036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:46.225036Z digest=sha256:951925ed8c2c49df3748f908c825049eb5e945538327865c1633751fd34139cf

Observation b5384dc5-75c5-4305-92b2-3f1b2d52ecbd · outbound

This paper cites Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding Object-Aware Multi-Branch Relation Networks for Spatio-Temporal Video Grounding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:46.357529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:46.357529Z digest=sha256:6f420b680820d1af2a93067170a27d29ea38860762d725f40945ac7c02119d7a

Observation 9f05630d-a843-4d86-a813-9f87112a3abb · outbound

This paper cites In: CVPR.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: CVPR

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:46.502729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:46.502729Z digest=sha256:d557ca069bb2dafe9cfd0f12863135087356eaef052780bab39ac4eae90a19f1

Observation fbc05517-7d50-4e30-86d6-ea086a89211f · outbound

This paper cites arXiv preprint arXiv:2602.13313 (2026).

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding arXiv preprint arXiv:2602.13313 (2026)

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:46.678128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:46.678128Z digest=sha256:7526e3860cae5a69db19e1a2e72157a95b78b182666a7589c59d8c4dda3f8763

Observation 5ba1d294-ebe0-4644-9933-fa8fbc3eed6e · outbound

This paper cites In: ECCV.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding In: ECCV

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:46.771301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:46.771301Z digest=sha256:b5b7606692a300758a1e4421747bcbb7ff41421d722beedbb8b49555050d94b2

Observation bdf53ae4-c518-48c1-b5d5-67657288b62d · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:46.866540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:46.866540Z digest=sha256:b92d508fd9ac4597fae3a698bfebb8970a44a9a4d1ea068360088d48bc0a68f9

Observation 3b6c098c-87ad-4fae-9f06-242326e05841 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

ScanFocus: A Coarse-to-Fine Framework for Spatio-Temporal Video Grounding Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:47.022013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:47.022013Z digest=sha256:3acc63cf969d1e6630617a3617782df64d36ca5e6efef5db54fa305ccbc967f6

Pith citing papers

No inbound Pith citation observations are available.