Pith. sign in

Paper Citation Record · LEDGER

Number it: Temporal Grounding Videos like Flipping Manga

As of 19 August 2026, this Paper Citation Record lists 100 of 101 outbound references and 6 inbound Pith citation observations for arXiv:2411.10332.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10332 v3

Coverage vector

measured 100 of 101 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:49:43.782509Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:59.262348Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T09:43:49.389473Z

Reference resolution

100 of 101 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved65
  • parse uncertain0
  • malformed identifier4
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e27ab6e-8d49-4f7e-a5b2-f40d22d73eb7 · outbound

This paper cites Qwen Technical Report.

Number it: Temporal Grounding Videos like Flipping Manga Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.271932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.271932Z digest=sha256:f3406bc9289d8b246001145b24fbae611c1cb0a263cb77e67650614a32ebf12a

Observation 72e74eea-765b-4db3-9a77-090e307a35c5 · outbound

This paper cites RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search.

Number it: Temporal Grounding Videos like Flipping Manga RaSa: Relation and Sensitivity Aware Representation Learning for Text-based Person Search

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.277756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.277756Z digest=sha256:4ea3f7d28ece873d2e4191dd2aa90776d69d7e1e81cabd1413e95a9449972e82

Observation 51485249-0540-4699-b758-f7d90e4a49e0 · outbound

This paper cites The surprising effectiveness of multimodal large language models for video moment retrieval.

Number it: Temporal Grounding Videos like Flipping Manga The surprising effectiveness of multimodal large language models for video moment retrieval

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.283074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.283074Z digest=sha256:7105e0256e45afe9256912c36a2a4db489c09565b77eeccfee7aa171974c8ce0

Observation 6d94de45-1e3f-4ae3-8853-9f4724b5450d · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Number it: Temporal Grounding Videos like Flipping Manga Activitynet: A large-scale video benchmark for human activity understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.288171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.288171Z digest=sha256:26a07803b147807cdf279a321202f753cb0d7ee043759b26d56dee4252ca11dd

Observation 69faa14f-5026-4bd5-a722-8b1a8fb90c15 · outbound

This paper cites Vip- llava: Making large multimodal models understand arbitrary visual prompts.

Number it: Temporal Grounding Videos like Flipping Manga Vip- llava: Making large multimodal models understand arbitrary visual prompts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.293876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.293876Z digest=sha256:b50248d08a7d95fc35f16311035bd56a7746b5cc4a60ff3d71e210c552641e8f

Observation 7096712d-dc08-4519-973f-04955a4c4edb · outbound

This paper cites Progressive bilateral-context driven model for post-processing person re-identification.

Number it: Temporal Grounding Videos like Flipping Manga Progressive bilateral-context driven model for post-processing person re-identification

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.298855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.298855Z digest=sha256:e73a4dc01e69a61803260e354596b0d4f68a9d9dd409d4474977a05a2ed8f166

Observation 4654a0f1-076a-47c9-94ac-674fa74120fc · outbound

This paper cites Image-text Retrieval: A Survey on Recent Research and Development.

Number it: Temporal Grounding Videos like Flipping Manga Image-text Retrieval: A Survey on Recent Research and Development

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.304365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.304365Z digest=sha256:65dd2465b83d9f50980065e994c9647cd1796de2e4684e997a40cf119ff3024e

Observation 2c1476d4-0379-45ec-ab13-99bb0fa5505a · outbound

This paper cites An empirical study of clip for text-based person search.

Number it: Temporal Grounding Videos like Flipping Manga An empirical study of clip for text-based person search

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.309359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.309359Z digest=sha256:b66f650e0307caa55974a83735871c7ad664ea2ca7ce843bd539c982b40c4598

Observation 7b1695a7-7ab0-4b01-96ba-62029f41cea9 · outbound

This paper cites An empirical study of clip for text-based person search.

Number it: Temporal Grounding Videos like Flipping Manga An empirical study of clip for text-based person search

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.314089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.314089Z digest=sha256:149b722c33783ba0e54c467a8ed7a797c8933092cc8f5d2ec4bcaf5b54d5e02a

Observation 7aefe83a-b8f4-464e-bd6c-c8a679f34b62 · outbound

This paper cites End- to-end multi-modal video temporal grounding.

Number it: Temporal Grounding Videos like Flipping Manga End- to-end multi-modal video temporal grounding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.319675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.319675Z digest=sha256:08ba9fe62ec274266fe92a41144f7d5f140ea96439d765c2cacc157b72b9b68c

Observation de3edf37-4da3-4ce3-b00c-c91aa7b91e7b · outbound

This paper cites InstructDET: Diversifying Referring Object Detection with Generalized Instructions.

Number it: Temporal Grounding Videos like Flipping Manga InstructDET: Diversifying Referring Object Detection with Generalized Instructions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.324512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.324512Z digest=sha256:9db4c59f7c82bd2b074d1b173a81110e59ba44cfae20585338f5152acf995a30

Observation 30a36257-8ce9-444f-929c-9cc9ffd8e0e0 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Number it: Temporal Grounding Videos like Flipping Manga An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.329699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.329699Z digest=sha256:3b6d25d8a9aee8ac0ad8c5846ae5c2c69ed50cc1e0f052d18baca94e3f67aa79

Observation 9b2b6afd-e249-426a-9884-5d899c6e2245 · outbound

This paper cites VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models.

Number it: Temporal Grounding Videos like Flipping Manga VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.334478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.334478Z digest=sha256:3a40eabb417566788642eb32c364611b5ac4ac99545485849fb5ff5ac5eef4be

Observation 55ff0f91-8ea2-412b-9bc5-b194286f9130 · outbound

This paper cites Cityllava: Efficient fine-tuning for vlms in city scenario.

Number it: Temporal Grounding Videos like Flipping Manga Cityllava: Efficient fine-tuning for vlms in city scenario

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.339509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.339509Z digest=sha256:759e4846ff303c8e3546ed4752cbba86cd60e1e578cec663418b672bec7a008e

Observation 2b825821-aa6b-4a71-bc83-42cdc2fafece · outbound

This paper cites System- status-aware adaptive network for online streaming video un- derstanding.

Number it: Temporal Grounding Videos like Flipping Manga System- status-aware adaptive network for online streaming video un- derstanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.343963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.343963Z digest=sha256:ca44be8de6233eb24dc0c59512dbd3dbfed8397b278307930ddcd92810db8e68

Observation 19655bc8-82cc-496c-8605-de77949e833b · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Number it: Temporal Grounding Videos like Flipping Manga Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.348499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.348499Z digest=sha256:91849cc531cbd46a919b716852cfd7bc3b604c5c0fb26581184a3a8282c95768

Observation 5c093acf-5c4d-4376-9079-b5a4e9ff882b · outbound

This paper cites Fast video moment re- trieval.

Number it: Temporal Grounding Videos like Flipping Manga Fast video moment re- trieval

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.353673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.353673Z digest=sha256:c79e741e51f48a54a0e6c7221b2b25ff9aaa7f44f26f966bea6fc6191fb66921

Observation 23988d94-7d0b-46c1-8ed7-326ba4e2628d · outbound

This paper cites Tall: Temporal activity localization via language query.

Number it: Temporal Grounding Videos like Flipping Manga Tall: Temporal activity localization via language query

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.358620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.358620Z digest=sha256:2acf6d5d73bd3d3d0d072202c7f3b322213f97ce5d22b24d5bc4db7c5b409959

Observation bd6fdd50-76ce-4130-83e9-4fb054b96c1a · outbound

This paper cites Scaling New Frontiers: Insights into Large Recommendation Models.

Number it: Temporal Grounding Videos like Flipping Manga Scaling New Frontiers: Insights into Large Recommendation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.364700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.364700Z digest=sha256:12754aebc27929ed5c497a0e15158040d0e529a64fd810585f31d8d8f94db336

Observation 94d173e7-b44a-440d-987d-b0353c541ff7 · outbound

This paper cites VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding.

Number it: Temporal Grounding Videos like Flipping Manga VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.369881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.369881Z digest=sha256:2033e4b6f306b4aba4738b1ef36fd616f233a6dbf0b7ab30b6f9e01f945df2b1

Observation 63b88a5e-6050-47ba-a9c7-d690d850d67e · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

Number it: Temporal Grounding Videos like Flipping Manga TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.374768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.374768Z digest=sha256:2e8d2f734fe7f24cb118150f35c3226e011090d5a8f60a977d9ddde7ebd2c409

Observation 28b9663f-688b-420b-8df6-2c973fdb6a44 · outbound

This paper cites Homeomorphism prior for false posi- tive and negative problem in medical image dense contrastive representation learning.

Number it: Temporal Grounding Videos like Flipping Manga Homeomorphism prior for false posi- tive and negative problem in medical image dense contrastive representation learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.380500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.380500Z digest=sha256:6a9af4d3c6004f41694cdc26a0ff3581253973b39fcf5049dfd9f920e68d49a9

Observation e54aaa00-c38d-4514-a277-f47c263a2bd6 · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

Number it: Temporal Grounding Videos like Flipping Manga CogVLM2: Visual Language Models for Image and Video Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.385753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.385753Z digest=sha256:249c77f266f8376f6df4ee3a5e8c1c152a00826b44bce14263a51c3038244894

Observation 7033bc6f-0678-4306-9ec1-1ea8e26ab70c · outbound

This paper cites Lora: Low- rank adaptation of large language models.

Number it: Temporal Grounding Videos like Flipping Manga Lora: Low- rank adaptation of large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.390812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.390812Z digest=sha256:1ab959af4648870b4280d6fd8d873e77aa342a82407ae24a2d55737620a4d55f

Observation 7e740943-f226-4117-ae00-bc2d36886610 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

Number it: Temporal Grounding Videos like Flipping Manga Vtimellm: Empower llm to grasp video moments

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.395369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.395369Z digest=sha256:635bb60615cff64cc7d3c8e81b8f0fb1236096823bee12c333979491ff6f89eb

Observation 5e2bf25e-8a6b-4cb6-aa13-0a76dce4666a · outbound

This paper cites LITA: Language Instructed Temporal-Localization Assistant.

Number it: Temporal Grounding Videos like Flipping Manga LITA: Language Instructed Temporal-Localization Assistant

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.401035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.401035Z digest=sha256:2c304d6976a72ad8f4730367f9c6eacaafb070c17a68763c695097fcb7fb9c1c

Observation 506ce60d-03f8-44f0-bde1-05b22e3252e7 · outbound

This paper cites Dg- pic: Domain generalized point-in-context learning for point cloud understanding.

Number it: Temporal Grounding Videos like Flipping Manga Dg- pic: Domain generalized point-in-context learning for point cloud understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.405965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.405965Z digest=sha256:5ad5e2c2928a42fc08c31434fbbfeb00d16958e08927d18638518c217d44bfc0

Observation 1a7d9609-adf3-48a2-8e71-24552839217c · outbound

This paper cites Do you remember? dense video captioning with cross-modal memory retrieval.

Number it: Temporal Grounding Videos like Flipping Manga Do you remember? dense video captioning with cross-modal memory retrieval

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.410751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.410751Z digest=sha256:d808d7fa41dfde129acd368968c877e8068989cbf8001f2818e1ec04d0774a45

Observation 624db7aa-a8ab-42bd-9fdc-c622c3847fe8 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Number it: Temporal Grounding Videos like Flipping Manga Adam: A Method for Stochastic Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.415479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.415479Z digest=sha256:cce10aa6298095287d353bcc300c52a3805b111dacc32734c77ed05c881b3d32

Observation 91cf0829-593d-4968-b7d4-0358d816f9f2 · outbound

This paper cites Multi-scale spatial-temporal attention networks for functional connectome classification.

Number it: Temporal Grounding Videos like Flipping Manga Multi-scale spatial-temporal attention networks for functional connectome classification

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.420390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.420390Z digest=sha256:d3aa01ce32d9d993eb515c3319390b3cc3dca694c176ae3de91022e4ce888db9

Observation 8f4931d5-9099-49bb-ad9c-1c7491b00e05 · outbound

This paper cites Dense-captioning events in videos.

Number it: Temporal Grounding Videos like Flipping Manga Dense-captioning events in videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.424914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.424914Z digest=sha256:0eef4f0f74837897979180bb7fb3675686070fb98f6fc9d7540994f0127d613e

Observation e4b763a7-98c6-4ce9-939a-2904564c83d9 · outbound

This paper cites CoLLaVO: Crayon Large Language and Vision mOdel.

Number it: Temporal Grounding Videos like Flipping Manga CoLLaVO: Crayon Large Language and Vision mOdel

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.430211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.430211Z digest=sha256:655b538eab9d4bed8ea6ffaa32bee572123563f6507931f17116f9ecbca13705

Observation 507369f4-f904-498e-9ab2-1e0469f59ca3 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

Number it: Temporal Grounding Videos like Flipping Manga Detecting mo- ments and highlights in videos via natural language queries

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.435396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.435396Z digest=sha256:f7de02e1c4a896d858c57afd0b8d533cd45d7d03c7a6a12b91a7c5f06e9ced1e

Observation 59561a61-f99b-4cab-94b1-5289163a3801 · outbound

This paper cites Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition.

Number it: Temporal Grounding Videos like Flipping Manga Frame Order Matters: A Temporal Sequence-Aware Model for Few-Shot Action Recognition

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.440209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.440209Z digest=sha256:53eb2ee9b825f4306713dd3f256448637b9a850d98c2a2e12ec1b136e7d5d6d4

Observation f957d1ad-6ae8-4ea1-b745-919a0d9f1367 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Number it: Temporal Grounding Videos like Flipping Manga LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.444946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.444946Z digest=sha256:2b7cb441cd2fe40b6e3dc0e7a348db7debcf4613825fadc3b39b5b5aacf3b094

Observation 0d53bb1b-9d8d-4f2d-893c-463e4dea03e7 · outbound

This paper cites LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding.

Number it: Temporal Grounding Videos like Flipping Manga LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.450680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.450680Z digest=sha256:11cbb03cf93d3959eb85a57ce76098a843750292d2ba437781f5b8cbf8dd4dc4

Observation 09a2978a-0dab-4de4-80b6-81a87fd1baf6 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Number it: Temporal Grounding Videos like Flipping Manga Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.455776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.455776Z digest=sha256:d9f90b94feaa4e37d9fc4893f1aed62dfbb50c4ca0fd523c597fcbf8cf62b172

Observation c62509b5-0f60-4f87-b29d-4971869331d7 · outbound

This paper cites Learning semantic- aligned feature representation for text-based person search.

Number it: Temporal Grounding Videos like Flipping Manga Learning semantic- aligned feature representation for text-based person search

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.460835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.460835Z digest=sha256:ead326308791f718e7725efe7d5d9c878e27f6ec26daa2b3d297ef9a5383a3ab

Observation b68bf659-b23a-4b9d-9e04-349e81444b5a · outbound

This paper cites Tea: Temporal excitation and aggregation for action recognition.

Number it: Temporal Grounding Videos like Flipping Manga Tea: Temporal excitation and aggregation for action recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.465697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.465697Z digest=sha256:4e88c12abfd022296eab8c98de56410c4da41035e7b7e8bc8a498b7a3d4c8b9f

Observation a4355b3d-2707-4492-aa59-d4862cef34b0 · outbound

This paper cites GroundingGPT:Language Enhanced Multi-modal Grounding Model.

Number it: Temporal Grounding Videos like Flipping Manga GroundingGPT:Language Enhanced Multi-modal Grounding Model

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.470442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.470442Z digest=sha256:20dc2f41cca69bb84b11fed72fb58d3112b1c5cdc28622558444730e83c3f404

Observation 2fbc9ea8-719a-4696-b2e4-2b25d120c80e · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

Number it: Temporal Grounding Videos like Flipping Manga Univtg: Towards unified video- language temporal grounding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:45.217964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.475415Z digest=sha256:4c1580c13b93725feeda31b9f58394e937d2eb0a9ae3724a15e7fe237cd73164

Observation bcdd80b3-b468-44fa-b79a-3af037fee7c3 · outbound

This paper cites Microsoft coco: Common objects in context.

Number it: Temporal Grounding Videos like Flipping Manga Microsoft coco: Common objects in context

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.480438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.480438Z digest=sha256:5a6b600c9ce92802e91caaf75496cfa94c17ff992a14778861ce2ef3edd9a780

Observation 2a41c13f-4a9c-4a74-b462-dbe3afec3864 · outbound

This paper cites Visual instruction tuning.

Number it: Temporal Grounding Videos like Flipping Manga Visual instruction tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.485503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.485503Z digest=sha256:1b38559aa768d2e9b589206267001f577270454ead63663deab53ddf9342cbe7

Observation cb8adf0e-8be9-40d2-becf-9fa385417219 · outbound

This paper cites OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning.

Number it: Temporal Grounding Videos like Flipping Manga OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.490402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.490402Z digest=sha256:6ee7aaf4ed247297aff1e7baec09e21b2362c66a2a8400793cdf206792c393e5

Observation 6b993de5-2630-46c2-9c65-e5c372cc7511 · outbound

This paper cites General- ized video anomaly event detection: Systematic taxonomy and comparison of deep models.

Number it: Temporal Grounding Videos like Flipping Manga General- ized video anomaly event detection: Systematic taxonomy and comparison of deep models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:45.179638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.495792Z digest=sha256:5bf49f1e6e93e304e154764b2f46500f0d7f1bd347a3d04ecc60ed99d39b521e

Observation f832ca3a-3360-413d-bd6b-fff18b756119 · outbound

This paper cites Zero-shot model diagnosis.

Number it: Temporal Grounding Videos like Flipping Manga Zero-shot model diagnosis

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:45.161962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.500708Z digest=sha256:4e7806f47f28afe5c10263377b1b79ea6235121a3711b792be56f359cb830faf

Observation d37a9f99-af29-49d6-93c9-fc5a7b5a20d0 · outbound

This paper cites Groma: Localized visual tokenization for grounding multimodal large language models.

Number it: Temporal Grounding Videos like Flipping Manga Groma: Localized visual tokenization for grounding multimodal large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:45.143222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.505481Z digest=sha256:0875f9502ac91c0fd0a2cbd056a8eb6eed431a5ce8d9fe8f2ea1444f476d37d3

Observation 6c8126da-127e-4b33-bcbd-c8c00fd2848e · outbound

This paper cites Visual Perception by Large Language Model's Weights.

Number it: Temporal Grounding Videos like Flipping Manga Visual Perception by Large Language Model's Weights

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.510176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.510176Z digest=sha256:dfaf0ef0c727fc83fc8f4c5e29c8874a8a0dd446db99b38f385915219e1d33c7

Observation 39867ad3-f3eb-4a93-884d-f204bbca527c · outbound

This paper cites Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model.

Number it: Temporal Grounding Videos like Flipping Manga Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-12T19:49:44.112918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.515735Z digest=sha256:ac1f6959c50463eef6e00b73528d6c61f8b6b52b191143c1bacaf1613dfed439

Observation 792d7724-6f6b-4491-b9ad-593aef9b054c · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Number it: Temporal Grounding Videos like Flipping Manga Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.521301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.521301Z digest=sha256:4ab108750e2d55407994ec93954ea8a5efb67c6bb9157e2186a1a03aba06bf15

Observation 005a43c2-eed4-49bb-bea5-6cebb5c9849f · outbound

This paper cites Weakly supervised video moment retrieval from text queries.

Number it: Temporal Grounding Videos like Flipping Manga Weakly supervised video moment retrieval from text queries

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:45.126539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.526201Z digest=sha256:cc714dbfdfc87f2796f82799034d00a263a8c064fdba26be2dc050778ed39d5e

Observation 7ea4ed77-5b01-4a45-92fd-709ead73e3b9 · outbound

This paper cites Interventional video ground- ing with dual contrastive learning.

Number it: Temporal Grounding Videos like Flipping Manga Interventional video ground- ing with dual contrastive learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:45.109006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.530802Z digest=sha256:a3248ebe7c92f3538f35953a9acf09f13d752264a81d8041d2b64ff517b179c6

Observation b589c9df-8b6f-4af6-930a-2ecb8547af93 · outbound

This paper cites Hello gpt-4o.

Number it: Temporal Grounding Videos like Flipping Manga Hello gpt-4o

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:45.092223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.535375Z digest=sha256:81553edfec4f0dc170117aaece2d40296cea305c6bc2b76d362bf7e3ad4c1812

Observation d81928c2-bb55-4722-96b2-9db197dd2aa5 · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

Number it: Temporal Grounding Videos like Flipping Manga LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.539858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.539858Z digest=sha256:bc4291bbde06ab824bba3151bee03fecb7de4ef426e01cc9a36a6e2e4bd3b433

Observation 253ded45-7cf1-49f3-9645-b0508957ab5b · outbound

This paper cites Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning.

Number it: Temporal Grounding Videos like Flipping Manga Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.544990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.544990Z digest=sha256:79af1ac867974abfe47bddaa65be510b22fd9fe9fc648d934a98a9593f521e3d

Observation 8f88397d-ae84-459c-a6e7-fde74adeb53c · outbound

This paper cites Controllable augmentations for video representation learning.

Number it: Temporal Grounding Videos like Flipping Manga Controllable augmentations for video representation learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:45.074341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.550723Z digest=sha256:17d01770f0f8c6a467f05bd4dc46c1939fc49dabf94beebf9b31362ba90b7b87

Observation 970d9ade-5a9f-4d92-88ff-73c048c44c0b · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Number it: Temporal Grounding Videos like Flipping Manga Learning transferable visual models from natural language supervi- sion

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.555412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.555412Z digest=sha256:ab48d19a1bfb00a6bb099fdd37c56b7ab787b68b2cd41030bfea7e6f2dfc0b8e

Observation 55398ac1-70a7-4692-9547-ff71c03a510f · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

Number it: Temporal Grounding Videos like Flipping Manga Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:45.046475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.561089Z digest=sha256:39a845079727e83c5ad93353768744cbee8a3f49fcf566f49c36dc0af7244d79

Observation deb07073-8690-4d58-bb61-5a630f4dac5a · outbound

This paper cites Towards more unified in-context visual un- derstanding.

Number it: Temporal Grounding Videos like Flipping Manga Towards more unified in-context visual un- derstanding

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:45.027166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.565725Z digest=sha256:2c887e0c15ccd8bc11fbfa2fadc40f195f05fdb660501bd6158597aa0ed6a52c

Observation 0c444450-b4f7-4d1d-9dbc-13a5bae980b2 · outbound

This paper cites What does clip know about a red circle? vi- sual prompt engineering for vlms.

Number it: Temporal Grounding Videos like Flipping Manga What does clip know about a red circle? vi- sual prompt engineering for vlms

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:45.009434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.570449Z digest=sha256:ec5621a4db6f2306faa440676a9dd51fa17e4c8e4bc352fed1abb6d52c5c0ad0

Observation 0361c958-e6d7-46a7-894f-acddc024dae5 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Number it: Temporal Grounding Videos like Flipping Manga Gemini: A Family of Highly Capable Multimodal Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.575901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.575901Z digest=sha256:a786a3a9a7398f0f256e0e7ba4376e1d9c0cb1ed8ccaba5a727336fb37208734

Observation 6f30a34f-00be-44f9-a43e-390dfcaac459 · outbound

This paper cites Add-it: Training-free object inser- tion in images with pretrained diffusion models, 2024.

Number it: Temporal Grounding Videos like Flipping Manga Add-it: Training-free object inser- tion in images with pretrained diffusion models, 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.580783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.580783Z digest=sha256:f8abf6811b9ecaecad8f21f33f884252ba4e27ee29a0d066ddd8f8bebc7c9c90

Observation 640ad148-7c08-4f97-b350-e19e4b5c85f7 · outbound

This paper cites Towards Open-World Grasping with Large Vision-Language Models.

Number it: Temporal Grounding Videos like Flipping Manga Towards Open-World Grasping with Large Vision-Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.586214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.586214Z digest=sha256:e2b823fe165b9ba292437fc1e65028f19a031ac0ee0b4cb3ccf16484e6ed0bfc

Observation c8b08ce7-467b-493d-a67e-5a418909482d · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

Number it: Temporal Grounding Videos like Flipping Manga Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.592295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.592295Z digest=sha256:cd6841a99bcc2aa5c58c0941539038f1fb445effb47d10918ff55550290913f3

Observation a039a1d3-518a-447e-b081-8ed6ad15de62 · outbound

This paper cites Temporal segment networks for action recognition in videos.IEEE transactions on pattern analysis and machine intelligence , 41(11):2740– 2755, 2018.

Number it: Temporal Grounding Videos like Flipping Manga Temporal segment networks for action recognition in videos.IEEE transactions on pattern analysis and machine intelligence , 41(11):2740– 2755, 2018

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.980623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.598449Z digest=sha256:23fe6ef7c7a01ff2566da99d341db3455a9e37b3ec11d19fab2576f513a50d71

Observation d7919a77-a927-49c0-8fcd-2534bdace045 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Number it: Temporal Grounding Videos like Flipping Manga Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.604285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.604285Z digest=sha256:585ce1cb2ab7931f3e0c8a73912af39c293f686da8ecebed4f287ee392e1ae2b

Observation a8003260-5202-4e4f-85b8-9412403eb99f · outbound

This paper cites Visual- semantic network: a visual and semantic enhanced model for gesture recognition.

Number it: Temporal Grounding Videos like Flipping Manga Visual- semantic network: a visual and semantic enhanced model for gesture recognition

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.961107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.609223Z digest=sha256:b457be0198e1de235d54bf03eef25b1ea7b97ecd16d3dd5e38bc7cb5c3d853a7

Observation 5a20b3e7-ecab-4f1d-a804-bd3b2d5cc14e · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

Number it: Temporal Grounding Videos like Flipping Manga HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.614209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.614209Z digest=sha256:2c60c8353dd2263440cfdee179d3434582c063cbbccd4e08c128b2427025f9d0

Observation 2928fbed-084a-49ff-914d-306fff890e28 · outbound

This paper cites Omniedit: Building image editing generalist models through specialist supervision, 2024.

Number it: Temporal Grounding Videos like Flipping Manga Omniedit: Building image editing generalist models through specialist supervision, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.943556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.619235Z digest=sha256:1c539cd406a90010c800f72739b83089051ddd0931c440f3ad750965046cf87d

Observation 186fc9f2-b1e4-402a-bfff-25de3fcdb879 · outbound

This paper cites Visual Prompting in Multimodal Large Language Models: A Survey.

Number it: Temporal Grounding Videos like Flipping Manga Visual Prompting in Multimodal Large Language Models: A Survey

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.624115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.624115Z digest=sha256:b05968b990f6a7c3925b5cc0aac325ebaa24287e62277caf257ddb694eb7eb0d

Observation 198c9a2a-42d1-475e-a6fa-0bee1598c4af · outbound

This paper cites Data-Efficient 3D Visual Grounding via Order-Aware Referring.

Number it: Temporal Grounding Videos like Flipping Manga Data-Efficient 3D Visual Grounding via Order-Aware Referring

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.629655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.629655Z digest=sha256:596e355797d841926135d960fc5cea231b694ee2cfe7d85f9875d671f1ea2e57

Observation 15b4b763-163a-45b5-a43f-487c092bbcb3 · outbound

This paper cites A glance at in-context learning.

Number it: Temporal Grounding Videos like Flipping Manga A glance at in-context learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.925113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.634816Z digest=sha256:4ffbd978361b6244fa3a4f1560abaf0d604e7cc727aa717def5ac1a16bffe38b

Observation 6e877104-f1bb-4bc0-aa50-cf16b322ff04 · outbound

This paper cites DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM.

Number it: Temporal Grounding Videos like Flipping Manga DetToolChain: A New Prompting Paradigm to Unleash Detection Ability of MLLM

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.639447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.639447Z digest=sha256:fbf7d03bb1e2f34b79720373360d6a340063f12680df845479c6e9c6ca990a5f

Observation e1389df3-b9f4-46ed-9419-c89b6e100bb9 · outbound

This paper cites Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark.

Number it: Temporal Grounding Videos like Flipping Manga Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.644813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.644813Z digest=sha256:a41d1e7508aa98dad34dbf7f97fb71f10039c6fdac73761cd81c35780acfd326

Observation 1d9beef1-7b34-4fe4-b639-afb6a8227267 · outbound

This paper cites Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events.

Number it: Temporal Grounding Videos like Flipping Manga Sutd-trafficqa: A question answering benchmark and an efficient network for video rea- soning over traffic events

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.908762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.650031Z digest=sha256:837289feac62110b42554c3cdbf134b5819af9e35b3ccbf93487d5e334a2a51e

Observation c99c834b-cb81-439c-aaf2-c9d88864af74 · outbound

This paper cites Meta spatio-temporal debiasing for video scene graph gen- eration.

Number it: Temporal Grounding Videos like Flipping Manga Meta spatio-temporal debiasing for video scene graph gen- eration

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.891670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.655524Z digest=sha256:c74b6d690bae0db58ee569346a2ac91cdccd805ddada99fa269c961b4a14a1d1

Observation 6b132d9e-1fa5-467d-8084-1732054ac599 · outbound

This paper cites Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning.

Number it: Temporal Grounding Videos like Flipping Manga Vid2seq: Large-scale pretraining of a vi- sual language model for dense video captioning

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.660472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.660472Z digest=sha256:530a2d1ccdc7d1d35cc73dbe23a37eeadb992ddbbf82ef1ca2f012e4851e90e1

Observation 0d761af4-a624-420f-90d2-a8572c02712a · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Number it: Temporal Grounding Videos like Flipping Manga Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.666512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.666512Z digest=sha256:ef7f7aacee2d5ed0279fb5a6c52d75cb56e3991d031ca96a815419d8e6609972

Observation fd5fb59f-df58-49c6-9cbb-12e8d403fb8c · outbound

This paper cites Exploring diverse in-context configurations for image captioning.

Number it: Temporal Grounding Videos like Flipping Manga Exploring diverse in-context configurations for image captioning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.864498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.671820Z digest=sha256:0e997361c5116f6b5b172982136ffc0bd0779ceba82fdafa5e4bd1bd572dba7b

Observation eec21f06-3239-4544-8be1-58efdbbe9566 · outbound

This paper cites Cpt: Colorful prompt tuning for pre-trained vision-language models.

Number it: Temporal Grounding Videos like Flipping Manga Cpt: Colorful prompt tuning for pre-trained vision-language models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.677405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.677405Z digest=sha256:63bba00aa12167b2abd81a3e4b552455c40a1cb25cde13dca0c5b4d11471ea2c

Observation 0ac32572-b24b-4b83-a87a-a84f1530da4d · outbound

This paper cites Visual place recognition via local affine preserving matching.

Number it: Temporal Grounding Videos like Flipping Manga Visual place recognition via local affine preserving matching

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.836887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.682207Z digest=sha256:37f0f422a3aea3114661ac2ed45db1e15623e889767579fbd69c72db73dfd59d

Observation 375c7cc9-c1cb-4b33-8270-aa6aea22f330 · outbound

This paper cites Neighborhood manifold preserving matching for visual place recognition.

Number it: Temporal Grounding Videos like Flipping Manga Neighborhood manifold preserving matching for visual place recognition

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.820197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.687779Z digest=sha256:40e67014b9cdc6c55fb6c5f073ee6e40fc404bf2aff434f73f3e0629a3628183

Observation dd6aaa0c-f58e-4d35-b0aa-e71c86ce244e · outbound

This paper cites Vqne: Variational quan- tum network embedding with application to network align- ment.

Number it: Temporal Grounding Videos like Flipping Manga Vqne: Variational quan- tum network embedding with application to network align- ment

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.803501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.692662Z digest=sha256:cb75014d4df492dc977b177956a7c9b1b57044b17eff9b72325bf8b2e66732b7

Observation 0adc9ecf-94f3-49ba-a647-5a8dd444734b · outbound

This paper cites Bridge the modality and capability gaps in vision-language model selection.

Number it: Temporal Grounding Videos like Flipping Manga Bridge the modality and capability gaps in vision-language model selection

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.787295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.697250Z digest=sha256:f9a58c5f58326462615b96ae9bad307d7c6607bc39d00fa6a2a0fa145b336f0c

Observation 2562db88-e2e8-4df8-aa42-7e69e0b33615 · outbound

This paper cites Dataset regeneration for sequential recommendation.

Number it: Temporal Grounding Videos like Flipping Manga Dataset regeneration for sequential recommendation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.768468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.702773Z digest=sha256:87cb939e7d76587eb00b29ccdf92c632d7642c286d4649846072c2266501c9e5

Observation 9878b8de-46dc-46ad-a9bb-a51c3af5f986 · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

Number it: Temporal Grounding Videos like Flipping Manga Hierarchical video-moment retrieval and step-captioning

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.751440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.707228Z digest=sha256:c65df08a0b49760cb6f6b8efbc45345f54a801049722676f5ecd15c115733ea3

Observation 1054c9b0-a73c-4304-839d-7d6752d46d73 · outbound

This paper cites Long Context Transfer from Language to Vision.

Number it: Temporal Grounding Videos like Flipping Manga Long Context Transfer from Language to Vision

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.711824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.711824Z digest=sha256:07fabcf2b35e6268e7a61ff598256f8c48503ae7354e56437a83d670558adcfd

Observation a6586124-9cc2-4de1-b69d-0be5a0440220 · outbound

This paper cites Pixel adapter: A graph-based post-processing approach for scene text image super-resolution.

Number it: Temporal Grounding Videos like Flipping Manga Pixel adapter: A graph-based post-processing approach for scene text image super-resolution

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.734674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.716781Z digest=sha256:a9bbde6a97063883d2bf13700ddc85f4bc1594ec640bbd1e1f13bdaad37de5bf

Observation db67889a-4ef9-412d-8c69-c44365f0aa2e · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Number it: Temporal Grounding Videos like Flipping Manga LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.721713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.721713Z digest=sha256:24d0741e80a53f77b6715e920d30fdc6a1a8c2af97179e57d7fb42e0d443f233

Observation 4259b911-136d-4604-b2e9-9b7ebb1dce2b · outbound

This paper cites Temporal action detection with structured segment networks.

Number it: Temporal Grounding Videos like Flipping Manga Temporal action detection with structured segment networks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.727289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.727289Z digest=sha256:2fd54a617f6f155aa61c1660ee7423bf2a4a261af9aafa3f21b2b2485ba9f1b0

Observation 3ed2a69d-cf91-43cd-88f1-5f4f38caada0 · outbound

This paper cites Multi-modal in-context learning makes an ego-evolving scene text recognizer.

Number it: Temporal Grounding Videos like Flipping Manga Multi-modal in-context learning makes an ego-evolving scene text recognizer

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.707038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.732725Z digest=sha256:3fc799a3028adf185c42d8b77bd6b1077c94c84a0dd1f4a3fa888ded0ed69f44

Observation e2ff5e7c-1844-4af8-b586-f0cd12f33379 · outbound

This paper cites MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World Control.

Number it: Temporal Grounding Videos like Flipping Manga MineDreamer: Learning to Follow Instructions via Chain-of-Imagination for Simulated-World Control

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T19:49:43.738235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:49:43.738235Z digest=sha256:069a8a579d2121cc03d4c9bb1d495ae354cb6743af7c197244ce596c84bbf49d

Observation c6e286cc-dbcc-4042-b0ac-841dd8c3d13f · outbound

This paper cites Moment Retrieval In the training-free (NumPro) setting, we extract frames from videos at 1 FPS, with each frame resized to a reso- lution of 336 × 336.

Number it: Temporal Grounding Videos like Flipping Manga Moment Retrieval In the training-free (NumPro) setting, we extract frames from videos at 1 FPS, with each frame resized to a reso- lution of 336 × 336

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.691069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.743874Z digest=sha256:65ebd0e1063d17f33a744f97f4a1fe860a30538fb8c934827f3cde9b1a4bdfbc

Observation 8388b3a9-1aa1-4bb9-ac04-c541875b6ddc · outbound

This paper cites from frame 000 to frame 200.

Number it: Temporal Grounding Videos like Flipping Manga from frame 000 to frame 200

Reference 94

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T19:49:44.675063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.748806Z digest=sha256:2b4e0641989ac082ddbf48d422c718f5534617f98bd36cf689b217d4587da31d

Observation 2f45607d-5dab-4c63-ab69-f38d8f5eb138 · outbound

This paper cites from 2 to.

Number it: Temporal Grounding Videos like Flipping Manga from 2 to

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.658471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.754945Z digest=sha256:a7bc52155a9ffe11bd9c9290acd7e16b1e6f86ae34752cb4dcacf284629b16dd

Observation 04f2c49e-fe5f-4251-96e8-5e20ea4bc4e8 · outbound

This paper cites We use 1FPS as the sampling rate, and adopt a design of red color, font size 40, and bottom right positioning for the number prompt.

Number it: Temporal Grounding Videos like Flipping Manga We use 1FPS as the sampling rate, and adopt a design of red color, font size 40, and bottom right positioning for the number prompt

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.639960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.759990Z digest=sha256:df6b589c9d6c95ae9174ba1cf399f43a8a4c1510360e870a2190f71ab5c21091

Observation 62fd4d25-418d-44b6-8af2-9bd64dc760c0 · outbound

This paper cites The results show that our method generalizes well across various General Vid-LLMs, achieving notable improvements in both mAP and HIT@1 metrics.

Number it: Temporal Grounding Videos like Flipping Manga The results show that our method generalizes well across various General Vid-LLMs, achieving notable improvements in both mAP and HIT@1 metrics

Reference 97

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T19:49:44.621085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.765727Z digest=sha256:2eefce3b6cadeeed30f66e7dcb4d6e9e93d6a65b599e8defb8b740c29e69347b

Observation 2b4ca4a2-edf2-43f8-8179-dce471846824 · outbound

This paper cites While a font size of 60 achieves better Number Accu- racy than a size of 40 (Figure 5 of the main paper), it reduces Caption Accuracy and introduces more outliers.

Number it: Temporal Grounding Videos like Flipping Manga While a font size of 60 achieves better Number Accu- racy than a size of 40 (Figure 5 of the main paper), it reduces Caption Accuracy and introduces more outliers

Reference 98

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T19:49:44.603843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.771309Z digest=sha256:5b854301d90d871709958ba71a7de9b543e7e1adaeac2bc7823954b0f675888f

Observation f9fb595a-597d-452a-9daf-29dea2b4b37d · outbound

This paper cites 10.5s”) may introduce decimals, which can increase parsing complexity for Vid-LLMs. In Tabel 9, we compare minute-level temporal annotations (e.g., “01:10.

Number it: Temporal Grounding Videos like Flipping Manga 10.5s”) may introduce decimals, which can increase parsing complexity for Vid-LLMs. In Tabel 9, we compare minute-level temporal annotations (e.g., “01:10

Reference 99

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T19:49:44.587142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.776940Z digest=sha256:1d864d37466abb28ebd65d348f8ca727656cfe930fc6426f73c10e62ecba0bc0

Observation ac78c114-2a4a-4414-aad3-d5b181948711 · outbound

This paper cites Dialogue Figure 11 illustrates a real-world application of our NumPro method within the Qwen2-VL-7B model, highlighting its ability to handle complex video-based dialogue tasks.

Number it: Temporal Grounding Videos like Flipping Manga Dialogue Figure 11 illustrates a real-world application of our NumPro method within the Qwen2-VL-7B model, highlighting its ability to handle complex video-based dialogue tasks

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:49:44.570644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-12T19:49:43.782509Z digest=sha256:245f06798af23e3b638bf1da096df9b8fb89ef7e1a5e8791c85d2c07039f753f

Pith citing papers

Observation 4e3c51c5-4ff5-45b9-bf7d-a5636c50b639 · inbound

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning cites this paper.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Number it: Temporal Grounding Videos like Flipping Manga

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.262348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.262348Z digest=sha256:898642a91840392870a3d56d674bc02475c8f02d2fa8ff0dac99e94a93a88c40

Observation e62bb179-dca9-4451-af74-35c82c258d88 · inbound

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought cites this paper.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Number it: Temporal Grounding Videos like Flipping Manga

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:45.397760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:45.397760Z digest=sha256:e4fe09a9709b06c2b02c478cdadd16125403c4fa01e06cf3a2eef2237fa2cd4e

Observation 7e1acb8e-75d2-4094-aa64-372b6ae565b4 · inbound

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism cites this paper.

VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism Number it: Temporal Grounding Videos like Flipping Manga

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:22.277066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:09:22.277066Z digest=sha256:12d3338e925bca7a889a2daa745b57db8049097e07e1f088d28edaa8f12b931f

Observation c79c5776-cd85-4233-a118-800e5393bb60 · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Number it: Temporal Grounding Videos like Flipping Manga

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.943418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.943418Z digest=sha256:7843dcdb65210518686d0f4dd93d72c50fd4f8b00e334dcad5496b82acc78bbb

Observation 64179649-6c66-45a4-8c0a-944396226658 · inbound

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation cites this paper.

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation Number it: Temporal Grounding Videos like Flipping Manga

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:43:49.391833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:10:44.608959Z digest=sha256:5d5c4065ab6ddf2144a4d8b8b743b909be9190d870a17e771a51ddea097b30e6

Observation 236e8bce-f0ed-463a-895b-a3947d3cf26b · inbound

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding cites this paper.

TimePLE: Rethinking Temporal Representation for Video Temporal Grounding Number it: Temporal Grounding Videos like Flipping Manga

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T23:32:54.533949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:32:54.533949Z digest=sha256:484cae1b5dc7cc43bbd4fc474da3e4dc7f75548f8592a48fd51a48e33571a776