Pith. sign in

Paper Citation Record · LEDGER

Object-centric Video Question Answering with Visual Grounding and Referring

As of 17 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2507.19599.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.19599 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:19:40.237501Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy44
  • unresolved35
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f4a2b7bb-2e0c-4a06-b9fb-9f8b22e30d21 · outbound

This paper cites GPT-4 Technical Report.

Object-centric Video Question Answering with Visual Grounding and Referring GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:30.864771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:30.864771Z digest=sha256:ae24a943b7bf9ce8e7af88b511969d00cddf8dcc3a287d6ba590177923b470d6

Observation 15a59f5c-3d8f-4f09-ab56-e0507cb122af · outbound

This paper cites Burst: A benchmark for unifying ob- ject recognition, segmentation and tracking in video.

Object-centric Video Question Answering with Visual Grounding and Referring Burst: A benchmark for unifying ob- ject recognition, segmentation and tracking in video

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:52.299469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:30.941003Z digest=sha256:ebbdc0f52a0e8a430961df734ee2053aa3de9784d7c83951965e33d298a4566b

Observation 72cd616b-060b-4990-93c0-ea6b3f62f79b · outbound

This paper cites Qwen2.5-VL Technical Report.

Object-centric Video Question Answering with Visual Grounding and Referring Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.060322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.060322Z digest=sha256:a1f64612dac1ef0a3975a62d5aff273b01b5b1b4103f683456ebdebae73c6d32

Observation 64102306-6b7f-4f91-80c3-f32aa1652cfa · outbound

This paper cites One token to seg them all: Lan- guage instructed reasoning segmentation in videos.

Object-centric Video Question Answering with Visual Grounding and Referring One token to seg them all: Lan- guage instructed reasoning segmentation in videos

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:52.160107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:31.203054Z digest=sha256:9246066b48e958043f144d896f63f52f1c0d17d75342254ecc95477d65426314

Observation de2d2fdc-fc69-4e84-a048-286ac09946bf · outbound

This paper cites Xmem++: Production-level video segmentation from few annotated frames.

Object-centric Video Question Answering with Visual Grounding and Referring Xmem++: Production-level video segmentation from few annotated frames

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:52.024268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:31.309544Z digest=sha256:d8b9a921c8bf9d64c9e5550668946043f3d02bd038e29d7d99ce9de957502c16

Observation 3897ae52-368e-4cf8-a23e-9833ac40268c · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

Object-centric Video Question Answering with Visual Grounding and Referring Coco-stuff: Thing and stuff classes in context

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.766255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:31.414465Z digest=sha256:77af7b8475734cf118efd3b3f2b180a740295414c3da4af4d361aa46f76d64b2

Observation 6fca08f5-55d2-42a4-8ad2-c38d9365bb64 · outbound

This paper cites Vip-llava: Making large multi- modal models understand arbitrary visual prompts.

Object-centric Video Question Answering with Visual Grounding and Referring Vip-llava: Making large multi- modal models understand arbitrary visual prompts

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.512367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:31.529633Z digest=sha256:21c07b2dbc85faceb5eac09fbbc0b564a82ada81a07dc8a9f752cef5231087cb

Observation 5d06ad8b-31d9-4502-86fc-bb077eb6f79d · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Object-centric Video Question Answering with Visual Grounding and Referring Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.659606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.659606Z digest=sha256:6f076934c9b1b9f9f3e5ac0964d27177c42a95f17388cff4ff41c5881beef3f9

Observation 1ee91e18-9a87-407e-9fd5-9213426db69f · outbound

This paper cites Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos.

Object-centric Video Question Answering with Visual Grounding and Referring Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.776850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.776850Z digest=sha256:6d7bb1113f48ccfc66cded77bda9e3917ef223b409c0f72a9485b62f13a7ae1b

Observation c8d0e2f8-5dd9-4fac-8c69-b65e8e9c55ea · outbound

This paper cites Detect what you can: Detecting and representing objects using holistic models and body parts.

Object-centric Video Question Answering with Visual Grounding and Referring Detect what you can: Detecting and representing objects using holistic models and body parts

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.309218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:31.883436Z digest=sha256:e477843e49e704f6d12fe7d2ccfc7a0b2a39b7fe7ee80fcea5ce3f94f64c6be5

Observation 87cfd5da-20c0-4dd2-be86-c82accb6bea4 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Object-centric Video Question Answering with Visual Grounding and Referring How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:31.985474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:31.985474Z digest=sha256:af7eb369660349f111dd8f64a5ed648a63fa9c15554469a16fbe9d83f9c51ee9

Observation 08867776-fa48-4a4a-afc0-7032580eddf9 · outbound

This paper cites Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model.

Object-centric Video Question Answering with Visual Grounding and Referring Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:51.085303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:32.101781Z digest=sha256:b9e9761ed9a05856986628210f1709bdfc1d666496273a9e4fcb62a67be5cb7a

Observation 467cceab-2790-4e27-8aab-b4c40e40fda3 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Object-centric Video Question Answering with Visual Grounding and Referring VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.183705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.183705Z digest=sha256:a9d2cb45617d40aa2a0cf58270ec51bbac702ae61814ee0809f34493bfffc236

Observation 6e4f8a82-6789-47d6-ab6f-b8fe37747739 · outbound

This paper cites Grounded question- answering in long egocentric videos.

Object-centric Video Question Answering with Visual Grounding and Referring Grounded question- answering in long egocentric videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:50.848944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:32.316400Z digest=sha256:57af02d191e58411b3e5a9e6eb731e8c4cc5905403630c94e69032762ecf012c

Observation b9e99d39-f6dc-4276-9c58-823895421b22 · outbound

This paper cites Mevis: A large-scale bench- mark for video segmentation with motion expressions.

Object-centric Video Question Answering with Visual Grounding and Referring Mevis: A large-scale bench- mark for video segmentation with motion expressions

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:50.598223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:32.471391Z digest=sha256:9818017efb4906cdc27c81f13132a228fcb2641245cedf652620512fa66e1197

Observation d3ab6386-277c-40ad-a059-8b69694a63cc · outbound

This paper cites Mose: A new dataset for video object segmentation in complex scenes.

Object-centric Video Question Answering with Visual Grounding and Referring Mose: A new dataset for video object segmentation in complex scenes

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:50.279506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:32.592931Z digest=sha256:337c14ba3aace326cea7b12bd72991004ef394b7e7691878ba521b79b6bdc60c

Observation a69ad915-4348-4f40-b026-8ab69739055a · outbound

This paper cites LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring LVOS: A Benchmark for Large-scale Long-term Video Object Segmentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.702316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.702316Z digest=sha256:5e6d9f8a761ebe36beb0581da2e3db238c389c0d863fe7c1efa3480b05332a29

Observation 615ca1a8-b839-481e-86be-2a949d326ed0 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Object-centric Video Question Answering with Visual Grounding and Referring LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.811828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.811828Z digest=sha256:02c07dd53a24988b2bfed99a82ccb43b042319626293e501c03f40d34329ae6c

Observation 2a6d0135-259a-4498-af7e-5c2be488fbd3 · outbound

This paper cites GPT-4o System Card.

Object-centric Video Question Answering with Visual Grounding and Referring GPT-4o System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:32.918892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:32.918892Z digest=sha256:145f7d2f2c23699efd8196031866bea961eac6f575a23f0e3d74ce5ab353c1f0

Observation a4d82382-ce65-4f2f-90f6-875055bd318e · outbound

This paper cites Cotracker3: Simpler and better point tracking by pseudo-labelling real videos.

Object-centric Video Question Answering with Visual Grounding and Referring Cotracker3: Simpler and better point tracking by pseudo-labelling real videos

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.949105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:33.036045Z digest=sha256:f8dee933bbb014c888eba3550a7d1e638dc837c0f5cfc10937905118284b8fa7

Observation 8e024cfc-dbb8-4abf-99fb-9249ed304db2 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Object-centric Video Question Answering with Visual Grounding and Referring Referitgame: Referring to objects in photographs of natural scenes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.720019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:33.140267Z digest=sha256:e93da48f5fe8a2e423a1bfe31cdcc4e43bacd1783a8f6bbe85ee6da649190263

Observation 29c9ab82-2837-4706-add4-46a147faccd3 · outbound

This paper cites Video object segmentation with language referring ex- pressions.

Object-centric Video Question Answering with Visual Grounding and Referring Video object segmentation with language referring ex- pressions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.479956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:33.238590Z digest=sha256:c54c9453f8ba1b28b4050e47985077b7fa6262bef65304ae5420e081b1d5ebbf

Observation a306dbb0-bdd4-4aa2-8e53-6b96dab7dc20 · outbound

This paper cites Segment anything.

Object-centric Video Question Answering with Visual Grounding and Referring Segment anything

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:49.210901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:33.380141Z digest=sha256:61437b357a6e50707f8b87f18dc15b59b5ce5858212cfe6a5ecf250bc73e8e74

Observation 79e37907-115d-4a89-8f14-eeb865c5259e · outbound

This paper cites Grounding language models to images for multi- modal inputs and outputs.

Object-centric Video Question Answering with Visual Grounding and Referring Grounding language models to images for multi- modal inputs and outputs

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.929539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:33.502374Z digest=sha256:3d69f7bd95063abda6a36f99644403dcd598fa163ca472845a8aa6c0f1956e30

Observation 58447ff6-88b7-4e0d-ab6b-135f1d9fc163 · outbound

This paper cites Generating images with multimodal language models.

Object-centric Video Question Answering with Visual Grounding and Referring Generating images with multimodal language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.636508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:33.605622Z digest=sha256:088a12b4aaf211ec7cbd675de2fda18ed42d8143cc8d0c14afb0f6ff22747f93

Observation d3f1ad95-2d6d-41cb-b41c-26b35e07aed9 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Object-centric Video Question Answering with Visual Grounding and Referring Lisa: Reasoning segmentation via large language model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.380741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:33.708532Z digest=sha256:5ad702d4b570db240c07a4d05598ca4b793e63a59796169e2972d25f114f8e4a

Observation f50cf220-2210-43f9-852c-2d9ad08ed121 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

Object-centric Video Question Answering with Visual Grounding and Referring MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:33.841194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:33.841194Z digest=sha256:233422484885affeb93410a6907791ca74fee5ce50706b736d9fd4f5ea3c54bb

Observation 95ea52ed-4eec-4bce-990b-893c9990c20e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.009655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.009655Z digest=sha256:e46c8fd5e6cb3aec3d0b04b9512ad9a3fad5fad430f7b417421c146012408eef

Observation f7cad20f-21af-4df1-92e8-4210147ffecf · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.124422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.124422Z digest=sha256:0c234f78ff9b0ef141e3e688a90281638973334e96d0ecc43f2fcfd9223ec2c3

Observation 172277c2-04cd-4578-9645-535d88e70850 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:48.097986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:34.242376Z digest=sha256:47b1ea57d9a036518c05a06aa27f43e6f321988f78ba8a728ed3517e6e60905c

Observation 3a772343-879c-4ae3-8d14-0f5d729134dc · outbound

This paper cites Towards Robust Referring Video Object Segmentation with Cyclic Relational Consensus.

Object-centric Video Question Answering with Visual Grounding and Referring Towards Robust Referring Video Object Segmentation with Cyclic Relational Consensus

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.362625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.362625Z digest=sha256:3d57bb8f81dcadac9469438edd53dcfeabe983df3f6f2aa05ee123f8d443eef8

Observation 0b453265-6047-4a43-867c-f4127e8ee3d2 · outbound

This paper cites Describe anything: Detailed localized image and video captioning.

Object-centric Video Question Answering with Visual Grounding and Referring Describe anything: Detailed localized image and video captioning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.846987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:34.543370Z digest=sha256:45d2d5ffe1134f0d15f0c9dd3887b3048a329556fb698d0fd67981c62f0cfdc0

Observation b81e087f-251f-40a7-ae09-ceb99b331e7a · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Object-centric Video Question Answering with Visual Grounding and Referring Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:34.701074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:34.701074Z digest=sha256:daa534f6c16cbecd365931bae9db7281060d0c3f7c5d89bdab15c0f8473628ef

Observation 1237dd00-ff9e-45c2-adab-a977ca078d9e · outbound

This paper cites Rouge: A package for automatic eval- uation of summaries.

Object-centric Video Question Answering with Visual Grounding and Referring Rouge: A package for automatic eval- uation of summaries

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.568836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:34.792454Z digest=sha256:3b9a4b51a100622d764417aa3ba88cda322c988f99b7bb0b14d87740c2c6cd0a

Observation a155d143-04a0-4678-ac25-5bd7cbdcefed · outbound

This paper cites Glus: Global-local reasoning unified into a single large language model for video segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring Glus: Global-local reasoning unified into a single large language model for video segmentation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.292410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:34.885264Z digest=sha256:3249578c0052fb7d7684ef9a4d0252a24570027b0972f4ab6f8ec1e81183cab3

Observation a14696b6-30a5-4cce-8ba6-08e350b0cc02 · outbound

This paper cites Visual instruction tuning.

Object-centric Video Question Answering with Visual Grounding and Referring Visual instruction tuning

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:47.019287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:34.986936Z digest=sha256:d063c79402a447b24d392cab814bc7a14a34e056b05ee39dbe1c14d2b23c2d07

Observation 47c75236-174c-43bb-82be-0cc419d7d744 · outbound

This paper cites Lamra: Large multimodal model as your ad- vanced retrieval assistant.

Object-centric Video Question Answering with Visual Grounding and Referring Lamra: Large multimodal model as your ad- vanced retrieval assistant

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.770684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:35.099805Z digest=sha256:35417553f8720e80d1ecb75249eabfb27682e99758d47e1d5346a5a612f7689c

Observation 2dc44f75-013a-403d-be4b-9560b923ed50 · outbound

This paper cites Decoupled Weight Decay Regularization.

Object-centric Video Question Answering with Visual Grounding and Referring Decoupled Weight Decay Regularization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:35.242260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:35.242260Z digest=sha256:e26eff827a03f5e665edc0a59e32af2b43cbc11e2a5b221d32c8e548b0bbf6f0

Observation aa765c0b-ad33-427e-93f2-d3503289bb97 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Object-centric Video Question Answering with Visual Grounding and Referring Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:35.344758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:35.344758Z digest=sha256:d17149da8a04d715273881321f2616c58d3dc4e560403ddff863c9392b643dce

Observation 6cbe70e5-59c8-4e31-98c6-eee78a3e78ec · outbound

This paper cites Generation and comprehension of unambiguous ob- ject descriptions.

Object-centric Video Question Answering with Visual Grounding and Referring Generation and comprehension of unambiguous ob- ject descriptions

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.495180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:35.462523Z digest=sha256:463876b46549a9a0d389d11275d10f4f0b94104d6f3f62156cf0d799dc50894f

Observation 1ea753bb-c75a-4a3c-af41-dcc1cd2756be · outbound

This paper cites Large-scale video panoptic segmentation in the wild: A benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Large-scale video panoptic segmentation in the wild: A benchmark

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.303905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:35.567110Z digest=sha256:9be56f4e59d2315b69708d3703e5372bb693ea3fa3c512c710cb2bd0c992d2c7

Observation 18aacfb3-e0e4-4215-bf55-5ba9bff3d63f · outbound

This paper cites V-net: Fully convolutional neural networks for volumetric medical image segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring V-net: Fully convolutional neural networks for volumetric medical image segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:46.043914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:35.730518Z digest=sha256:375a52d5c97f72b9868895db72f3e6c83f13151a289fb17bf96be7196ee252a5

Observation 0ed6dda5-284f-482d-af20-571d707c01f7 · outbound

This paper cites Bleu: A method for automatic eval- uation of machine translation.

Object-centric Video Question Answering with Visual Grounding and Referring Bleu: A method for automatic eval- uation of machine translation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.806001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:35.851977Z digest=sha256:14c36c225e3f058028a454903bb32a073a2f44013ae38952656c9ab2da01856b

Observation 4fc8125a-1305-44cd-931f-26d4ddd572c4 · outbound

This paper cites Perception test: A diagnostic bench- mark for multimodal video models.

Object-centric Video Question Answering with Visual Grounding and Referring Perception test: A diagnostic bench- mark for multimodal video models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.588351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:35.974409Z digest=sha256:397ced523942d72e9d4c9704a1244a87135c181ba40d2bf36628a0aaa0bcd257

Observation 07e3f2f3-ce0a-4fc7-b345-467de8bfbfab · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Object-centric Video Question Answering with Visual Grounding and Referring Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:36.065603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:36.065603Z digest=sha256:d492385a90d175d1f99003e42a661020ffb90c87b6945a8e7ac3016f0b96727a

Observation fa54ea9a-a496-4f7e-aace-b97d167440a3 · outbound

This paper cites Occluded video in- stance segmentation: A benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Occluded video in- stance segmentation: A benchmark

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.342672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:36.172270Z digest=sha256:36b346067a342a2ba23e3bd75acca471a957392d1e5df0c148cd3585ae84183e

Observation c5968d3d-6c7a-4dee-8860-bda3c817aa7e · outbound

This paper cites Artemis: Towards referential un- derstanding in complex videos.

Object-centric Video Question Answering with Visual Grounding and Referring Artemis: Towards referential un- derstanding in complex videos

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:45.082859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:36.327772Z digest=sha256:878fccf94eb8fe2562377f2f42a8770931753e925f880f72a6598dd575bb614b

Observation fc7d881a-c0f5-4e22-b453-0570dc9067a9 · outbound

This paper cites Paco: Parts and attributes of common objects.

Object-centric Video Question Answering with Visual Grounding and Referring Paco: Parts and attributes of common objects

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:44.828041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:36.435318Z digest=sha256:2cefefa48e70b05a794d7657ef475fc0671d00af145c9e00418bd7b0fd0f47fa

Observation ca092a3f-da64-4fc2-aee2-c549055601d9 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Object-centric Video Question Answering with Visual Grounding and Referring SAM 2: Segment Anything in Images and Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:36.552673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:36.552673Z digest=sha256:c0ae0a38810e6e13ebba76986c7fdd28af91735d0b7ae797a134baaf78d8f89d

Observation 5c03be5e-3b00-4980-88c4-78066345f1ff · outbound

This paper cites Hiera: A hierarchical vision transformer without the bells-and-whistles.

Object-centric Video Question Answering with Visual Grounding and Referring Hiera: A hierarchical vision transformer without the bells-and-whistles

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:44.575409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:36.644218Z digest=sha256:31a5af0dcc45e7615ee4adac30cd92f3b3668d3bd1ea12f77ff4a9fd2e27d01f

Observation 297c4461-aa4a-413f-8cf5-41104126d5c6 · outbound

This paper cites Urvos: Unified referring video object segmentation network with a large-scale benchmark.

Object-centric Video Question Answering with Visual Grounding and Referring Urvos: Unified referring video object segmentation network with a large-scale benchmark

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:44.344740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:36.749964Z digest=sha256:b0b1e9ca166e551bbc32c15a4c8eb5e20b48be1e359eb5ea810a7b0ca92f9e9d

Observation 624e5f94-9f0c-470f-acef-f2167f960c44 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Object-centric Video Question Answering with Visual Grounding and Referring Emu: Generative Pretraining in Multimodality

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:36.979072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:36.979072Z digest=sha256:8258bf038cd5bef5d1889297a4838b9a96a30281869e635456c43e77ac65bf05

Observation 1b464f74-7deb-44c6-8ad2-d543e84718e8 · outbound

This paper cites Cider: Consensus-based image descrip- tion evaluation.

Object-centric Video Question Answering with Visual Grounding and Referring Cider: Consensus-based image descrip- tion evaluation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.856383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:37.123600Z digest=sha256:ed6d093d46b504290be9dbe6dde745dbdf99781523276e85600b88ca45c87f23

Observation bc9be667-cf9b-423a-94e6-5f3465b07219 · outbound

This paper cites Ov-vis: Open-vocabulary video instance segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring Ov-vis: Open-vocabulary video instance segmentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.599672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:37.236690Z digest=sha256:5f18483732fc565579bd41bf20c2d97c9709a720775ffa01913d96f834b92620

Observation a4c34e4a-92bc-43f1-8a2c-09516a9bef1c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Object-centric Video Question Answering with Visual Grounding and Referring Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:37.356424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:37.356424Z digest=sha256:ffbf5e84e999db0da3d7ed9654c6663b3f72e3bb75ce66140be15a8229aa60a0

Observation ff284c6b-ee5d-44f0-b956-a21b58c65482 · outbound

This paper cites Unidentified video objects: A benchmark for dense, open-world segmentation.

Object-centric Video Question Answering with Visual Grounding and Referring Unidentified video objects: A benchmark for dense, open-world segmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.351712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:37.497702Z digest=sha256:3b18066d48b8ba4eb7e291240e706912ada84996f927081068ab758e8d59380b

Observation 6a9aaa68-a589-4f06-9676-baa83d4ef10e · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Object-centric Video Question Answering with Visual Grounding and Referring InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:37.661534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:37.661534Z digest=sha256:6550f65f735241114f52600e4cc4a65d6eedd6904851146ef6b444799ad42578

Observation 177d8f52-5c82-4d46-86ca-5d8315692384 · outbound

This paper cites Internvideo2: Scaling foun- dation models for multimodal video understanding.

Object-centric Video Question Answering with Visual Grounding and Referring Internvideo2: Scaling foun- dation models for multimodal video understanding

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:43.127670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:37.795167Z digest=sha256:021e6eb9dbccd4942f3de4da83e89aca3a9404a07682acf71b728ec648bda82c

Observation ff67cfcd-c3e3-4eb0-bb9c-114846a0b438 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

Object-centric Video Question Answering with Visual Grounding and Referring Next-qa: Next phase of question-answering to explaining temporal actions

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.919131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:37.926811Z digest=sha256:61b3dca284d0f66380e816c5898b9aae5339f6ee2309d2060a814c376b32e96e

Observation 872524ac-7cfc-420a-a0ea-c0d3541c37e5 · outbound

This paper cites Visa: Reasoning video object segmen- tation via large language models.

Object-centric Video Question Answering with Visual Grounding and Referring Visa: Reasoning video object segmen- tation via large language models

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.729703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:38.024130Z digest=sha256:310447369f2cd2ea05cd99666113a692669d03ea110345f45a2c845d8e5ec9ba

Observation 455e99f2-4c90-42d3-be2f-6e87c41904f0 · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Object-centric Video Question Answering with Visual Grounding and Referring VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.109831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.109831Z digest=sha256:cf32ba331a20c709cda14dbda4944467188d53924efe998179509b8eea3b4f3c

Observation 44a2257f-8dea-4cd1-8131-6d0e80f5ee68 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Object-centric Video Question Answering with Visual Grounding and Referring Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.229151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.229151Z digest=sha256:b1abcf69100eb5e50079a1e3ca05fc21f9784c905b9c64902eb2bb3498b7036b

Observation 827c5a9c-da2c-4dd7-80b6-f910464b5d2e · outbound

This paper cites LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model.

Object-centric Video Question Answering with Visual Grounding and Referring LISA++: An Improved Baseline for Reasoning Segmentation with Large Language Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.352762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.352762Z digest=sha256:7e41aa17d943959dda058b30d4dc89f5ad00b9cbb7b2cc29f880f85b3fbc7ec6

Observation cb2f858e-5f53-4f4b-b6dc-fac94a69787c · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Object-centric Video Question Answering with Visual Grounding and Referring mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.498533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.498533Z digest=sha256:a67e0cede97817d9488075d1946af6d34182fab689cb1b090fd80f1c40b226cb

Observation dd125f56-671f-4185-abee-9e1900f90ef0 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Object-centric Video Question Answering with Visual Grounding and Referring Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.603030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.603030Z digest=sha256:4db8f62a53caccc2e8cd7e5b1ed0592d1fb9de0d3771b14fe94b28cfe95ff1ff

Observation 450421f4-78e9-495d-b65b-092441bda764 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

Object-centric Video Question Answering with Visual Grounding and Referring Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:38.745855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:38.745855Z digest=sha256:ab15f46672ad7a95f4c3703063b432d7ef7e0a0e886224f2f7c1c76a4548626a

Observation 972d9ea1-6637-4f61-8713-1686d8528fd1 · outbound

This paper cites Sa2va: Marrying sam2 with llava for dense grounded understanding of im- ages and videos.

Object-centric Video Question Answering with Visual Grounding and Referring Sa2va: Marrying sam2 with llava for dense grounded understanding of im- ages and videos

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.467851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:38.813448Z digest=sha256:e91335a46aa296f297729da6d8128ddd4c3650aca6ea6673ad16497b0d236243

Observation a545401d-1949-49a9-8bc6-1d7869e20b88 · outbound

This paper cites Os- prey: Pixel understanding with visual instruction tun- ing.

Object-centric Video Question Answering with Visual Grounding and Referring Os- prey: Pixel understanding with visual instruction tun- ing

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.256487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:38.918347Z digest=sha256:bcec6eec3a11663c04d51cc590927700e430f7623ed75ea2aa9767fade411c32

Observation ccf18085-1bfe-4023-b3b1-542b44531262 · outbound

This paper cites Videorefer suite: Advancing spatial-temporal object understanding with video llm.

Object-centric Video Question Answering with Visual Grounding and Referring Videorefer suite: Advancing spatial-temporal object understanding with video llm

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:42.021520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:39.024526Z digest=sha256:f2fbe597ac632264cc23d05925a6c8c3b357d1217c2943939e74bd34791ce11c

Observation 025328a4-de54-48dd-bb8d-b7f8679de765 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Object-centric Video Question Answering with Visual Grounding and Referring VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.140269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.140269Z digest=sha256:4cb2114c8e688f63fcbed98a2437e3def23ead5f84850216786ec7bbc7566e4e

Observation b7f6f91c-70a2-46e0-adda-b72c587b9531 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Object-centric Video Question Answering with Visual Grounding and Referring Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.258349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.258349Z digest=sha256:5a3f3b634e0de876c43e73a4ed32c8769223457482e8b624a309b55426a93973

Observation 99f4e4f6-d6ea-447e-b0c2-7f198124978f · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Object-centric Video Question Answering with Visual Grounding and Referring LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.360265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.360265Z digest=sha256:90e592ea63a52d635a19fc15accf7ec8cebffd57971055f7de81100c1ba42d11

Observation b5ee245e-d564-4974-a412-7de86e962e88 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Object-centric Video Question Answering with Visual Grounding and Referring GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.474732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.474732Z digest=sha256:4f0b6910732127993235a94a8e92d8818575def8a77724797fa423471f01cf6c

Observation 67cb2719-8b81-48f7-adeb-ca73373abd9c · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Object-centric Video Question Answering with Visual Grounding and Referring LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.606964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.606964Z digest=sha256:b96ecbdb9a6694e16234ce5d4251f6072d3f0ac7816a5bbd9ea19d26b6e448aa

Observation 0d770155-febd-4a35-a8c6-8d9ae90f717e · outbound

This paper cites Scene parsing through ade20k dataset.

Object-centric Video Question Answering with Visual Grounding and Referring Scene parsing through ade20k dataset

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:41.775584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:39.713362Z digest=sha256:e5223c07bd9fba3b240d8b022a0bbdfb962b3183409c02e48c05b4e3d55b4f57

Observation f836cefc-0b8a-427a-9a92-f9483ff2bffc · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Object-centric Video Question Answering with Visual Grounding and Referring MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:19:39.834178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:19:39.834178Z digest=sha256:982e877922dfc91ad0878c65fc097cefc4afaa1c44ddea8963f4ccfa54816e17

Observation b64863f9-c6e2-441b-8d25-9f17e8265ac1 · outbound

This paper cites The complete list of used datasets in training is presented in Tab.

Object-centric Video Question Answering with Visual Grounding and Referring The complete list of used datasets in training is presented in Tab

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:19:41.512014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:39.976573Z digest=sha256:585994d12053c919a70fb76cafb64bb36a460137f5a63564624afd7eefdc2150

Observation 129b29ba-459e-434b-8336-2757de218c6a · outbound

This paper cites an unresolved cited work.

Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:19:41.268991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:40.126377Z digest=sha256:d4b8cfb2b5c605f2dc79e9170ef16a8873254a11a8b9f3d63c87e6ef8332c091

Observation b5b5aa3e-f768-44d4-b4c0-9b6077d8db28 · outbound

This paper cites an unresolved cited work.

Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:19:40.965085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:40.237501Z digest=sha256:d06bc7fbe202e33052550c90258e45ceadfa15ccfac371244e30787d83700f7a

Observation c21f2be1-ac58-4b45-9031-bf31ee57b6f1 · outbound

This paper cites an unresolved cited work.

Object-centric Video Question Answering with Visual Grounding and Referring Unresolved cited work

Reference 223

Resolution
parse uncertain
raw_fallback, observed 2026-08-06T14:19:44.107775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T14:19:36.874703Z digest=sha256:79c320add97357b2610840106d48add56d79ca9fcb7ca83f0b26221320abde35

Pith citing papers

No inbound Pith citation observations are available.