Pith. sign in

Paper Citation Record · LEDGER

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2507.04976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04976 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:41:39.062407Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8dc3bbba-9988-4504-9af2-da643588d748 · outbound

This paper cites Figure 14:Prompt used for Evaluation:(a) Evaluation prompt for answerable dataset, and (b) Evaluation prompt for our unanswerable dataset.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Figure 14:Prompt used for Evaluation:(a) Evaluation prompt for answerable dataset, and (b) Evaluation prompt for our unanswerable dataset

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.468104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:38.870093Z digest=sha256:9d20c2acc5f3de9540ef4a8625139cc470f6d9b6ac7b53c5cf3355a8f8eb9c26

Observation 02f5da86-7927-41f2-a283-6efff1c09946 · outbound

This paper cites an unresolved cited work.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:41:41.346617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:37.065756Z digest=sha256:42aa507eb5aa7cd770ce27516d11c97e69db4eb080b4d72c4021b0a80c225e8f

Observation 57943758-8e5e-41ea-94d9-3e882fba7ed4 · outbound

This paper cites A.3 ETHICSSTATEMENT In our study, we utilize Large Language Models (LLM) to generate our UVQA dataset and evaluate video-LLMs, which may result in unintended outcomes.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models A.3 ETHICSSTATEMENT In our study, we utilize Large Language Models (LLM) to generate our UVQA dataset and evaluate video-LLMs, which may result in unintended outcomes

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.705794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:38.744180Z digest=sha256:1e477c30dd47d1a0424ad82f60721291adb898178195f21f08623935e2b435d8

Observation 85137c54-8704-444c-bb7d-3a88e310606d · outbound

This paper cites The Llama 3 Herd of Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.255667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.255667Z digest=sha256:56133a24b4e41eaf99cb2dade872af131be9b220f4829a1346d976ded3d8a205

Observation 73f65394-be23-4b3f-a8e5-b7a101e5b96d · outbound

This paper cites UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.335716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.335716Z digest=sha256:beb5713755cf4c1593b3479a9788d824f6205cfffec5d2fb0967b1a5d60b4f28

Observation 62189781-c344-4e24-b858-45b5902f6ba8 · outbound

This paper cites URL https://aclanthology.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models URL https://aclanthology

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:41.204900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:37.912431Z digest=sha256:0c5fde0892059867886ce7f2e7113f5e0e057930a182ddd6737cda4a64e6896b

Observation 00f23829-0ca4-4a92-8137-c43ac7228a64 · outbound

This paper cites Language Models Can See: Plugging Visual Controls in Text Generation.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Language Models Can See: Plugging Visual Controls in Text Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.021957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.021957Z digest=sha256:1d53b3b512170f89d73075f8ea7dd98f341d4078c08d72c9f00d0be143345d57

Observation 9118d1cc-2d16-46f2-b112-37ed7af3a42a · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.043123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.043123Z digest=sha256:77a7d3ad723f8a2ea79812c9bd569e7c7f2c7331de731f278ad64edb22e1b7cc

Observation 0347a6c8-223a-4acd-985c-78e758eeced3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.101205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.101205Z digest=sha256:f1b81efadadbd6e2b8d371d0a86300a1f306eb34a92841e6bf06dc62747334e6

Observation 0683f6b1-37ae-4a44-a710-5083f293a2a1 · outbound

This paper cites Know Your Limits: A Survey of Abstention in Large Language Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Know Your Limits: A Survey of Abstention in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.196093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.196093Z digest=sha256:29efc928224413ad7ba9d4c621d26b387b718750d0f341bb142e39552ea990fa

Observation 8881582b-3d4b-45c7-9a09-7e66de6adf72 · outbound

This paper cites Reliable visual question answering: Abstain rather than answer incorrectly.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Reliable visual question answering: Abstain rather than answer incorrectly

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:41.101947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:38.247478Z digest=sha256:c7171b94aef72548371bf545131ba239fde07dd4a6171a312768be44123f5812

Observation 94a70e9d-42ba-4c91-8346-677d493d6595 · outbound

This paper cites TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.377164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.377164Z digest=sha256:086f88e08b42d55b5358a0106a712bc48c4214a9c690eb12ccf0ddccd4ff8641

Observation 161abb5c-963a-480d-9f2d-f8cc0305dab7 · outbound

This paper cites doi: 10.18653/v1/2022.emnlp-main.280.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models doi: 10.18653/v1/2022.emnlp-main.280

Reference 19

Resolution
verified exact
doi, observed 2026-08-06T19:41:39.212989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:38.442554Z digest=sha256:27fc9458ce6bc66bbc145b388b1d073a9585578cbd2560e5e66ba538cff04c61

Observation 6eb2d644-0901-4600-9cfb-d1dd8dfff21a · outbound

This paper cites doi: 10.18653/v1/2023.findings-emnlp.797.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models doi: 10.18653/v1/2023.findings-emnlp.797

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.481770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.481770Z digest=sha256:35cf6fb541ae6f26ddd8cf5b5ee7716ff056f32b8ca0b543cf11f4e8de41cf40

Observation 80aecfca-840e-43ab-83a4-dc31abeba65e · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 22

Resolution
malformed identifier
no resolver link, observed 2026-08-06T19:41:38.650060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.650060Z digest=sha256:8633ce896bd0abe0f73233d2d20ed6482a1ed232777143beafae9d657a8e74cb

Observation f6e78513-1439-4059-a12c-58f917db02c8 · outbound

This paper cites an unresolved cited work.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:41:40.862940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:38.693004Z digest=sha256:2949c3103d90468c8c151d77ea6a3dbd96566532d731c155e7e07d5d2395fd3a

Observation c83c0aad-ef3c-41b9-846e-4d605e0a9699 · outbound

This paper cites (2023); Li et al.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models (2023); Li et al

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.579369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:38.838760Z digest=sha256:6bdca42315b34ce2efbd742212dba62d3e35b1b73561b33a10c24202c80c3bbd

Observation 062e422e-1d8d-48b5-9331-3c8c8a40b6d8 · outbound

This paper cites an unresolved cited work.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:41:40.309993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:38.914823Z digest=sha256:10d98b9e75bf04d318f2aa5a2b1a8bba74d1e4e847ee786126fb7b9553e67467

Observation 2b3a4891-e017-4a31-ac0b-74358f2b3dec · outbound

This paper cites If the question cannot be answered using the video content, state that it is unanswerable and provide a reason.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models If the question cannot be answered using the video content, state that it is unanswerable and provide a reason

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.088734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:39.013143Z digest=sha256:c03bcaf849b6afabf54a535fc3a95b180696a794c7e7999639aab0f9a0fbe390

Observation d5262a23-4d2f-48bc-a386-d802029116e7 · outbound

This paper cites 5All annotators have TOEFL iBT scores above 100 and hold at least a bachelor’s degree.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models 5All annotators have TOEFL iBT scores above 100 and hold at least a bachelor’s degree

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:39.948586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:39.062407Z digest=sha256:29d4a1acaa5c7777e1c054fe900db458cbc9c907ece98aa81275f21f0a09587f

Observation df799896-7081-41b3-a7e3-1a1f51c2dac9 · outbound

This paper cites Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A

Reference 2008

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T19:41:39.608405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:37.587634Z digest=sha256:076f69843b6a063de7268bcd2d61077e3ca91236b20a76144c2cc70cd14d22ad

Observation 6d3a0d12-eb22-46e0-990d-cc9176b83382 · outbound

This paper cites Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.666437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.666437Z digest=sha256:27099ba9d25d4c809531f4bfe8a79f811e75acdcb9430f2fad580486382c384a

Observation 75009a38-1cf7-40b2-9eeb-61476426c4a0 · outbound

This paper cites Alignment for Honesty.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Alignment for Honesty

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.287739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.287739Z digest=sha256:caf5830bb02549de453e3e4c56dbf4c4599f75502c7d2753ae8d29a003d7af4a

Observation 29e42f01-c0a9-4a27-b57b-40a538265d7d · outbound

This paper cites Mapping Images to Scene Graphs with Permutation-Invariant Structured Prediction.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Mapping Images to Scene Graphs with Permutation-Invariant Structured Prediction

Reference 2018

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:41:39.743993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:37.410814Z digest=sha256:dc2439c78623142622849f0d6564e748398a151b757919e838dd670236f1ad8d

Observation e4ea3966-e96b-4e0e-93ab-c2b5f76afa5a · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video understanding.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video understanding

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.964592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T19:41:38.533379Z digest=sha256:083e069f62ddc6c35089db255271c8b487ab5f77a23bfef72d3cfd9534705785

Observation 6d9a7dca-3d00-45b4-b331-d8a61213e178 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.146585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.146585Z digest=sha256:6353d534f01d58e7c7aed3f0de809a91a648a0906cefef7c6aa9bbb40f25aafc

Observation 0087c442-8ec5-4d59-879e-a0b2f32f4ba8 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.813687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.813687Z digest=sha256:2789ead71f05d94c571e678589060bf334b05eca5dacb6ab4a7c26edeaee05d3

Observation 08062c73-ad7b-42cc-8184-5b9fbc832a40 · outbound

This paper cites doi: 10.18653/v1/2023.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models doi: 10.18653/v1/2023

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.497820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.497820Z digest=sha256:b793b3e367158b5398a879dc8cbe6e386a43d7c3d0bbdcb2e5e48d8fcd2f1d50

Observation 9242563d-580b-4c2c-9616-b4cf89155b5e · outbound

This paper cites PaLM 2 Technical Report.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models PaLM 2 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.015458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.015458Z digest=sha256:e51f46458cfdcdbcf3765e3f838ee52d4b3dcb952a053f3935395bd23f2b851d

Pith citing papers

No inbound Pith citation observations are available.