Pith. sign in

Paper Citation Record · LEDGER

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models

As of 13 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2507.04976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04976 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:41:39.062407Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8dc3bbba-9988-4504-9af2-da643588d748 · outbound

This paper cites Figure 14:Prompt used for Evaluation:(a) Evaluation prompt for answerable dataset, and (b) Evaluation prompt for our unanswerable dataset.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Figure 14:Prompt used for Evaluation:(a) Evaluation prompt for answerable dataset, and (b) Evaluation prompt for our unanswerable dataset

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.468104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:38.870093Z digest=sha256:5499caf6e231f2208c5209b2475af13388d1b9b4f7f688cd73b46f14636ad1b6

Observation 02f5da86-7927-41f2-a283-6efff1c09946 · outbound

This paper cites an unresolved cited work.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:41:41.346617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:37.065756Z digest=sha256:1a35d7ddf9ff7defc8b7bda09405940f0248c3b88d20009f40a03dd8f84b6859

Observation 57943758-8e5e-41ea-94d9-3e882fba7ed4 · outbound

This paper cites A.3 ETHICSSTATEMENT In our study, we utilize Large Language Models (LLM) to generate our UVQA dataset and evaluate video-LLMs, which may result in unintended outcomes.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models A.3 ETHICSSTATEMENT In our study, we utilize Large Language Models (LLM) to generate our UVQA dataset and evaluate video-LLMs, which may result in unintended outcomes

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.705794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:38.744180Z digest=sha256:11a917ba8f6de6f702cc0aad155b5453c2e5c11600fce770243310ff0ff4719f

Observation 85137c54-8704-444c-bb7d-3a88e310606d · outbound

This paper cites The Llama 3 Herd of Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.255667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.255667Z digest=sha256:6313fc2d3a949fcc6eea9c3a6689a3d32de9becf0d697e25031cccc271c72d50

Observation 73f65394-be23-4b3f-a8e5-b7a101e5b96d · outbound

This paper cites UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models UNK-VQA: A Dataset and a Probe into the Abstention Ability of Multi-modal Large Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.335716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.335716Z digest=sha256:33936efbee8bd0fe4d6e7d03ec821a9945d2cc20256957e648815e880786d8dc

Observation 62189781-c344-4e24-b858-45b5902f6ba8 · outbound

This paper cites URL https://aclanthology.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models URL https://aclanthology

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:41.204900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:37.912431Z digest=sha256:b35e77c014ad5b8e24948e571ba87620ce9c1d442b802657ea8283868eaef274

Observation 00f23829-0ca4-4a92-8137-c43ac7228a64 · outbound

This paper cites Language Models Can See: Plugging Visual Controls in Text Generation.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Language Models Can See: Plugging Visual Controls in Text Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.021957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.021957Z digest=sha256:f3e520717486e4ad8b8036dee53178030733f035e6583c182057098d744c1066

Observation 9118d1cc-2d16-46f2-b112-37ed7af3a42a · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.043123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.043123Z digest=sha256:c493dcc895c014b23c1c74d0d1da0bc7ad602e7ec01fca455f434bd07a2dd080

Observation 0347a6c8-223a-4acd-985c-78e758eeced3 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.101205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.101205Z digest=sha256:0f808b1de350f51e0d3b3e0cbc4b0158b5a2762e5baf1349aaafcfa3129123ee

Observation 0683f6b1-37ae-4a44-a710-5083f293a2a1 · outbound

This paper cites Know Your Limits: A Survey of Abstention in Large Language Models.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Know Your Limits: A Survey of Abstention in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.196093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.196093Z digest=sha256:9ba6930d63ab5ef718a2e3fdd2708bac7d1376701c5eb5ac91bd11a8588d0838

Observation 8881582b-3d4b-45c7-9a09-7e66de6adf72 · outbound

This paper cites Reliable visual question answering: Abstain rather than answer incorrectly.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Reliable visual question answering: Abstain rather than answer incorrectly

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:41.101947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:38.247478Z digest=sha256:61054f100697072ace52ca00e426a9b6752cd554031a945e3f6e7608f63af790

Observation 94a70e9d-42ba-4c91-8346-677d493d6595 · outbound

This paper cites TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models TLCR: Token-level continuous reward for fine-grained reinforcement learning from human feedback

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.377164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.377164Z digest=sha256:bbe1655abdb01572619885f6a1aa941f8024ac25fb4b8980b2245e5047b3f34d

Observation 161abb5c-963a-480d-9f2d-f8cc0305dab7 · outbound

This paper cites doi: 10.18653/v1/2022.emnlp-main.280.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models doi: 10.18653/v1/2022.emnlp-main.280

Reference 19

Resolution
verified exact
doi, observed 2026-08-06T19:41:39.212989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:38.442554Z digest=sha256:d2435a17ad7e53fa7dcac6a4a695bfb0e914abb67c70f222647fafb45e91f6a6

Observation 6eb2d644-0901-4600-9cfb-d1dd8dfff21a · outbound

This paper cites doi: 10.18653/v1/2023.findings-emnlp.797.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models doi: 10.18653/v1/2023.findings-emnlp.797

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.481770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.481770Z digest=sha256:fe76536d666eeb80519c75de25c23739829f96bf7f6fb9c4378db8cad5da680b

Observation 80aecfca-840e-43ab-83a4-dc31abeba65e · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 22

Resolution
malformed identifier
no resolver link, observed 2026-08-06T19:41:38.650060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.650060Z digest=sha256:c70ae4352ab1b748d37e2387b6d0cfd266382ae2ce33adc097e6c8fd5dbfed29

Observation f6e78513-1439-4059-a12c-58f917db02c8 · outbound

This paper cites an unresolved cited work.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:41:40.862940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:38.693004Z digest=sha256:ec9dff170d49e05f4a4ad4ce06661672c4a8ed28ceb9346037dbeeb7a6476573

Observation c83c0aad-ef3c-41b9-846e-4d605e0a9699 · outbound

This paper cites (2023); Li et al.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models (2023); Li et al

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.579369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:38.838760Z digest=sha256:f614fb571839d70250ad9442f137b3b9811902d61d6440a1e6c567d956a6d52c

Observation 062e422e-1d8d-48b5-9331-3c8c8a40b6d8 · outbound

This paper cites an unresolved cited work.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:41:40.309993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:38.914823Z digest=sha256:e9aa14dd17d9d00e828791669c6aeb416110a18606fd6b716fa14ecc9e8db947

Observation 2b3a4891-e017-4a31-ac0b-74358f2b3dec · outbound

This paper cites If the question cannot be answered using the video content, state that it is unanswerable and provide a reason.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models If the question cannot be answered using the video content, state that it is unanswerable and provide a reason

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.088734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:39.013143Z digest=sha256:dab953c7ff8cb9cda954a411c5e0fa5fcd98460b2ff1b8edc09c5710d62d0cbb

Observation d5262a23-4d2f-48bc-a386-d802029116e7 · outbound

This paper cites 5All annotators have TOEFL iBT scores above 100 and hold at least a bachelor’s degree.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models 5All annotators have TOEFL iBT scores above 100 and hold at least a bachelor’s degree

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:39.948586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:39.062407Z digest=sha256:650ea301a08df892bbf6bec910eca70a9ca630352c4ead0970ae4c0a22d8af27

Observation df799896-7081-41b3-a7e3-1a1f51c2dac9 · outbound

This paper cites Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David A

Reference 2008

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T19:41:39.608405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:37.587634Z digest=sha256:1166eedcc2d88d4d0454106b8548f4afdbd4843c90bd94de1c387dca9f582b49

Observation 6d3a0d12-eb22-46e0-990d-cc9176b83382 · outbound

This paper cites Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Sicong Leng, Hang Zhang, Guanzheng Chen, Xin Li, Shijian Lu, Chunyan Miao, and Lidong Bing

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.666437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.666437Z digest=sha256:3d9d4d998ae77e88beeab5518a74a01c44ad7c786b84ce3f0111d5021c92a794

Observation 75009a38-1cf7-40b2-9eeb-61476426c4a0 · outbound

This paper cites Alignment for Honesty.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Alignment for Honesty

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:38.287739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:38.287739Z digest=sha256:f1b4368f230bc6269f7009c23103f878b9cccedb6ea6e5056d2d6f9cd3e8255c

Observation 29e42f01-c0a9-4a27-b57b-40a538265d7d · outbound

This paper cites Mapping Images to Scene Graphs with Permutation-Invariant Structured Prediction.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Mapping Images to Scene Graphs with Permutation-Invariant Structured Prediction

Reference 2018

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:41:39.743993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:37.410814Z digest=sha256:adace60a310eceab57ac94e86ce55daedd7a1c7a7732aa491e92cb488705803e

Observation e4ea3966-e96b-4e0e-93ab-c2b5f76afa5a · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video understanding.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models Video-LLaMA: An instruction-tuned audio-visual language model for video understanding

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:41:40.964592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T19:41:38.533379Z digest=sha256:ad598e86a7eaca82fda80d618ecdb234831827107cee5de9688f6102f8a9012f

Observation 6d9a7dca-3d00-45b4-b331-d8a61213e178 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.146585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.146585Z digest=sha256:a1da7633d93c4328c1924f7bdb1022c1f7ca96a40f979e3da189f18e3bc473be

Observation 0087c442-8ec5-4d59-879e-a0b2f32f4ba8 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.813687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.813687Z digest=sha256:6c8c1254553bcce28691fa3ec033820b71fb9f2795f52577d88ac0d8958da456

Observation 08062c73-ad7b-42cc-8184-5b9fbc832a40 · outbound

This paper cites doi: 10.18653/v1/2023.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models doi: 10.18653/v1/2023

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.497820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.497820Z digest=sha256:bb7b744955a138a0076bc77d1c02eb895090077989d58bc8e18c5f39eb9cae1f

Observation 9242563d-580b-4c2c-9616-b4cf89155b5e · outbound

This paper cites PaLM 2 Technical Report.

Can Video LLMs Refuse to Answer? Alignment for Answerability in Video Large Language Models PaLM 2 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T19:41:37.015458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:41:37.015458Z digest=sha256:805952a7f22903904c79c52db8e8c3571fc047b0d07d35678f3e40a13c106aa6

Pith citing papers

No inbound Pith citation observations are available.