Pith. sign in

Paper Citation Record · LEDGER

BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2309.15785.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.15785 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:53:01.291124Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T20:21:57.995580Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2c18829c-482a-4a61-8c12-3e5b88644105 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:57.998742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:e45e3f8d92df64394a8eb54c330ec175756abc3fbb5566250face8e0d43ed0c6

Observation a3b9a45d-6725-4bc9-b4b6-83eb4af1d8a4 · inbound

Generative Timelines for Instructed Visual Assembly cites this paper.

Generative Timelines for Instructed Visual Assembly BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:47:39.823182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:47:39.823182Z digest=sha256:04c5fc4c3abce740a4f431cd53a34070becb5cb24038042a580d473fcee91989

Observation 60a8e0cb-1022-4047-a7ed-825647eab5e4 · inbound

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding cites this paper.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.214641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.214641Z digest=sha256:614242244ad8c33b93adf73dae07f5e123c908d2e3b8af5d288376261fc82ec7

Observation 974cd515-e231-4f49-a557-bf9411acb935 · inbound

ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models cites this paper.

ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T05:29:55.438142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:29:55.438142Z digest=sha256:bfcab0382ac56ec20956c92376208c90390cef192d841ed3b559c12cdf72c03d

Observation b91f4669-2ba2-43a9-b417-d6bd7b745f1e · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 151

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.476595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.476595Z digest=sha256:4258594cca310dddc3676f0bc419c13cabd2b9fa9c14690b8113decfce8d7b93

Observation c6196b34-0323-4b0c-a59a-c4796f0303b6 · inbound

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness cites this paper.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.087918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.087918Z digest=sha256:63e803586a0c0c92d1eb1c3b74bf10ae3aeb75dc83cf8f36909e0cecbf29ae42

Observation 4e1ec515-1f6b-4fdd-98bb-f55350a443f3 · inbound

ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations cites this paper.

ResNetVLLM-2: Addressing ResNetVLLM's Multi-Modal Hallucinations BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:52:16.814156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:52:16.814156Z digest=sha256:ceae3bfb706ce50e3b135640446cc2a40a8128b431736f9c7ae78e2f2fd9d3a2

Observation 0ba76b65-e53d-49a9-8a97-f5deaed3158c · inbound

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task cites this paper.

ResNetVLLM -- Multi-modal Vision LLM for the Video Understanding Task BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T11:53:01.291124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:53:01.291124Z digest=sha256:f4e939d5891dcd68b6a997a3ddba4ce9cf100fc02412f0c6725c51ed5531f99c

Observation 610e1dfd-bbd6-458f-b3f3-251567ad856b · inbound

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model cites this paper.

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:26:37.966867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:26:37.966867Z digest=sha256:78056a3e95846a2203388cb4633185aaa3cf7f20faa0c41b8c0405d4c6dcbaf5

Observation 4b5a927f-3563-448b-b4f1-cbf3e29cf05d · inbound

Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback cites this paper.

Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:43:26.262167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:43:26.262167Z digest=sha256:ea3c736b9b221a5a86792a0f47a084653781268e1e4c1a7e8114dc1a27ee5570