Pith. sign in

Paper Citation Record · LEDGER

BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2309.15785.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.15785 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T17:47:39.823182Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T20:21:57.995580Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2c18829c-482a-4a61-8c12-3e5b88644105 · inbound

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning cites this paper.

PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:21:57.998742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T20:21:57.873354Z digest=sha256:88cdee9c69d5d44655e709bd7f689e1faf1533a35c010ee0fddeef9aed4a5df0

Observation a3b9a45d-6725-4bc9-b4b6-83eb4af1d8a4 · inbound

Generative Timelines for Instructed Visual Assembly cites this paper.

Generative Timelines for Instructed Visual Assembly BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:47:39.823182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:47:39.823182Z digest=sha256:93396a38b47f170d8b3b69e900c9e254f5c979ecb5e617d26e49ebe6f32138a4

Observation 60a8e0cb-1022-4047-a7ed-825647eab5e4 · inbound

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding cites this paper.

DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T17:43:03.214641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:43:03.214641Z digest=sha256:549dcbe93542c949734cb65b5291c736c56c773859a84524440b537038b0474f

Observation 974cd515-e231-4f49-a557-bf9411acb935 · inbound

ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models cites this paper.

ATP-LLaVA: Adaptive Token Pruning for Large Vision Language Models BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T05:29:55.438142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:29:55.438142Z digest=sha256:c4e54264b68c893d6164f3e0c0c4cc3481daccc17fe3293012f9868a56c31a21

Observation b91f4669-2ba2-43a9-b417-d6bd7b745f1e · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 151

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.476595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.476595Z digest=sha256:35c800066067897d5687f3152d2b6c48526daebdeeef5f8eb930cbc326279555

Observation c6196b34-0323-4b0c-a59a-c4796f0303b6 · inbound

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness cites this paper.

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:33:54.087918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:33:54.087918Z digest=sha256:3474cff3e3d9a681817aef4564313ee00cb25281a18d87cd9444069624e60ff2

Observation 610e1dfd-bbd6-458f-b3f3-251567ad856b · inbound

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model cites this paper.

Period-LLM: Extending the Periodic Capability of Multimodal Large Language Model BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:26:37.966867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:26:37.966867Z digest=sha256:4f454e5886e17ce37e9af12f1923fb355c3e5a93d707321cb07abcba1c9cc556

Observation 4b5a927f-3563-448b-b4f1-cbf3e29cf05d · inbound

Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback cites this paper.

Oracle-RLAIF: An Improved Fine-Tuning Framework for Multi-modal Video Models using Reinforcement Learning from Ranking Feedback BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T12:43:26.262167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:43:26.262167Z digest=sha256:28f7ad0b9cc73c1c36335bcc6d712a7d2f27ac1932a67d94d6ed8e476c1fcfff