Pith. sign in

Paper Citation Record · LEDGER

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens

As of 15 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2412.09919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09919 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:38:23.397829Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved19
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e269592-11af-445e-aa6e-a0cf055d1045 · outbound

This paper cites Accessed: 2024-09-30.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Accessed: 2024-09-30

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.534530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.085961Z digest=sha256:ed3e2259ee5d33e5c90a3efedf5dab0f866b68161a69fcbc4758f9274dde6087

Observation 5cce60a5-a4a9-4ed6-90ef-b513cafce434 · outbound

This paper cites Flamingo: A Visual Language Model For Few-shot Learning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Flamingo: A Visual Language Model For Few-shot Learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.520243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.090969Z digest=sha256:614a2c3fa6f51f6c5a37cff9fa15d6650faeb642b51547b35dd9f826907ceca4

Observation 5b1a2745-d8f5-40b8-8072-e37ceeccd44d · outbound

This paper cites Deepspeed-Inference: Enabling Efficient Inference of Trans- former Models at Unprecedented Scale.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Deepspeed-Inference: Enabling Efficient Inference of Trans- former Models at Unprecedented Scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.505394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.095827Z digest=sha256:cd0f77e31aa2ce32ab7ff47b62b64fa03feddd8f2aa352c39fb9187949da20c2

Observation 0ed46894-0977-4b78-b1a0-fcc333a4411a · outbound

This paper cites Qwen Technical Report.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.100815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.100815Z digest=sha256:d23626af3c7e5c20c9c23c0dd15b1c5001694ec91ab991002ccdb7bc47def38d

Observation 24e6e467-6e9d-48d9-84df-282eb0896330 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.489843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.106163Z digest=sha256:e597deb243244e1e7983be011fda3f6dc175ea58f5afcdd1f02c2c4e1ebf2949

Observation 66863b66-26b1-4de9-999c-5059196ed2d0 · outbound

This paper cites Token Merging for Fast Stable Diffusion.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Token Merging for Fast Stable Diffusion

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.472817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.111065Z digest=sha256:2956db4f995ae4edcb3f995db310887622722ce334d89aa19ce5e500ba6f49ad

Observation 165db5d1-38ec-480f-bb5e-65fb6fdd29d3 · outbound

This paper cites Token Merging: Your ViT But Faster.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Token Merging: Your ViT But Faster

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.456825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.116301Z digest=sha256:c5867149f0c760a2cf105af73da1130c0890b1984c10d3dbf954a212f39eddc9

Observation fb4e4d0b-64f3-44c5-bbd6-066c0d878c26 · outbound

This paper cites Dif- fusiondet: Diffusion model for object detection.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Dif- fusiondet: Diffusion model for object detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.440804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.120851Z digest=sha256:328db0576567ee43694df13c04dd81763072d3ab11e15c6f145b0894881f4ded

Observation 97cfd238-8e33-4990-bfe2-38f7dded1efd · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.125642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.125642Z digest=sha256:b04c9a3bf292db74397d7db88782895703d8c686c9b9dc3bd4e96c9232e34ef8

Observation 4a20ac0e-7c12-4615-828d-320b0218b5bb · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.130731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.130731Z digest=sha256:84f58517f3e22707718cd9ff72806cdd6e1c25d0ef09302f2e89fdb62d02db6f

Observation 7c5aca54-ace4-42a8-8118-b1de65e80b0e · outbound

This paper cites The Llama 3 Herd of Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.135791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.135791Z digest=sha256:719811ec4654205617de88dff96109c8b2a1616d259d62e91046f521b0fb126c

Observation 2330e1df-0570-4baf-9e50-13fe6758d348 · outbound

This paper cites ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.422363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.141476Z digest=sha256:f3ceb68b3f831ad8aa0715ecfa5b0ff7bcf782c7342544e668255199e1425f1e

Observation 6c8bf51d-4b97-4f92-8a27-f03d3374c844 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.146380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.146380Z digest=sha256:4190722b17aec66de51569239aa2ffe553140409517239a19f01ec023d3c147a

Observation 8781f6cb-a499-4491-b7e3-42c24537ec2a · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.407378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.151391Z digest=sha256:9d2a035804268cdae7a0a27a1690bb4d48ab537cd928505d5d75d71e3f83de76

Observation 3f706293-107d-4993-bbd4-e83e824ee93d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.155758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.155758Z digest=sha256:c8338df6ad7c453a4f3792f6ac1e4b049353f1182047f441a89def4dfb1dc502

Observation e8f1a4dc-1609-4fb8-9d2b-03606732a319 · outbound

This paper cites Making the V in VQA Matter: Ele- vating the Eole of Image Understanding in Visual Question Answering.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Making the V in VQA Matter: Ele- vating the Eole of Image Understanding in Visual Question Answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.392962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.160786Z digest=sha256:1fbb5aed29a79daf911bcf94693a326245272de1cba874dddfa793d58212a13b

Observation f1228d8d-5a4c-4164-9e72-a3f607853442 · outbound

This paper cites Vizwiz Grand Challenge: Answering Visual Questions from Blind People.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Vizwiz Grand Challenge: Answering Visual Questions from Blind People

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.378034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.165606Z digest=sha256:6899b69fa0f1d19e26d32189aafe5186604c142606eddc0a7c984a8650c32139

Observation 6a00835a-f980-4756-a9df-70f5f278827b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.170289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.170289Z digest=sha256:2e4f1f0bed899af50ba1e70e702dca6b1f46c1e3608d3c22229fcb94ae6555f2

Observation 858c5d64-c8a9-495f-be26-980af22e3251 · outbound

This paper cites Vision-based freezing of gait detection with anatomic patch based representation.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Vision-based freezing of gait detection with anatomic patch based representation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.363992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.175362Z digest=sha256:71038b40cc20393b67d19bec7de6f802753255b52090dd70b694ab65f53b72b5

Observation 5b56b643-4a71-4cfc-9c8a-74f0e0f86360 · outbound

This paper cites GQA: A New Dataset for Real-World Visual Eeasoning and Compositional Question Answering.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens GQA: A New Dataset for Real-World Visual Eeasoning and Compositional Question Answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.349286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.179941Z digest=sha256:48ea3d13bb2845cddee3c2603d4472beef68e80597ea74554192b22d1b40f8a6

Observation 7794ffa5-db67-469b-b98c-b57fb734b362 · outbound

This paper cites LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.184668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.184668Z digest=sha256:2a8fa8748cffa97e4179c8703b098f82d06c98aba86348f0e0bf658acb94cf63

Observation d3431e70-d808-426f-8882-589bf490cab8 · outbound

This paper cites Chat-univi: Unified Visual Representation Em- powers Large Language Models with Image and Video Un- derstanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Chat-univi: Unified Visual Representation Em- powers Large Language Models with Image and Video Un- derstanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.334087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.189447Z digest=sha256:d1c0519e57924729d8607b32edfa296b44f662391657f7e24df0f7b14b3446a9

Observation e95e7ffa-4f76-4850-ac58-5a023780b521 · outbound

This paper cites ReferItGame: Referring to objects in pho- tographs of natural scenes.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens ReferItGame: Referring to objects in pho- tographs of natural scenes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.319453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.193513Z digest=sha256:fe069fd7c088026fcb60c5ed9426a7b18feaf98f21b7c1fd147d85c39b582dfe

Observation 050cfcae-719a-4d2b-aa14-f9261a73546f · outbound

This paper cites Shamma, Michael S.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Shamma, Michael S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.303994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.197464Z digest=sha256:61033b9501e2e0f9505ce328cc8fab509305d91e922f78598f27856f7cd854d4

Observation 68eec932-30cf-47ef-bfd9-ffb94fc0e74a · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.201279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.201279Z digest=sha256:4a788af2edb440f3da73f9ddff1f16f3fb429486d43985f233d5a6ba59d9c3d5

Observation 988303e2-7d51-4828-bbcf-77e5927944d9 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.289534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.205522Z digest=sha256:cdf4ccc831eb6e62e6f074242b6d0117e6a9662ba1eb89bfc8eae21478db30f3

Observation 06f1b6c1-cfe4-4fcd-b12b-9167853b30d2 · outbound

This paper cites MVBench: A Comprehensive Multi-Modal Video Under- standing Benchmark.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MVBench: A Comprehensive Multi-Modal Video Under- standing Benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.274403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.209832Z digest=sha256:7c6a9fec3c4c5da1a6afbba1535c8638ecbf1de6e403c0fcf1eb3ebdeb2a5a01

Observation 518762f7-23b3-4235-92d4-4c1344c7aecb · outbound

This paper cites VidToMe: Video Token Merging for Zero-Shot Video Edit- ing.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VidToMe: Video Token Merging for Zero-Shot Video Edit- ing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.258141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.213617Z digest=sha256:dac2fd14765da61dc05c74c756052c9389a3cd3b2915e062403dd4560e0f5e25

Observation 51dd0415-e64e-4783-9a9e-fda1710687c5 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Evaluating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.242917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.217336Z digest=sha256:063cec8e8ba76ffa8c2dd4c70ef9c784df9fcffc4fb934a26d3d3d9c13fd66bd

Observation 92a16398-b545-44e4-8979-57ae413ca2be · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.227794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.222025Z digest=sha256:6f42cc430b1d63fa9384e423f4300618679661e18bf0014204b7e26d119b9063

Observation 2e3f976d-c981-44d5-a858-17164d855aee · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.226456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.226456Z digest=sha256:26abeb13524fafa091911825d2a2edb37d8c959c45dfcefcb68e9eb5e66cb36d

Observation 3f5c0778-e792-4e59-8a65-efc0971f955e · outbound

This paper cites VILA: On Pre-training for Vi- sual Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VILA: On Pre-training for Vi- sual Language Models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.212782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.232050Z digest=sha256:aab0380606a88fc2580fdc16df09fa52438e27c388409462819dfb2197c5671c

Observation 32ebfc29-6c55-4996-aae1-a76ca8583fbc · outbound

This paper cites Visual Instruction Tuning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Visual Instruction Tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.199799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.237633Z digest=sha256:05318799f8b5aaf8ae82f838a1367423edf2ac28ba16cce2e9f3449cc9d5f2ec

Observation 07d0d442-44ae-4980-a08c-383de4b60c55 · outbound

This paper cites MMBench: Is Your Multi-Modal Model an All-Around Player? In ECCV, pages 216–233.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MMBench: Is Your Multi-Modal Model an All-Around Player? In ECCV, pages 216–233

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.186301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.242415Z digest=sha256:4f2dc16a33f4989e0116bd509a10bcbfea3dcc77bc1b61fb68ac59633a9a3b82

Observation 7a9a20a6-08bb-4b17-9cb3-637fd6a280f8 · outbound

This paper cites Decoupled Weight Decay Regularization.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Decoupled Weight Decay Regularization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.247628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.247628Z digest=sha256:f44fe89f23efc46264d254a5d09e6d2ef862c0a283e73e78ddc28ed92aa03e6d

Observation 7e65f279-9a63-403a-824c-3ac0e044ba28 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.171875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.252816Z digest=sha256:6e98cc12947b0a620f810d1472285ca15f9ae063d2714505a1e90e8282c0c943

Observation 548e78ee-8e54-4b6a-8702-cc6a77bc6639 · outbound

This paper cites Autoregressive omni-aware outpainting for open- vocabulary 360-degree image generation.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Autoregressive omni-aware outpainting for open- vocabulary 360-degree image generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.156331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.257615Z digest=sha256:b7bb00b57f8f4117ba69ece731f782d4acf9f40990ff3c14f8eb847e9146a14d

Observation e4ff90b6-9572-4e9e-bb04-3ebbee2a1dcc · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.262723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.262723Z digest=sha256:01cb544864f638cc28c988a83b0530396b082aa3c6cf42477ae407d56f18696a

Observation f691b681-a29b-4662-b668-5adb143c233a · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.141744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.267826Z digest=sha256:cb9d427eb846c07db0aac2dfae2ecd23890c9d4d43821cd45131ac46caafc265

Observation 1380e099-43ab-480b-8e23-5aa7f9838e04 · outbound

This paper cites Egoschema: A Diagnostic Benchmark for Bery Long-Form Video Language Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Egoschema: A Diagnostic Benchmark for Bery Long-Form Video Language Understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.126565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.273657Z digest=sha256:1379ae8cde82c4177f2358f5facaf673531569069e5bf94ce5675e050570ece6

Observation ddd3961b-6767-4666-8164-aac1912eaec5 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions, 2016.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Generation and comprehension of unambiguous object descriptions, 2016

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.111537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.278963Z digest=sha256:c29414fbb1ecf1722501545b2f7af4f24184d3bbeedda1d0ad9fadc2e67a4e87

Observation 185ae4d6-ef84-462f-a676-5b1346e15f11 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Ocr-vqa: Visual question answering by reading text in images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.096536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.284075Z digest=sha256:652fc34cd43e7e01fcee7006defe9f11a5cdd85c634e4b19604f29534c1a02b0

Observation 0ae5b099-655e-46a9-b95f-99c036a5e7cf · outbound

This paper cites Perception Test: A Diagnostic Benchmark for Multimodal Video Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.082201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.289095Z digest=sha256:7a6f79986bace1ef966b07e074283caba06945d230a98e369cc20f17d2906f8e

Observation f9c5f32d-f17c-4c41-a43f-f69bcfcbb382 · outbound

This paper cites Learning Transferable Visual Models from Natural Language Supervi- sion.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Learning Transferable Visual Models from Natural Language Supervi- sion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.065917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.294097Z digest=sha256:a5cfe219e9760a317f97568cbdc512eaa8ff55b2f16fc4aeff5058f305bf49f5

Observation f605098e-75ae-43b5-9fdf-74ae1e73dc5b · outbound

This paper cites A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge, 2022.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge, 2022

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.050404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.298810Z digest=sha256:9c01d8dc7cfe6e58fe45d600b6648af5af5375dfe55777c1b5af2bfe8a1723fc

Observation f567e276-c483-4e2b-ac6b-118259650460 · outbound

This paper cites Llava-prumerge: Adaptive Token Reduction for Efficient Large Multimodal Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Llava-prumerge: Adaptive Token Reduction for Efficient Large Multimodal Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.303588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.303588Z digest=sha256:4dfb52e74cae291738ed8a803b780f8cd9617865fdd53d701e360b98968ab28e

Observation b534e526-2aa0-4464-8f86-c197afa5c564 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.034898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.308314Z digest=sha256:7c583b3590085ff490469341f222511134539671caa3a7f51a21d09eca2a9e02

Observation 0203478c-b5fc-41e7-a016-a25cda20fc82 · outbound

This paper cites TextCaps: a Dataset for Image Captioning with Reading Comprehension, 2020.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens TextCaps: a Dataset for Image Captioning with Reading Comprehension, 2020

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.020566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.312965Z digest=sha256:1c126a4282f8cc909c7a4765ebbc107dfc613e995d40b3544275c1ae16530434

Observation be74fcc7-6dc7-4e10-a015-78bf9100fdce · outbound

This paper cites Towards VQA Models That Can Read.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Towards VQA Models That Can Read

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.006513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.317610Z digest=sha256:4eab942fbdc04e8e4090a15bffca7ce08ccd1c5b06a8795b22ffb2fee0d3edeb

Observation 4c438198-f11a-4cd1-a214-35405fb7ca7a · outbound

This paper cites Moviechat: From Dense To- ken to Sparse Memory for Long Video Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Moviechat: From Dense To- ken to Sparse Memory for Long Video Understanding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.992233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.322443Z digest=sha256:c0f6ed23ace9465718f5c16657dd754a7641146cd3559b92846b7f915b887bfe

Observation 2d20aec9-ae63-4483-a3d9-132da3fc07ce · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.326515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.326515Z digest=sha256:895590c6bd25974f6409dc80d4148c5120cc117d84085a3d2f723f5e2d8b57b9

Observation 927d1f5a-eb39-4d25-a0d6-5dd9d7e7c996 · outbound

This paper cites [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.330686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.330686Z digest=sha256:5a96a0d9f02c90df37b192d6a0de9e88defec9eaeac6b682b3803316a4ff4367

Observation a58606c6-b27b-4f42-bc8c-f950dcfea81b · outbound

This paper cites VisionLLM: Large Language Model is Also An Open-ended Decoder for Vision-Centric Tasks.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VisionLLM: Large Language Model is Also An Open-ended Decoder for Vision-Centric Tasks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.975446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.334738Z digest=sha256:7296f36aaf65a74d92e7375994ac46e47660e8607a8e07828eb9b6b7519e0837

Observation 40d9be43-8b8c-4ab3-b5f0-119fd11e7d43 · outbound

This paper cites LongVLM: Efficient Long Video Under- standing Via Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LongVLM: Efficient Long Video Under- standing Via Large Language Models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.959253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.338794Z digest=sha256:bf730f9c73741e435101ca2430daeacd7f2b12b2fb5d95b8a8febc180aff5323

Observation e7dd3e38-dc97-4301-ab75-fa784b31980b · outbound

This paper cites Video Question Answer- ing via Gradually Refined Attention over Appearance and Motion.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video Question Answer- ing via Gradually Refined Attention over Appearance and Motion

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.928832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.346774Z digest=sha256:eb27585283333cd70f73d854f86460440bdde2ea23d0e190782d13f660e841b3

Observation ad4deac9-1809-4d49-8119-ad02323728de · outbound

This paper cites Qwen2 Technical Report.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Qwen2 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.351205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.351205Z digest=sha256:7ad8c4582644e2f8ff587a537556fd156f0c699d8013fe87780dd707e10c817c

Observation 1440bf94-3c60-4b0a-a6ba-06e2acd78b34 · outbound

This paper cites SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Segmentation.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Segmentation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.355997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.355997Z digest=sha256:a170bcd216d3fdb4aa166fc25689722998f4ecb8e64f1f208e8307c5beb05bcd

Observation 8bf437b6-282a-4b7d-9fd5-9bfd9b010e8e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.360911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.360911Z digest=sha256:1699bec1698a26704aa91d243e14d395229fdf9a40b5207a536b7ce43b56772c

Observation de2b71cb-7949-42ce-9489-ce89004ba5e7 · outbound

This paper cites LLaV A-NeXT: A Strong Zero-shot Video Understanding Model, 2024.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLaV A-NeXT: A Strong Zero-shot Video Understanding Model, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.914436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.365795Z digest=sha256:6325fca5f3a0e446204fed77a8ee63fa545a642a3c9b63de0b312b4f3a4c0da1

Observation 2e7fbe65-0514-494c-b5a9-b5f9956e60c7 · outbound

This paper cites Needle In A Video Haystack: A Scalable Syn- thetic Framework for Benchmarking Video MLLMs.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Needle In A Video Haystack: A Scalable Syn- thetic Framework for Benchmarking Video MLLMs

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.897803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.370070Z digest=sha256:0269efb9560bde0a69d5b2b41a2527ec92bcddeff38f87a3de2ddd87ec0c04dc

Observation 1d57da18-8074-49fb-bc71-4fbfdcd2012d · outbound

This paper cites Clip in medical imaging: A survey.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Clip in medical imaging: A survey

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.880456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.374598Z digest=sha256:c5e1a51c301c03045625aefad611e1dcd90be0db154c6500e0b7127e7a696980

Observation 4e9f2669-99f0-4716-809f-fdfd699558a7 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.379009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.379009Z digest=sha256:36257905eae62217b063103568c0a228ed8a742fe9962f61b076014bee19d33d

Observation 4f8d223f-0972-4550-9f92-406505997ab6 · outbound

This paper cites A closer look at the cls token for cross-domain few-shot learning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens A closer look at the cls token for cross-domain few-shot learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.864704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.383842Z digest=sha256:650f0f09c64b477e1ed0639615985d39b7be5f28386876496bd89ece6618baf3

Observation 1a170a45-5858-417e-ad21-d9e763a9e8ce · outbound

This paper cites Training Details We adopt a two-stage training strategy [9, 30, 33], dividing training into pretraining for modality alignment and fine- tuning for instruction tuning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Training Details We adopt a two-stage training strategy [9, 30, 33], dividing training into pretraining for modality alignment and fine- tuning for instruction tuning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.849232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.388255Z digest=sha256:27baf47cab073c0779f119ccdd7d2420af8bec260496e0b74ef3e31f76408eae

Observation be9a685a-eca2-48a4-a5dc-bca19ff2bb72 · outbound

This paper cites Additional Discussion on Different Frame Se- lection Features.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Additional Discussion on Different Frame Se- lection Features

Reference 66

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T16:38:23.834550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.392641Z digest=sha256:6210b158c8cdb66ae47826084daf218a68c2f3052f247b3f88b8551a36a4f482

Observation f95bd0b0-8aba-4b16-b736-e932672b0924 · outbound

This paper cites Game Science.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Game Science

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.817966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.397829Z digest=sha256:663baa7bb705275ecdc6278f8112db393ee1ec1fcaa01e6281738417d4ed3431

Observation 3039831d-a05b-48b2-935a-8e8423126b96 · outbound

This paper cites an unresolved cited work.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Unresolved cited work

Reference 470

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:38:23.943843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:38:23.343023Z digest=sha256:609f6027597492efe76ecfcd88f7b2964eda699ed3d99322b8f27779bd9f88b2

Pith citing papers

No inbound Pith citation observations are available.