Pith. sign in

Paper Citation Record · LEDGER

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens

As of 15 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2412.09919.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09919 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:38:23.397829Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved19
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6e269592-11af-445e-aa6e-a0cf055d1045 · outbound

This paper cites Accessed: 2024-09-30.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Accessed: 2024-09-30

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.534530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.085961Z digest=sha256:692b82a5ac912482f734d8ed681cb271881414dfe6a95cdb032ac700e20d000e

Observation 5cce60a5-a4a9-4ed6-90ef-b513cafce434 · outbound

This paper cites Flamingo: A Visual Language Model For Few-shot Learning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Flamingo: A Visual Language Model For Few-shot Learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.520243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.090969Z digest=sha256:24a85200694ae780c46350ab2436a2fa35b4218747e0f33bd4e123f5d5459676

Observation 5b1a2745-d8f5-40b8-8072-e37ceeccd44d · outbound

This paper cites Deepspeed-Inference: Enabling Efficient Inference of Trans- former Models at Unprecedented Scale.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Deepspeed-Inference: Enabling Efficient Inference of Trans- former Models at Unprecedented Scale

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.505394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.095827Z digest=sha256:89957767f651240dd7e217c3d4eff82516ea60f8fd0b95e94ee3dcc24e797769

Observation 0ed46894-0977-4b78-b1a0-fcc333a4411a · outbound

This paper cites Qwen Technical Report.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Qwen Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.100815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.100815Z digest=sha256:d23626af3c7e5c20c9c23c0dd15b1c5001694ec91ab991002ccdb7bc47def38d

Observation 24e6e467-6e9d-48d9-84df-282eb0896330 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.489843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.106163Z digest=sha256:e1578a00a130db3b882e5916165b118c03a096024057eae4044e2f570529003e

Observation 66863b66-26b1-4de9-999c-5059196ed2d0 · outbound

This paper cites Token Merging for Fast Stable Diffusion.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Token Merging for Fast Stable Diffusion

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.472817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.111065Z digest=sha256:1931be55501c467bbabe063426426cf04e5aa7518bff8c96cd0232044e839982

Observation 165db5d1-38ec-480f-bb5e-65fb6fdd29d3 · outbound

This paper cites Token Merging: Your ViT But Faster.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Token Merging: Your ViT But Faster

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.456825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.116301Z digest=sha256:9a29b82e5c91009ad3d3333ded29cf229bae3378448f2d3bd01c2a08c4dfaf96

Observation fb4e4d0b-64f3-44c5-bbd6-066c0d878c26 · outbound

This paper cites Dif- fusiondet: Diffusion model for object detection.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Dif- fusiondet: Diffusion model for object detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.440804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.120851Z digest=sha256:e5f8f5317da529cb575891252dbe41004b62388cff87ddfb6cac6f38b85ee10d

Observation 97cfd238-8e33-4990-bfe2-38f7dded1efd · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.125642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.125642Z digest=sha256:b04c9a3bf292db74397d7db88782895703d8c686c9b9dc3bd4e96c9232e34ef8

Observation 4a20ac0e-7c12-4615-828d-320b0218b5bb · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.130731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.130731Z digest=sha256:84f58517f3e22707718cd9ff72806cdd6e1c25d0ef09302f2e89fdb62d02db6f

Observation 7c5aca54-ace4-42a8-8118-b1de65e80b0e · outbound

This paper cites The Llama 3 Herd of Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens The Llama 3 Herd of Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.135791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.135791Z digest=sha256:719811ec4654205617de88dff96109c8b2a1616d259d62e91046f521b0fb126c

Observation 2330e1df-0570-4baf-9e50-13fe6758d348 · outbound

This paper cites ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.422363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.141476Z digest=sha256:daff2eb93005c4c17ce08abe28a3e8aeda76e85e09c69d94c46669b60cad4d8b

Observation 6c8bf51d-4b97-4f92-8a27-f03d3374c844 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.146380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.146380Z digest=sha256:4190722b17aec66de51569239aa2ffe553140409517239a19f01ec023d3c147a

Observation 8781f6cb-a499-4491-b7e3-42c24537ec2a · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.407378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.151391Z digest=sha256:26af3f65b6ad08ba0f3c697a8b14f88fee720096edf5c2575ddedbaf5d94d832

Observation 3f706293-107d-4993-bbd4-e83e824ee93d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.155758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.155758Z digest=sha256:c8338df6ad7c453a4f3792f6ac1e4b049353f1182047f441a89def4dfb1dc502

Observation e8f1a4dc-1609-4fb8-9d2b-03606732a319 · outbound

This paper cites Making the V in VQA Matter: Ele- vating the Eole of Image Understanding in Visual Question Answering.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Making the V in VQA Matter: Ele- vating the Eole of Image Understanding in Visual Question Answering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.392962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.160786Z digest=sha256:e3c8b603eca42ced5ab81dc95a4b907a38763bd6e9d500317b99c58880c25faa

Observation f1228d8d-5a4c-4164-9e72-a3f607853442 · outbound

This paper cites Vizwiz Grand Challenge: Answering Visual Questions from Blind People.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Vizwiz Grand Challenge: Answering Visual Questions from Blind People

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.378034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.165606Z digest=sha256:1c99a4b939669cab65ceb0ccea9b155de6e19a3bccaa0916646218b630d7e182

Observation 6a00835a-f980-4756-a9df-70f5f278827b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LoRA: Low-Rank Adaptation of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.170289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.170289Z digest=sha256:2e4f1f0bed899af50ba1e70e702dca6b1f46c1e3608d3c22229fcb94ae6555f2

Observation 858c5d64-c8a9-495f-be26-980af22e3251 · outbound

This paper cites Vision-based freezing of gait detection with anatomic patch based representation.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Vision-based freezing of gait detection with anatomic patch based representation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.363992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.175362Z digest=sha256:9f9428f2cf15397a615ebded91f2be58b4aa7b8a85ee1054f9bc6243ecc9b4a5

Observation 5b56b643-4a71-4cfc-9c8a-74f0e0f86360 · outbound

This paper cites GQA: A New Dataset for Real-World Visual Eeasoning and Compositional Question Answering.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens GQA: A New Dataset for Real-World Visual Eeasoning and Compositional Question Answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.349286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.179941Z digest=sha256:193d9579361ed4e6f5186f59b3a964275756a5a3654193b7e7e1807965773382

Observation 7794ffa5-db67-469b-b98c-b57fb734b362 · outbound

This paper cites LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLMLingua: Compressing Prompts for Accelerated Inference of Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.184668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.184668Z digest=sha256:2a8fa8748cffa97e4179c8703b098f82d06c98aba86348f0e0bf658acb94cf63

Observation d3431e70-d808-426f-8882-589bf490cab8 · outbound

This paper cites Chat-univi: Unified Visual Representation Em- powers Large Language Models with Image and Video Un- derstanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Chat-univi: Unified Visual Representation Em- powers Large Language Models with Image and Video Un- derstanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.334087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.189447Z digest=sha256:017f1fea90a8b51d5a7224b25094b11426250b9760858611640e2c6c921c931e

Observation e95e7ffa-4f76-4850-ac58-5a023780b521 · outbound

This paper cites ReferItGame: Referring to objects in pho- tographs of natural scenes.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens ReferItGame: Referring to objects in pho- tographs of natural scenes

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.319453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.193513Z digest=sha256:e7ec82fac36bded56e601d31839a0725f03f84008a9d76f53fb0666e69f81750

Observation 050cfcae-719a-4d2b-aa14-f9261a73546f · outbound

This paper cites Shamma, Michael S.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Shamma, Michael S

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.303994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.197464Z digest=sha256:1b38db347196f57d43b7c4bef0ee0f775e66f6c31b295067c54cb77d567d1da0

Observation 68eec932-30cf-47ef-bfd9-ffb94fc0e74a · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.201279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.201279Z digest=sha256:4a788af2edb440f3da73f9ddff1f16f3fb429486d43985f233d5a6ba59d9c3d5

Observation 988303e2-7d51-4828-bbcf-77e5927944d9 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.289534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.205522Z digest=sha256:bea19cfa59caa80318b993ba23b51e2adc181c842d642c3759e374a161f6f58c

Observation 06f1b6c1-cfe4-4fcd-b12b-9167853b30d2 · outbound

This paper cites MVBench: A Comprehensive Multi-Modal Video Under- standing Benchmark.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MVBench: A Comprehensive Multi-Modal Video Under- standing Benchmark

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.274403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.209832Z digest=sha256:8d08b1401fed0c4ee03cc0dd69d6c4f3577986ff4ad930fc1d4241215cfaab54

Observation 518762f7-23b3-4235-92d4-4c1344c7aecb · outbound

This paper cites VidToMe: Video Token Merging for Zero-Shot Video Edit- ing.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VidToMe: Video Token Merging for Zero-Shot Video Edit- ing

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.258141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.213617Z digest=sha256:51dd2ba4d7fb775f805dd96d661fd7dfa3555b085558f672f0daa897da1bf9da

Observation 51dd0415-e64e-4783-9a9e-fda1710687c5 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Evaluating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.242917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.217336Z digest=sha256:cdb24a93e59daedac7bfda98b40f66731b275e25fc42fe5520d4d8985fcf1005

Observation 92a16398-b545-44e4-8979-57ae413ca2be · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.227794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.222025Z digest=sha256:f8dde140f9d7df7ab861b78537feb066872d174eb03a25b9dc1ad13e3d7a9afe

Observation 2e3f976d-c981-44d5-a858-17164d855aee · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.226456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.226456Z digest=sha256:26abeb13524fafa091911825d2a2edb37d8c959c45dfcefcb68e9eb5e66cb36d

Observation 3f5c0778-e792-4e59-8a65-efc0971f955e · outbound

This paper cites VILA: On Pre-training for Vi- sual Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VILA: On Pre-training for Vi- sual Language Models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.212782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.232050Z digest=sha256:4adeed76b0d2682c826faa24bfc4099fe0654d31cd23a6b6371f14e975610161

Observation 32ebfc29-6c55-4996-aae1-a76ca8583fbc · outbound

This paper cites Visual Instruction Tuning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Visual Instruction Tuning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.199799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.237633Z digest=sha256:343f2df92baed1051c745eed7a94002b6562d7f14e8fce7736eab0b095393e9d

Observation 07d0d442-44ae-4980-a08c-383de4b60c55 · outbound

This paper cites MMBench: Is Your Multi-Modal Model an All-Around Player? In ECCV, pages 216–233.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens MMBench: Is Your Multi-Modal Model an All-Around Player? In ECCV, pages 216–233

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.186301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.242415Z digest=sha256:7bffdc0a401403850d710f1e6e8a3b620b891dda70fc387714456207a9c397b4

Observation 7a9a20a6-08bb-4b17-9cb3-637fd6a280f8 · outbound

This paper cites Decoupled Weight Decay Regularization.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Decoupled Weight Decay Regularization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.247628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.247628Z digest=sha256:f44fe89f23efc46264d254a5d09e6d2ef862c0a283e73e78ddc28ed92aa03e6d

Observation 7e65f279-9a63-403a-824c-3ac0e044ba28 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.171875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.252816Z digest=sha256:dd5e09afec729b0691e5832bd357a2a9b5aebe0e8ab2984e3da5c3c650ccb739

Observation 548e78ee-8e54-4b6a-8702-cc6a77bc6639 · outbound

This paper cites Autoregressive omni-aware outpainting for open- vocabulary 360-degree image generation.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Autoregressive omni-aware outpainting for open- vocabulary 360-degree image generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.156331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.257615Z digest=sha256:37704a2dca5fa0825f690996bc2c905be44e297babdc87c7cc1025cd2d6a7458

Observation e4ff90b6-9572-4e9e-bb04-3ebbee2a1dcc · outbound

This paper cites Valley: Video Assistant with Large Language model Enhanced abilitY.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Valley: Video Assistant with Large Language model Enhanced abilitY

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.262723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.262723Z digest=sha256:01cb544864f638cc28c988a83b0530396b082aa3c6cf42477ae407d56f18696a

Observation f691b681-a29b-4662-b668-5adb143c233a · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Mod- els

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.141744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.267826Z digest=sha256:e4c95bf573e35bf703c6bb30b5299f11e464f4500740dd73b652e5c3ca70b6a4

Observation 1380e099-43ab-480b-8e23-5aa7f9838e04 · outbound

This paper cites Egoschema: A Diagnostic Benchmark for Bery Long-Form Video Language Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Egoschema: A Diagnostic Benchmark for Bery Long-Form Video Language Understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.126565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.273657Z digest=sha256:57714b26d0d6d6c50e9ece1d85f88ecee6dedcbdcbde694da0a5a9abbfe15cb3

Observation ddd3961b-6767-4666-8164-aac1912eaec5 · outbound

This paper cites Generation and comprehension of unambiguous object descriptions, 2016.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Generation and comprehension of unambiguous object descriptions, 2016

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.111537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.278963Z digest=sha256:48ce7514f470932d62359ca12ab4ca631fb1ce211b5a8fedb6c765c4741f9d11

Observation 185ae4d6-ef84-462f-a676-5b1346e15f11 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Ocr-vqa: Visual question answering by reading text in images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.096536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.284075Z digest=sha256:f34d7c0a51c06a84dc243b42d6ca9e220800be0ab220f60086d335ecbe3371be

Observation 0ae5b099-655e-46a9-b95f-99c036a5e7cf · outbound

This paper cites Perception Test: A Diagnostic Benchmark for Multimodal Video Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.082201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.289095Z digest=sha256:8a36a91da68986abc02b4345719082262cec8db471bd820906610856a165962c

Observation f9c5f32d-f17c-4c41-a43f-f69bcfcbb382 · outbound

This paper cites Learning Transferable Visual Models from Natural Language Supervi- sion.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Learning Transferable Visual Models from Natural Language Supervi- sion

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.065917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.294097Z digest=sha256:d82be0dea6b396dd04af10b28982a1e64484972ebce04adf6cd18d40e0645b8b

Observation f605098e-75ae-43b5-9fdf-74ae1e73dc5b · outbound

This paper cites A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge, 2022.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge, 2022

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.050404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.298810Z digest=sha256:883702f65c4eeef3cb246344f30853dd2da4f59aecb4d3fc022f82c65f626cd6

Observation f567e276-c483-4e2b-ac6b-118259650460 · outbound

This paper cites Llava-prumerge: Adaptive Token Reduction for Efficient Large Multimodal Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Llava-prumerge: Adaptive Token Reduction for Efficient Large Multimodal Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.303588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.303588Z digest=sha256:4dfb52e74cae291738ed8a803b780f8cd9617865fdd53d701e360b98968ab28e

Observation b534e526-2aa0-4464-8f86-c197afa5c564 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Conceptual captions: A cleaned, hypernymed, im- age alt-text dataset for automatic image captioning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.034898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.308314Z digest=sha256:e2824e395e65a9dfb73c5c348d3083f2326a4745fa53bd5d8b45b86e62513369

Observation 0203478c-b5fc-41e7-a016-a25cda20fc82 · outbound

This paper cites TextCaps: a Dataset for Image Captioning with Reading Comprehension, 2020.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens TextCaps: a Dataset for Image Captioning with Reading Comprehension, 2020

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.020566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.312965Z digest=sha256:c6f264a8c6c04f68c538e24ff54bd4d9c974abde79b565afb6fe2ca0f2a21226

Observation be74fcc7-6dc7-4e10-a015-78bf9100fdce · outbound

This paper cites Towards VQA Models That Can Read.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Towards VQA Models That Can Read

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:24.006513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.317610Z digest=sha256:81b12e7339d2d1ab31dd47ebc07ecf4b7fedfa60cf2913bf1037340f6c32fbfd

Observation 4c438198-f11a-4cd1-a214-35405fb7ca7a · outbound

This paper cites Moviechat: From Dense To- ken to Sparse Memory for Long Video Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Moviechat: From Dense To- ken to Sparse Memory for Long Video Understanding

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.992233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.322443Z digest=sha256:869db886295230af01b4e5a681382191a3da9184993356ad04c857dbd4e2f676

Observation 2d20aec9-ae63-4483-a3d9-132da3fc07ce · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.326515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.326515Z digest=sha256:895590c6bd25974f6409dc80d4148c5120cc117d84085a3d2f723f5e2d8b57b9

Observation 927d1f5a-eb39-4d25-a0d6-5dd9d7e7c996 · outbound

This paper cites [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens [CLS] Token Tells Everything Needed for Training-free Efficient MLLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.330686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.330686Z digest=sha256:5a96a0d9f02c90df37b192d6a0de9e88defec9eaeac6b682b3803316a4ff4367

Observation a58606c6-b27b-4f42-bc8c-f950dcfea81b · outbound

This paper cites VisionLLM: Large Language Model is Also An Open-ended Decoder for Vision-Centric Tasks.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens VisionLLM: Large Language Model is Also An Open-ended Decoder for Vision-Centric Tasks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.975446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.334738Z digest=sha256:4ff30146d9a6dab1b12b94b7a7df2e50fa35a99f515c3330fbaf3ab55991fa7e

Observation 40d9be43-8b8c-4ab3-b5f0-119fd11e7d43 · outbound

This paper cites LongVLM: Efficient Long Video Under- standing Via Large Language Models.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LongVLM: Efficient Long Video Under- standing Via Large Language Models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.959253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.338794Z digest=sha256:628b1f82903ff55f36cef6d18917fd6d8515f09b778e658924c617190b411903

Observation e7dd3e38-dc97-4301-ab75-fa784b31980b · outbound

This paper cites Video Question Answer- ing via Gradually Refined Attention over Appearance and Motion.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video Question Answer- ing via Gradually Refined Attention over Appearance and Motion

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.928832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.346774Z digest=sha256:21b8e371a40a4f3f2e54eeb70ea80ceca51a6c2aa6df0275d9e0bb3150abc394

Observation ad4deac9-1809-4d49-8119-ad02323728de · outbound

This paper cites Qwen2 Technical Report.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Qwen2 Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.351205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.351205Z digest=sha256:7ad8c4582644e2f8ff587a537556fd156f0c699d8013fe87780dd707e10c817c

Observation 1440bf94-3c60-4b0a-a6ba-06e2acd78b34 · outbound

This paper cites SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Segmentation.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens SurgicalPart-SAM: Part-to-Whole Collaborative Prompting for Surgical Instrument Segmentation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.355997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.355997Z digest=sha256:a170bcd216d3fdb4aa166fc25689722998f4ecb8e64f1f208e8307c5beb05bcd

Observation 8bf437b6-282a-4b7d-9fd5-9bfd9b010e8e · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.360911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.360911Z digest=sha256:1699bec1698a26704aa91d243e14d395229fdf9a40b5207a536b7ce43b56772c

Observation de2b71cb-7949-42ce-9489-ce89004ba5e7 · outbound

This paper cites LLaV A-NeXT: A Strong Zero-shot Video Understanding Model, 2024.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LLaV A-NeXT: A Strong Zero-shot Video Understanding Model, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.914436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.365795Z digest=sha256:4be6af318d43c75af64f7d4bc670b55995a29142b1dda4175418edfde608a246

Observation 2e7fbe65-0514-494c-b5a9-b5f9956e60c7 · outbound

This paper cites Needle In A Video Haystack: A Scalable Syn- thetic Framework for Benchmarking Video MLLMs.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Needle In A Video Haystack: A Scalable Syn- thetic Framework for Benchmarking Video MLLMs

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.897803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.370070Z digest=sha256:59e06641eee211e29b0fe71c7cb5ec48b7145fc6f609a95e21c6e50c17637d79

Observation 1d57da18-8074-49fb-bc71-4fbfdcd2012d · outbound

This paper cites Clip in medical imaging: A survey.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Clip in medical imaging: A survey

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.880456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.374598Z digest=sha256:ae53f9be9993ad47b1ea56d92cffe95e333f7bf87063288736397cc2eccc81ff

Observation 4e9f2669-99f0-4716-809f-fdfd699558a7 · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T16:38:23.379009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:38:23.379009Z digest=sha256:36257905eae62217b063103568c0a228ed8a742fe9962f61b076014bee19d33d

Observation 4f8d223f-0972-4550-9f92-406505997ab6 · outbound

This paper cites A closer look at the cls token for cross-domain few-shot learning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens A closer look at the cls token for cross-domain few-shot learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.864704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.383842Z digest=sha256:33a4e6bd47ea6ae496ab97a78dfdf427a666159335582151bcce96e78aaf05f4

Observation 1a170a45-5858-417e-ad21-d9e763a9e8ce · outbound

This paper cites Training Details We adopt a two-stage training strategy [9, 30, 33], dividing training into pretraining for modality alignment and fine- tuning for instruction tuning.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Training Details We adopt a two-stage training strategy [9, 30, 33], dividing training into pretraining for modality alignment and fine- tuning for instruction tuning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.849232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.388255Z digest=sha256:93d3f214a12f521da2aa829e0855b8b92c13f9f7c9c367c061382c9d51df34a4

Observation be9a685a-eca2-48a4-a5dc-bca19ff2bb72 · outbound

This paper cites Additional Discussion on Different Frame Se- lection Features.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Additional Discussion on Different Frame Se- lection Features

Reference 66

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T16:38:23.834550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.392641Z digest=sha256:e93ad558c07cd70b70113e04b9166a765e5d8aec8ecc115df52efdbddc75bf53

Observation f95bd0b0-8aba-4b16-b736-e932672b0924 · outbound

This paper cites Game Science.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Game Science

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:38:23.817966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.397829Z digest=sha256:16cbf5f19cd07d9d2a40fc315410e7410fd36b5d548e7e9f1e5e66c3e8a13ad4

Observation 3039831d-a05b-48b2-935a-8e8423126b96 · outbound

This paper cites an unresolved cited work.

B-VLLM: A Vision Large Language Model with Balanced Spatio-Temporal Tokens Unresolved cited work

Reference 470

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T16:38:23.943843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T16:38:23.343023Z digest=sha256:0f259d10471e35a5df8dcbd9fc7326b5ed474951754585bddc82bd91db24630b

Pith citing papers

No inbound Pith citation observations are available.