Pith. sign in

Paper Citation Record · LEDGER

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models

As of 18 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 2 inbound Pith citation observations for arXiv:2508.19650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19650 v3

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:40:40.564012Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:16:07.090870Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T08:16:47.773102Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved15
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65e74ff3-ca57-4317-b64b-72922a10f5e6 · outbound

This paper cites Pixtral 12B.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.375797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.375797Z digest=sha256:1c0e0a3daee3d894d4d28027ce5437e3ce7c52c9befd9c1cce0d77db00187d10

Observation b145b217-8f77-4661-b58a-012bfb679beb · outbound

This paper cites an unresolved cited work.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:40:41.324211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.379826Z digest=sha256:36f9d2718ad83ccfd144fae12dbeeffcbdd0391b5076270cfe91009723402fc5

Observation 37c1ab81-66d2-4275-8218-f814e6f8b074 · outbound

This paper cites Claude-sonnet-4.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Claude-sonnet-4

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.315764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.382974Z digest=sha256:2049286f1c8ed92542b66a4e60daaa75be6e6ca16ea637ff7ba3360dcee255e1

Observation 5689d440-2051-4c0b-bc89-e1e4b1c65b74 · outbound

This paper cites Minigpt4-video: Advancing multimodal llms for video understanding with interleaved visual-textual tokens,.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Minigpt4-video: Advancing multimodal llms for video understanding with interleaved visual-textual tokens,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.306761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.386685Z digest=sha256:655ae720257e361bffaa4310c07b2bddc1271715b33170ae8f266b8de578c982

Observation 2bece409-1543-4315-8abb-85aba8fc7f43 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.297697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.389989Z digest=sha256:1d7c28a13332cb4382c3a33da0784196dc7c6af81e1b1fb21cac389b905c71de

Observation 059fc12a-ced7-4d08-af87-903701ead891 · outbound

This paper cites Qwen2.5-VL Technical Report.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.393101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.393101Z digest=sha256:91611a87c4a28e7e4710e5cec69669b8a6f061ff6c6b8086b04eca3fd986c5df

Observation c353df1a-85d1-4771-aa1b-1739b7ecfa43 · outbound

This paper cites Doubao-seed-1.6.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Doubao-seed-1.6

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.288645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.396513Z digest=sha256:73fdbf47ef2919f8f956aeeda3be4c0edac1233cba16953bc67477afa4e941e0

Observation 8b13b3b2-8aa9-48d4-a1ed-33d456936220 · outbound

This paper cites an unresolved cited work.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:40:41.279774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.399914Z digest=sha256:ce053b9383761f5352a5224407747dad035640818059f080d533970778d5b4f2

Observation ce52ab39-d687-408f-801d-df85a4ecd06c · outbound

This paper cites LongVILA: Scaling long-context visual language models for long videos.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models LongVILA: Scaling long-context visual language models for long videos

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.270684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.402904Z digest=sha256:5a0fda465c9c439dc578b5daff5e55c271d605b8795b41e310c73e515d775136

Observation 1b2fed83-7c48-4c90-bd44-ae8aa70838ac · outbound

This paper cites Videollama 2: Advancing spatial-temporal modeling and audio understanding in video- llms, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Videollama 2: Advancing spatial-temporal modeling and audio understanding in video- llms, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.260969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.405886Z digest=sha256:e39abb16d842835d4a354547c1fb6e09f4ba195d0efa949dd25127bf449f59dd

Observation 9d43fce6-8ccf-46cc-b5e3-cc9bff5c269d · outbound

This paper cites Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic ca- pabilities, 2025.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic ca- pabilities, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.251670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.408752Z digest=sha256:de28cd7f758bf3c7df8d16bc3f1bd0c69ec567df7597b861e6780bdbb668a0e3

Observation f5ab3313-f21c-47f7-864e-7f20a81a650e · outbound

This paper cites Video- mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video- mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.242533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.411871Z digest=sha256:4e159898c8a8ea4c82a5f7d41544fea0c2e3e71a09991e8fed400770e6a411c7

Observation ea51068d-30b1-43a7-97ad-ad4c47deb42d · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Ego4d: Around the world in 3,000 hours of egocentric video

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.233387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.414832Z digest=sha256:a243ecf62a828336d227d3a2ffd3748390e03f78456b9c16dd7b861c699d579d

Observation 00566207-44b0-4dc9-9499-bec7e5bbe829 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.224035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.417531Z digest=sha256:8f1c9dd4b7031919cfeef5305c6613adf55591ae10787246968aaa68ce7aac26

Observation 2307a43e-3c88-4356-b064-ede6b14e023f · outbound

This paper cites Scaling Laws for Neural Language Models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Scaling Laws for Neural Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.420423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.420423Z digest=sha256:96677ec3468777d2045915352e91eff8a8b26c35778257b44450bbba1d47e2e8

Observation b9dc044b-cff9-4370-b2ce-6b458c9ccfb6 · outbound

This paper cites Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Rush, Douwe Kiela, Matthieu Cord, and Victor Sanh

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.423571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.423571Z digest=sha256:c62937a835b2a51eaf768e9053127a86cfaef60f08de3cc4fc230d6fcfde62c0

Observation 1f10baab-a03f-448b-bc7f-5ebb11321d01 · outbound

This paper cites LLaV A-onevision: Easy visual task transfer.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models LLaV A-onevision: Easy visual task transfer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.209634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.426686Z digest=sha256:4ae6b2ae7404f205536e9445326482fbd5ab3361a7cbce416d53d243945ebe49

Observation dbb12601-14cf-4791-b2e6-2e166ac5db85 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understand- ing benchmark.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Mvbench: A comprehensive multi-modal video understand- ing benchmark

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.201130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.429618Z digest=sha256:20bde3f724c022ed1d02bbc72e82c4fdb620713dc53c09740f1c776508bfb226

Observation 1fc36da3-17db-4a3c-826c-a94a1c94c0bd · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Llama-vid: An image is worth 2 tokens in large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.192619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.432539Z digest=sha256:192902e7b11a7f20c6d5ecada2203d17fb1701cf47443128f2c1f5d1da6b798b

Observation a871293f-60e7-4828-a97a-939032e264de · outbound

This paper cites Video-LLaV A: Learning united visual rep- 9 resentation by alignment before projection.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-LLaV A: Learning united visual rep- 9 resentation by alignment before projection

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.183525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.435609Z digest=sha256:90b1b423af153fa6106eeb6a7d92cea8c4422bbaa40b18778eef69f1e589247a

Observation adce1126-a8a0-4236-b4ca-617e4f826596 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.438516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.438516Z digest=sha256:85d20e101cd8735554dd3df88a32f66aba2cffea26f858bb3abff37f1c2f21c7

Observation 38c45cc1-d510-49ce-b333-a65a322c07fd · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.174770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.441934Z digest=sha256:35fb8d4cc026ac8af653f066d4b98d8501a29ee6e08ef7d4d99c9745c0dd1332

Observation af6c663a-86dd-4ae8-b784-9963b6deba10 · outbound

This paper cites Scaling laws of rope-based extrapola- tion, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Scaling laws of rope-based extrapola- tion, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.165563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.445184Z digest=sha256:b7d0f9796d5702971f45838a14b3b8e2f9ffd8e11dd59fb2e7f3db0acf454fb5

Observation 6753fc71-50bc-4a1c-9b66-c935817c839c · outbound

This paper cites TempCom- pass: Do video LLMs really understand videos? In Findings of the Association for Computational Linguistics: ACL 2024, pages 8731–8772, Bangkok, Thailand, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models TempCom- pass: Do video LLMs really understand videos? In Findings of the Association for Computational Linguistics: ACL 2024, pages 8731–8772, Bangkok, Thailand, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.156567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.448237Z digest=sha256:970b26494092ad4e98b42b7e01c3448324e61991e36b2d74830b3bd065b0107d

Observation d7681b96-e8d3-4c04-88e3-5b215a9226eb · outbound

This paper cites Nvila: Efficient frontier visual lan- guage models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Nvila: Efficient frontier visual lan- guage models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.147633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.451552Z digest=sha256:d4ea72ee447905a02d7786f8f4c63336422663ae4033830bd3adbaecb3b81b71

Observation a0edd6fb-bef5-477d-9122-9dd90f9cbd0e · outbound

This paper cites Nee- dle in a video haystack: A scalable synthetic evaluator for video mllms.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Nee- dle in a video haystack: A scalable synthetic evaluator for video mllms

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.138613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.454556Z digest=sha256:3186e1c64d289f5f1b5ac17e087a1a7a1489e0db33d77bd1d5f7225ea3e50b4b

Observation dbd00143-8289-4c02-8ecb-57b2ec42e662 · outbound

This paper cites Video-chatgpt: Towards detailed video un- derstanding via large vision and language models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-chatgpt: Towards detailed video un- derstanding via large vision and language models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.128830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.457639Z digest=sha256:b5c24a00a89e31dd0e0ad6d72cf9b0892be1a16474be24a1aac4fa4584fb138e

Observation b27ab64c-bbdf-4ef0-90ff-b9379f5b0262 · outbound

This paper cites The serial position effect of free recall.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models The serial position effect of free recall

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.118795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.460901Z digest=sha256:bfecaa9ecaf7c8f1870bc8e4f4f61f381452290f84e5aebb00f351a26cd688e4

Observation ab257ea7-ec9f-4d31-9354-ee63024834ad · outbound

This paper cites Needle in the Haystack for Memory Based Large Language Models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Needle in the Haystack for Memory Based Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.464235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.464235Z digest=sha256:1928494dec003b5c02f2d0b6388d6fb5eae97b7ba8bc1fa1301f62529be7768b

Observation d8901391-803d-4ff6-b76f-753967c912c4 · outbound

This paper cites an unresolved cited work.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-05T15:40:41.109398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.467639Z digest=sha256:7d202ccd13a2c8455414b59599ce3b9193ae77c0d9ff815d1cf8184c4689aff1

Observation 9c773163-96be-4efa-9531-554b984c2f0e · outbound

This paper cites Gpt-4 technical report, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Gpt-4 technical report, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.100555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.470740Z digest=sha256:fbb8cbd5dd19fb460e64fdd21e4f1034b492ee3793a61c273b221aabba953b56

Observation 4605b57d-522c-4e9d-b7b3-845e14079c8b · outbound

This paper cites Too many frames, not all useful: Efficient strategies for long- form video QA.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Too many frames, not all useful: Efficient strategies for long- form video QA

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.091034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.473979Z digest=sha256:68882b0b09887ee4d6c0f4db96b6348aafc427925bfa1a32951569c1a36ae0cb

Observation bb56a904-e9af-4aa6-bde5-e33822a78dcd · outbound

This paper cites Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.477147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.477147Z digest=sha256:ef636854a7c5eedd0cc0245f0f65af6630f2d1d337664ccf0d6a42976623cf8f

Observation 09913029-dae3-4a30-9500-59136aa2bbb5 · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Video-xl: Extra-long vision language model for hour-scale video understanding, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.081785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.480593Z digest=sha256:695c10a93e1ae06786145594284dddca085508fbe0b434d9a73bea64fe623211

Observation e6cce75b-3f48-4e2e-aa07-a121e0e2311c · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Moviechat: From dense token to sparse memory for long video understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.483719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.483719Z digest=sha256:58a695488262a74ff545001ea1f7a8bce0cafafee9e7c31403d804da3c8ea68d

Observation 99f41bf0-272c-438f-aee6-58fabe8059a7 · outbound

This paper cites Real-world anomaly detection in surveillance videos.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Real-world anomaly detection in surveillance videos

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.065754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.487027Z digest=sha256:b5d0ab4964c633422e7c62a9a4597c015a1db403c864c6115c3ebc907cac0d8f

Observation d376f8ca-6071-4fce-ae16-dbcbd90f2658 · outbound

This paper cites Mimo-vl technical report, 2025.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Mimo-vl technical report, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.055225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.490189Z digest=sha256:17b6c0b460aadf7c4ef3b1ec1cc2924ee37198f6dde82b0168a23e5e8791b459

Observation e1c435f9-c960-4abd-ac46-0f9a93d2e30d · outbound

This paper cites Glm-4.5v and glm-4.1v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning, 2025.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Glm-4.5v and glm-4.1v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.045168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.493475Z digest=sha256:22e846ed2371a8d8a9506486bb4b37df756ad1e3961a16121d2b0c61bb42311f

Observation 5170d3ad-81ff-486e-9603-2d99133f9c50 · outbound

This paper cites Multimodal needle in a haystack: Benchmarking long-context capability of multimodal large language models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Multimodal needle in a haystack: Benchmarking long-context capability of multimodal large language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.034671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.496721Z digest=sha256:a8cb62e0c904b37744b8902fc48b6bd7ae0be815f8f04fd3ce25be826dcc61cb

Observation 9ce82354-d169-42d4-b055-f69ac1ffb12e · outbound

This paper cites Lvbench: An extreme long video understanding benchmark, 2024.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Lvbench: An extreme long video understanding benchmark, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.024585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.499800Z digest=sha256:50a73fee2ad4dadb48c3d8fdafedfd82e08eb2eb51742f162cf7188620a38d16

Observation 32b649c9-7fce-441a-b8e3-b0e6c0690088 · outbound

This paper cites Needle in a multimodal haystack.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Needle in a multimodal haystack

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.013993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.503121Z digest=sha256:71d6801ce56dbf4abc3c2d3249dc5fc606fd34f068cfc8c7e82b6c2331b71718

Observation a6e007d2-b029-4615-9b19-1fba9a0470be · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:41.004212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.506176Z digest=sha256:b0c39ea3cb30197184dd953c54e0c9369e588edd903d623d98e558f71999c592

Observation 92174007-f620-44f5-b9a0-82522511e19b · outbound

This paper cites Mit- igating object hallucination via concentric causal atten- tion.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Mit- igating object hallucination via concentric causal atten- tion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.993509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.509168Z digest=sha256:165fa18c5789d5e0286d5116534bc6ec59ff08a7e435686b86cffd5f31efd46a

Observation 84336aff-ced5-476d-8666-c87a70868ed4 · outbound

This paper cites Qwen3 technical report, 2025.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Qwen3 technical report, 2025

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.982894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.512623Z digest=sha256:c50774a884077b71647c21d0ddf49c10d8f15523d4236a80e7c61f66e9b59084

Observation f5b55030-f5c0-4c8c-9071-6ce23f26e5fd · outbound

This paper cites Re-thinking temporal search for long-form video understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Re-thinking temporal search for long-form video understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.972907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.516157Z digest=sha256:d2de77ef4511f26eaa1b80c91fa562dad46ec92cade81729cc8a5ad45ad8bf06

Observation 1e268707-996f-4d21-abc5-7fba0a759ac6 · outbound

This paper cites Lv-eval: A balanced long-context bench- mark with 5 length levels up to 256k.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Lv-eval: A balanced long-context bench- mark with 5 length levels up to 256k

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.519503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.519503Z digest=sha256:e192a028755d0a02beac605d183b0acd856cb2463135798d6cf4249e65e3fd26

Observation 9da279b6-3d77-42df-979f-b468a4b7351b · outbound

This paper cites Videorefer suite: Advancing spatial- temporal object understanding with video llm.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Videorefer suite: Advancing spatial- temporal object understanding with video llm

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.962573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.522801Z digest=sha256:098b0978459cf3073cca63a84fbd5dd80439b12f830215816592da1ab8fb2625

Observation a6aeea15-5916-4ff5-9296-8ba06b9838ef · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.525881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.525881Z digest=sha256:62b24bc806850a1e2fba7de445cc08b9cd73fb32de8466ba31361585384eae04

Observation f74b40df-1c1c-4347-b75b-3322c6cf4af1 · outbound

This paper cites A simple LLM framework for long-range video question-answering.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models A simple LLM framework for long-range video question-answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.952156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.529685Z digest=sha256:0a6e2a760e9a35dbdaff1e2d3a26e56512d57f6314d12fbc0d95ef33670ca5e2

Observation 5ef51ecc-6e95-4b19-a3be-6fbefba61492 · outbound

This paper cites Long context transfer from language to vision.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Long context transfer from language to vision

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.941578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.533027Z digest=sha256:432b725326f32c266fb61547bb0040e161737793887ede0eb5cec814b36108ce

Observation 495fcdb7-43d6-4704-b7bb-7734549ef176 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.536185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.536185Z digest=sha256:cd41877ee601de3bb43ef84259e1c9068972dce75e20280c9df2d4323cb0bdb8

Observation c4ad9bb9-4dd1-4262-9ca1-62d36a96c2d2 · outbound

This paper cites Mmvu: Measuring expert-level multi- discipline video understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Mmvu: Measuring expert-level multi- discipline video understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.931073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.539944Z digest=sha256:51a2e927ecab0ec005fc52e85fc90a8672630e22a3e99781c71b3f7bbe53f8ff

Observation d365e751-59a2-4134-a938-4fcb7e40192b · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Mlvu: Benchmarking multi-task long video understanding

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.919974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.543150Z digest=sha256:659aa654002660685986e2030478ae729dfefd9deee7de865078c2d7e74b68d2

Observation 5b09bca6-5e5e-4d1d-8e5a-92b3ea07ebcc · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T15:40:40.546410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:40:40.546410Z digest=sha256:ab992f5d2241c32fe864856d4324e1ffd78a5c11acaa15c21978ea299149f775

Observation e7f879f7-f741-436c-a306-c62beb11157c · outbound

This paper cites Detection and tracking meet drones challenge.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Detection and tracking meet drones challenge

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.908840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.549951Z digest=sha256:dd2fc1f2563df66b64fbfa99208f586a9b7e20a69a3e50621ea27d39776466cf

Observation 7b3cf1e6-74a7-4c33-a225-b56aaea33ff1 · outbound

This paper cites Hlv-1k: A large-scale hour-long video benchmark for time- specific long video understanding, 2025.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models Hlv-1k: A large-scale hour-long video benchmark for time- specific long video understanding, 2025

Reference 56

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T15:40:40.897972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.553142Z digest=sha256:150a38daa71a1177b91a9e73031b698d66ed4a871aa8c74b70a7e72872f014fd

Observation 2f25c3d7-ceb1-4d3f-b2cb-85be3c223669 · outbound

This paper cites three chairs,.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models three chairs,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.886292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.557088Z digest=sha256:274ffff634da58c61ab60747e45e32b0dbc1bbea29f1cdb84421ba754394f754

Observation 2747719c-ecce-4797-a12a-6d8f5f63f1d0 · outbound

This paper cites the lamp is on top of the desk,.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models the lamp is on top of the desk,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.768054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.560393Z digest=sha256:5a54211699889a603f5ed31093e3014de38f7db4bed5f7ace88ff9511dcaab5b

Observation b2ae1e6c-6403-459c-8c95-e8b2abe56a88 · outbound

This paper cites nearby" can be replaced with.

Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models nearby" can be replaced with

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:40:40.757346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T15:40:40.564012Z digest=sha256:b8868fe111fa877a6978a7b869a94a85b4a09239c8487803c7ffb4b8b1b625dc

Pith citing papers

Observation 1bfb9d1e-a881-429b-9703-92f7ff47c0af · inbound

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs cites this paper.

DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:58:05.983666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T06:55:01.619441Z digest=sha256:e6d50dc1a2c4b16e5b19659a30a557381e9b8738706843c78b9448dfbf290f58

Observation d29541e9-3524-4b6a-a74a-0cb96578d36d · inbound

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks cites this paper.

M$^3$Eval: Multi-Modal Memory Evaluation through Cognitively-Grounded Video Tasks Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:16:47.774503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T06:16:07.090870Z digest=sha256:6c959dea3d88f8619e25a4857b6f0c2024bc17a1c85c6405438f4ad32aad4d59