Pith. sign in

Paper Citation Record · LEDGER

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times

As of 14 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2506.00928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00928 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:08.643991Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact5
  • verified fuzzy6
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 661074a2-ba1d-4c2a-bb75-d26358803388 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.065790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.065790Z digest=sha256:89de8db901bf22d41816c178e0ac815917ee824056132f5cb27dab08e5f4405c

Observation 6fbbaa43-276b-4ea5-8110-f20991041e7c · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.753863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.079290Z digest=sha256:dcc7cd53016d80f78b41ce6f821a5669c48d8bde0d6c2afcbac8797d5d3225d8

Observation f205f64e-d814-4608-b617-3f2535bf5cc7 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.090193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.090193Z digest=sha256:27c2e3ff1782a9d18703fcf447b0416fc165f46abd8dc918796d291825270b80

Observation 8149cf40-87c9-4658-8393-38ccfb40a981 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.712770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.100925Z digest=sha256:22ff85756c21273be301732bed2c5827a27b352c38d9030e62be1ac3c6ccdfde

Observation 1f05cdf1-6cb3-4bc6-b8dc-6e29a695e0b9 · outbound

This paper cites Qwen Technical Report.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.107125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.107125Z digest=sha256:a43408337d9edbace64e469398d4e72bfa03b7444967fa32c9e8857e437992e6

Observation 190c08d5-ef08-4c04-bad2-3aaafccb369c · outbound

This paper cites Bosch, Mathilde Chailleux, and Francesca Foppolo.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Bosch, Mathilde Chailleux, and Francesca Foppolo

Reference 6

Resolution
verified exact
doi, observed 2026-08-07T11:58:08.773160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.113594Z digest=sha256:5b621aa30a3e323514dee7c0ba010680c8f4ccf754b5e887648b5127bc29c3ff

Observation 305f6973-6c73-49c7-a805-d19b14918540 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.121411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.121411Z digest=sha256:29f06bda5ca9b26b4f3878c1194298ceff20b6ff714e87bda856ec5893e3505c

Observation 5174ea98-94dd-4396-b95f-d4ecb38eeee4 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.684219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.137586Z digest=sha256:59de1bc8351ce65f99b3bfd0d0ab97362891edaab761ff943b0e779f2965ae40

Observation 45f4a0c7-581e-43c9-9e16-0f09708d677b · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.148172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.148172Z digest=sha256:3b89bb2eb54994e4f5897e2e242ff97c79164543e1d8bae71a4d7a04e4b9468f

Observation 7af4940d-214d-4cf0-b8e1-93b564dfb6be · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.163052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.163052Z digest=sha256:862c1f8ef5537d9eb1a052eb753b881b994cd6291c56c5c113fe476780ad7327

Observation 780dd187-28cd-426c-b1c3-4b8d4be3d655 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Gonzalez, Ion Stoica, and Eric P

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.169555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.169555Z digest=sha256:dae4750029059ec0729eba79563bfaf36751ea6d7d5cf149cc7fa404f7cc1e70

Observation b3ea757f-cda6-4708-969b-56f7866b53d9 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Scaling Instruction-Finetuned Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.177631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.177631Z digest=sha256:d8ed2ad07b9c95de1a6fb6a7bb3e2baeafdd908887ea02790822ccbb9ecd198b

Observation abe87411-74c7-4815-b49a-6688ee090634 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.613626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.185701Z digest=sha256:34f3ff415c69038a8b505fd6825dbe081a56150d7e76c6ba43f309173c7d41b4

Observation ea391800-8fe8-4cec-ac86-e26b7f1da069 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.589722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.194140Z digest=sha256:2053b04550b7485c59bbf73f4cc2f135fad0599a3b7e773e31db4fa9c43257e8

Observation a67c9357-e855-4797-a47e-c00cf4454665 · outbound

This paper cites Bosch, Ciro Greco, Maria Nella Carminati, and Francesca Panzeri.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Bosch, Ciro Greco, Maria Nella Carminati, and Francesca Panzeri

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.548424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.205597Z digest=sha256:0a1eed4d9031f4f59bc8f6a66a9f70aef0258cbce855458073775138d476fd5f

Observation 3123e7e9-6174-4333-8504-59bf6df14862 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.511873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.222972Z digest=sha256:91a6759b6303f9ae3e85fd9c99006a5dcede356708d9ea7734cf781d04c487de

Observation 88f3d769-d1dd-4af2-870a-aa6bc7679af3 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.229550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.229550Z digest=sha256:47da8cccd8d810788e56e86bcecb8edfa1550599f4217197e1f12871db13cff5

Observation 33d5291c-fc49-4bb6-ac6b-ff35595f70e3 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.243662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.243662Z digest=sha256:19b2eb5a2a889058ed55f03e5801d834e92b969a1623085575aea60be641878d

Observation 88dececd-f3ee-419b-8eab-d98ae24b3419 · outbound

This paper cites Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.452676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.251953Z digest=sha256:7fb989d92dee5cd0deba27f185e90629ec09d4c321bb8c7fcfb5d7ca8087fb91

Observation 101a6953-0694-4eef-9cb4-f98d0e2002ab · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.409327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.261181Z digest=sha256:e19064f85d02f3397208167f8f3cbcd12bd90be957696ea6fbb6c20d5b5d0d46

Observation 7a86d6d0-763f-4431-9bb2-a47548082387 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times The Kinetics Human Action Video Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.269901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.269901Z digest=sha256:3a9c717f9d43997bc9ca723f5a3f7efe9fd854e1efa0210e640a8d56efa6f8c6

Observation 8ab32041-ba94-48b4-9e0e-184aec77ae79 · outbound

This paper cites Richard Landis and Gary G.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Richard Landis and Gary G

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.379407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.279360Z digest=sha256:47ca19b9a56500c1cdd02f6265b8a8e1f1f5c58f75ac791f597ff740061852f3

Observation 6ba89717-7e6b-444d-8765-2e136566d15a · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.355305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.294485Z digest=sha256:f22e5bdd421a3c7a402dc03b9eadd994381c9ff8db55afa5c92b676e63e8f6e7

Observation 8d77b7a2-50c7-406a-9503-d99203cf17a7 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.328245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.304150Z digest=sha256:5041db9c4b7e8a7f0a919a63cbe5ead9f8b52e5024ef9ce5c1e87205c516fd24

Observation 84b0703e-e012-4999-b0df-827ebc4bc01f · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.299380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.309019Z digest=sha256:91a4e8196a487338178671b0ace799fff765794e8c4433b298e22d2405d655c1

Observation dbb4b5e4-3a7e-4e76-b17d-325ebe888241 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.318031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.318031Z digest=sha256:96cfe9ac5f04e8736d92ff26faf55784adcff11f27a0f8dc829c3c84db9dbae5

Observation 003702f3-d3bb-4edb-aeea-0450b3168719 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.270819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.326369Z digest=sha256:b7cbff99a225bbb01d54f4d1a08d1ee38a69bcc7b311ef7e09eb3e985e523d27

Observation eda7a662-bae1-4f30-a5ad-1902a40b912d · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times TempCompass: Do Video LLMs Really Understand Videos?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.340059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.340059Z digest=sha256:a634fc647113839c7fb2dee438ca6bc317cbde83e8f9f0c692d11b910344a441

Observation 57753b8c-62ef-4678-b831-78e636faa163 · outbound

This paper cites Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.355931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.355931Z digest=sha256:1a7e4b1c398cfec986b76a1e1708bafcd02b37231358c7eb1ae0123fd3f9123f

Observation 77c7fac6-a0ae-4e54-812f-90d8282ef350 · outbound

This paper cites Agentivit\`a e telicit\`a in GilBERTo: implicazioni cognitive.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Agentivit\`a e telicit\`a in GilBERTo: implicazioni cognitive

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:58:09.253700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.363088Z digest=sha256:48db2874fee04895704e9e59463dec6a76399851eea063ea56d42e021c73b555

Observation df6d49b8-2b3a-4ee0-aa00-408b69daaa68 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.371169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.371169Z digest=sha256:d98bbc0e20fa758a2fa7b707e9094e80d1e079d897f615ee6b76342c22e0dfaf

Observation 7f14d5e1-ef21-48b7-91bb-e9ded914a0e7 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 32

Resolution
verified exact
doi, observed 2026-08-07T11:58:08.730314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.380118Z digest=sha256:a8e4b7da3247b603ee09713948f269313baff892d207125ba6f780be3952dd5c

Observation 65869796-290d-4a5f-ad21-318311b7bf4b · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.240152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.391886Z digest=sha256:a99d4a084143c7525b5980bac38d866caabe7b488cff3f082b544de89dff8947

Observation d6e8a864-4a98-46ec-9279-f20853c0ec43 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.203402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.398684Z digest=sha256:b6ff3ff73b53ccec77e07b83a05d0c4b8ed084571e82c1a9a58eda2aaf510c02

Observation 357174b5-9dd7-47cb-931d-64eadb196113 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.163004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.405770Z digest=sha256:a3b7a04a74b8e96a926e86bd66fc9699d97bbc0906b46722fd05a44f184df012

Observation 5814384c-e745-4a4d-ae37-4a4db5be6698 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.414163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.414163Z digest=sha256:a034805946c462bb9b630a6ff7ca976ca2bcb34d8278448dc4d303296c093a6f

Observation 97b4ab7a-f919-43e0-b666-cb0904b0fbda · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.420869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.420869Z digest=sha256:ad3be60fc58d41f46e8d581f3834b280883379b492b23f733ce0cd210948224c

Observation 9911d2db-23b7-4b05-af9d-22fde91c33ed · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.067753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.427772Z digest=sha256:122f50bdb1da89dbf6a5ed9e3acd4c36af62051f3a20701f2534f34818bd4dd8

Observation 8a725ab7-a015-451f-beec-04fa9aeb1814 · outbound

This paper cites Anwer, Tim Baldwin, Michael Felsberg, and Fahad S.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Anwer, Tim Baldwin, Michael Felsberg, and Fahad S

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.011609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.437777Z digest=sha256:e4c399377fb0b52a26e736a40c9778e05f09672c75cba93df7ac4979f6ddbfd9

Observation 3b6ff0a6-05be-4876-a83e-04b6bbee94d9 · outbound

This paper cites Video Question Answering with Phrases via Semantic Roles.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video Question Answering with Phrases via Semantic Roles

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:58:09.185982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.449747Z digest=sha256:753c756bf0146ad8cda11607870079a6de784ddef149370787fb3d01d6e1a51d

Observation 7b4c77ae-1998-49f9-96b9-b7dfba8e7405 · outbound

This paper cites Sigurdsson, G \"u l Varol, X.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Sigurdsson, G \"u l Varol, X

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:09.970795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.463803Z digest=sha256:b3a5523a5e560877e61153dc289aa2af7b4affebb9ff2a5e460471d17b8357eb

Observation 47f487db-cb26-4286-83b1-8ad8e5541a67 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Gemini: A Family of Highly Capable Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.473338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.473338Z digest=sha256:6f2cfbd35ae701cc9fb5ca70765c9b8efd57e721ad1900c7d7695a45ea9ce1d5

Observation 99e957c6-d17e-40dd-832f-b4d4bd4c4be1 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.940208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.484690Z digest=sha256:2b6c6b3ab25b15a8bb2dd4781b93bb06e7313e54d0083996b5c908b0bc17f131

Observation 6f28e24c-fd11-4442-b9a1-5dc98113b0fe · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.490321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.490321Z digest=sha256:4e7279f3dd367991e90f5f534a96c70949e7961ee0c8f52434424d11134f789d

Observation f9412774-1016-4c9d-ae3d-b17c4d922938 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.907561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.498573Z digest=sha256:f813d04d31a83792093a7b125312c482b27c0a4fc0220b7dd5b186ee145005b5

Observation 142db803-8e86-4a28-bcc5-e8239baec08e · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:09.879043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.509837Z digest=sha256:a37be15eed443f3f7a1f7370bb04497fa103935d667e6cd94dda2e1f06765757

Observation 71817efc-3217-43ed-ad6e-3f631e95fc82 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.520364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.520364Z digest=sha256:832d1115a35b4f9d1f886cc570b77477bb139d9616075639a57de7193fa8b8ac

Observation 3a903595-c68d-4c9e-9d41-4151345d5212 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.844726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.531169Z digest=sha256:17d03956e8e679ee628ef1c9757c8e9ef09082f7f7c98a70781bbdc522add0da

Observation 7de2259f-8e1f-466d-8cfb-4bea5307059a · outbound

This paper cites InstructionBench: An Instructional Video Understanding Benchmark.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times InstructionBench: An Instructional Video Understanding Benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.540817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.540817Z digest=sha256:799a3dbd9d39f1127115736cbebc9ffd193360e1615472e7ec4f6f3f6d68ae7c

Observation 73c04c96-407e-490a-a3fd-f5f0974c377e · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.813619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.551390Z digest=sha256:f1df272c9a832408ab48024d528e64e3c9b3dc7a8a5e44a39c4c452743cad555

Observation f87a7a24-3391-4e34-ad8c-22d14662f196 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.559129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.559129Z digest=sha256:cf6a6a773bd0209ff4ef7e52b3021582eec50d6e4fb1cd2aa5a61b133663a2d2

Observation dfd73e6f-180f-468e-9b48-2121b184fcd2 · outbound

This paper cites Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.572193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.572193Z digest=sha256:d586d452ed43950988dbeb32ea18208c60abae929536472c8622fe382fc6e901

Observation 583eb29f-cf09-4dcb-a04d-1b18f0d70251 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.580265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.580265Z digest=sha256:77ed49e171d0e4673ab0ccb21756a0d57c752aa453a4faf85e796e291ba124dc

Observation 0f910687-a0df-43de-a1f3-02378c2032ba · outbound

This paper cites ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.586217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.586217Z digest=sha256:1a6db842561fcd3b79edbe925a9a7cb2abd8a95d79aa758d32a855b54dbdc7fa

Observation 4a733a98-b3dc-49d6-b2e2-f4ad12fd0db0 · outbound

This paper cites Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.593276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.593276Z digest=sha256:0af3c08820849ac84ba071797154442b70df4c2b5e11890baf8ac47ebed74d38

Observation a858a6da-6e70-4864-8619-3aaf6e95981d · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Sigmoid Loss for Language Image Pre-Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.598707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.598707Z digest=sha256:1c78f7a02f18af77781c63fb265e7cb159e2d49d0a3ea0bb7ae05fdb65534f61

Observation 5762f209-5a99-4b11-9c1c-55cd14564bbe · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.604138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.604138Z digest=sha256:c392c5abdccfab8eaaaa2fedb7e6e69c31da28a5d682941094893c8df8f37afc

Observation 41957909-cf19-4920-ac01-d848ae17fb67 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 58

Resolution
verified exact
doi, observed 2026-08-07T11:58:08.704124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.611110Z digest=sha256:c7c4981464173a94b92dafce35700946196391980bc00259ceb3cb9883e93284

Observation 8896bb95-5376-4b02-a7d7-7c0f7a5e5aa9 · outbound

This paper cites Video Question Answering: Datasets, Algorithms and Challenges.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video Question Answering: Datasets, Algorithms and Challenges

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.626602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.626602Z digest=sha256:684976b98e6a675d5c0b8c4b2844f6f0f32db3fd6359cf52b2965906c3766ab3

Observation f4ff0b5f-be66-46ea-94b5-cd49718bdd99 · outbound

This paper cites online" 'onlinestring :=.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times online" 'onlinestring :=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.634847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.634847Z digest=sha256:a1defbf397e544296356940430c4c76b63a5351e20a474fac2b85fe0b7af5ae1

Observation 7ee018d2-2192-4ea8-be87-1c20cc788438 · outbound

This paper cites write newline.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.643991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.643991Z digest=sha256:cd866bd001b240067ae174b8aac30789a1ede6261baa8494e45fe852e0587044

Pith citing papers

No inbound Pith citation observations are available.