Pith. sign in

Paper Citation Record · LEDGER

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times

As of 15 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2506.00928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00928 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:08.643991Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact5
  • verified fuzzy6
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 661074a2-ba1d-4c2a-bb75-d26358803388 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.065790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.065790Z digest=sha256:d733a7a1f48a550b8c5e4935f9531b0d56f040e1c03d78ff56c76c14eda3bc67

Observation 6fbbaa43-276b-4ea5-8110-f20991041e7c · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.753863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.079290Z digest=sha256:47d2949669ebe4e548da4b3f36c42c0f3285ccabe42688971a16ec9206c6a1f9

Observation f205f64e-d814-4608-b617-3f2535bf5cc7 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Flamingo: a Visual Language Model for Few-Shot Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.090193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.090193Z digest=sha256:293e492682004f23c4a73d2fd78d5ee2969a3a7f187c69758b434a077379290b

Observation 8149cf40-87c9-4658-8393-38ccfb40a981 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.712770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.100925Z digest=sha256:79d83ec459ac01506ab54e50937823f430bbb29fc342bea1143250d1662b4ca3

Observation 1f05cdf1-6cb3-4bc6-b8dc-6e29a695e0b9 · outbound

This paper cites Qwen Technical Report.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.107125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.107125Z digest=sha256:53984a838ceb669af4dd59db15e9f78eee18b3e562be6c581830b8b0a8fd5dbd

Observation 190c08d5-ef08-4c04-bad2-3aaafccb369c · outbound

This paper cites Bosch, Mathilde Chailleux, and Francesca Foppolo.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Bosch, Mathilde Chailleux, and Francesca Foppolo

Reference 6

Resolution
verified exact
doi, observed 2026-08-07T11:58:08.773160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.113594Z digest=sha256:d9957806a471a54ccc4d488f73e1d9c3b780cb5074bf14afe242c7d2a56e1960

Observation 305f6973-6c73-49c7-a805-d19b14918540 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.121411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.121411Z digest=sha256:0dee4078073edd8c11159ac69812973cd6f0c3cec5846f50bbd1a686111e8510

Observation 5174ea98-94dd-4396-b95f-d4ecb38eeee4 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.684219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.137586Z digest=sha256:2ce3639ae403dff65bdd8e27b790ec92c3178f233e7d8dd7f82aa5180441fcc0

Observation 45f4a0c7-581e-43c9-9e16-0f09708d677b · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.148172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.148172Z digest=sha256:cfe09cc517fff765f1f9c6e07fe71d8d05db604e248aa79f30b8d033a89d1d71

Observation 7af4940d-214d-4cf0-b8e1-93b564dfb6be · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.163052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.163052Z digest=sha256:ee57065dcd0afae02243d8ca7e6b7cecc1089e40095de4568b09b20f4c0c46cc

Observation 780dd187-28cd-426c-b1c3-4b8d4be3d655 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Gonzalez, Ion Stoica, and Eric P

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.169555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.169555Z digest=sha256:690acb76665b8961d59cbf98acf1ad40624da04404187458f8c1306b80968af3

Observation b3ea757f-cda6-4708-969b-56f7866b53d9 · outbound

This paper cites Scaling Instruction-Finetuned Language Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Scaling Instruction-Finetuned Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.177631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.177631Z digest=sha256:6ab0151efe1be68d2b8d5b3812ac40a66d56b5b4bc12985c6f3754af35ec1c4d

Observation abe87411-74c7-4815-b49a-6688ee090634 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.613626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.185701Z digest=sha256:919d635e0092fe07b0b46c08fb0afadc5b148c1e08bf90c5451d39f46a72d283

Observation ea391800-8fe8-4cec-ac86-e26b7f1da069 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.589722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.194140Z digest=sha256:563fdc2b68a2c34c46eb0174d13c4b48aba9804558617e3d4d1104a6620c706f

Observation a67c9357-e855-4797-a47e-c00cf4454665 · outbound

This paper cites Bosch, Ciro Greco, Maria Nella Carminati, and Francesca Panzeri.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Bosch, Ciro Greco, Maria Nella Carminati, and Francesca Panzeri

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.548424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.205597Z digest=sha256:2e5741d3d2e56cf39208b37efd7fb2b0f016717ac23282c5dd4b4b426c521301

Observation 3123e7e9-6174-4333-8504-59bf6df14862 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.511873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.222972Z digest=sha256:fcd9b290e3897a894416254f76cc395a7aa91bd168727b0200898744824486ac

Observation 88f3d769-d1dd-4af2-870a-aa6bc7679af3 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.229550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.229550Z digest=sha256:d407abfc445eb792c0f81167af2c42d19659402e53b01df592594d23bbf588ba

Observation 33d5291c-fc49-4bb6-ac6b-ff35595f70e3 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.243662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.243662Z digest=sha256:a80d6768db0ed4be6aaa607cee63b5934ee1aec8466e32b339fb03ad110d6167

Observation 88dececd-f3ee-419b-8eab-d98ae24b3419 · outbound

This paper cites Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Jang, Yale Song, Chris Dongjoo Kim, Youngjae Yu, Youngjin Kim, and Gunhee Kim

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.452676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.251953Z digest=sha256:1f599b75ba1016238c74fe9f5a6f1a45456f98bec35c8efa729bc44bca6e5d5a

Observation 101a6953-0694-4eef-9cb4-f98d0e2002ab · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.409327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.261181Z digest=sha256:35925cae26ece15453898ca7fe9e94e598a222c819f2776851e00a36bc478be1

Observation 7a86d6d0-763f-4431-9bb2-a47548082387 · outbound

This paper cites The Kinetics Human Action Video Dataset.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times The Kinetics Human Action Video Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.269901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.269901Z digest=sha256:a931c5b1d35656a05585b95bda08b932fa1265f76ac8757f4829839fc44c9283

Observation 8ab32041-ba94-48b4-9e0e-184aec77ae79 · outbound

This paper cites Richard Landis and Gary G.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Richard Landis and Gary G

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.379407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.279360Z digest=sha256:1a555a337e2f13c928c996fd3b4fe050fe6979874178aa50057395f504c40806

Observation 6ba89717-7e6b-444d-8765-2e136566d15a · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.355305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.294485Z digest=sha256:d8fd8a0cb0ec7ee99981751e9e32fbc2b5d78bf296ea93aa4417338da0b81c5c

Observation 8d77b7a2-50c7-406a-9503-d99203cf17a7 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.328245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.304150Z digest=sha256:d426a4cb44fde38ec0f1d9c9985dc702e916b581cc34b28561c6c060ad02542f

Observation 84b0703e-e012-4999-b0df-827ebc4bc01f · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.299380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.309019Z digest=sha256:ba7fb936f92e84595e6d0c02d75bd2ddfa4b9e03e588fe76bb88a7fff0ae8b71

Observation dbb4b5e4-3a7e-4e76-b17d-325ebe888241 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.318031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.318031Z digest=sha256:1dd7175272316c7d4d920eaa6eb667128126eacf0821c883bb5446255c7faa48

Observation 003702f3-d3bb-4edb-aeea-0450b3168719 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.270819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.326369Z digest=sha256:5d4c019f8527b3b8b7d66cf4dc345111f5a80234b259692bdc7f2d27d5902743

Observation eda7a662-bae1-4f30-a5ad-1902a40b912d · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times TempCompass: Do Video LLMs Really Understand Videos?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.340059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.340059Z digest=sha256:e7a0e033cfc044ce66a2ee2a481f376177b49df6aa6ff8a2946d081f25e3729f

Observation 57753b8c-62ef-4678-b831-78e636faa163 · outbound

This paper cites Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Addressing Blind Guessing: Calibration of Selection Bias in Multiple-Choice Question Answering by Video Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.355931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.355931Z digest=sha256:54369a7656f79f3f168c0cd84120fdd1118c03b3ce4ceb4fd54e2ed392957e55

Observation 77c7fac6-a0ae-4e54-812f-90d8282ef350 · outbound

This paper cites Agentivit\`a e telicit\`a in GilBERTo: implicazioni cognitive.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Agentivit\`a e telicit\`a in GilBERTo: implicazioni cognitive

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:58:09.253700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.363088Z digest=sha256:dcddba5147b3dbcc4a58b29d9d64d1afed12a4d22bfaa49ccb141a5298e661de

Observation df6d49b8-2b3a-4ee0-aa00-408b69daaa68 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.371169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.371169Z digest=sha256:26ecd0a5bcc343ec9226e7dc0c1306983775ec9a08e275ef70bfb553db4f594c

Observation 7f14d5e1-ef21-48b7-91bb-e9ded914a0e7 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 32

Resolution
verified exact
doi, observed 2026-08-07T11:58:08.730314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.380118Z digest=sha256:2dce9d11ea931479d7b93bbecc918e47ae53ad0571071be44673cc46dcc2da6a

Observation 65869796-290d-4a5f-ad21-318311b7bf4b · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.240152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.391886Z digest=sha256:e6a3d3b5a7a75dce4ad887981a9a94a1b50ccf8b46f83f6c6443ee30d844680f

Observation d6e8a864-4a98-46ec-9279-f20853c0ec43 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.203402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.398684Z digest=sha256:dc9374dc95559c5b81d4701034c72cc86d06b981946716f98030c71601d5bb97

Observation 357174b5-9dd7-47cb-931d-64eadb196113 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.163004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.405770Z digest=sha256:5efb5bd219dd24679f8da9ebf1bbd71801bdb37428d6ed941f725c4d58f308b3

Observation 5814384c-e745-4a4d-ae37-4a4db5be6698 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.414163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.414163Z digest=sha256:0ab37ae8a8afd02f33a5cadcfef654f25c0f6fcc3e0598aced1e12e3d2f9c136

Observation 97b4ab7a-f919-43e0-b666-cb0904b0fbda · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.420869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.420869Z digest=sha256:41d01acb836591f68aa1b444f1ed30f20a61243a8adbe4286d8988515afd3ec0

Observation 9911d2db-23b7-4b05-af9d-22fde91c33ed · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:10.067753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.427772Z digest=sha256:5ef08e8a412e7e4451771c29ffb6b20613280565bd0768861d05dd1d8c1b2a58

Observation 8a725ab7-a015-451f-beec-04fa9aeb1814 · outbound

This paper cites Anwer, Tim Baldwin, Michael Felsberg, and Fahad S.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Anwer, Tim Baldwin, Michael Felsberg, and Fahad S

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:10.011609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.437777Z digest=sha256:28e0febedb56a895c09c864887553904df8a4275e2b9842ca9ea95645623674b

Observation 3b6ff0a6-05be-4876-a83e-04b6bbee94d9 · outbound

This paper cites Video Question Answering with Phrases via Semantic Roles.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video Question Answering with Phrases via Semantic Roles

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:58:09.185982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.449747Z digest=sha256:70d861505f92fcabb44771ea116cc717b7ab8f17d0a39c9553f22c7c9898b8ee

Observation 7b4c77ae-1998-49f9-96b9-b7dfba8e7405 · outbound

This paper cites Sigurdsson, G \"u l Varol, X.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Sigurdsson, G \"u l Varol, X

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:09.970795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.463803Z digest=sha256:8fa8e6188b0e0031b49eeb8ecb75f711d66a60c0b620b93e9c67dcdef03dbcf8

Observation 47f487db-cb26-4286-83b1-8ad8e5541a67 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Gemini: A Family of Highly Capable Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.473338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.473338Z digest=sha256:c94d98d5f2e7289e904172d06e340e9b7b87148a1fc6adf19077686aae871ff5

Observation 99e957c6-d17e-40dd-832f-b4d4bd4c4be1 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.940208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.484690Z digest=sha256:f4ad5eeb9224e278a8c9b9b8276889f66464cc06bbb531d2aeb44089ee373443

Observation 6f28e24c-fd11-4442-b9a1-5dc98113b0fe · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.490321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.490321Z digest=sha256:fd474998b37371f6b9daaa66c7116a0d1389cc56b65e585efc00eee04051b955

Observation f9412774-1016-4c9d-ae3d-b17c4d922938 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.907561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.498573Z digest=sha256:dba2f457093dd070c5643193cb7c92365569d2b13a23203fca1c53b3c8451ce4

Observation 142db803-8e86-4a28-bcc5-e8239baec08e · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:58:09.879043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.509837Z digest=sha256:c2ead7fda0ed530f2342b3b8d9502c3d7ef2c4f82e588a941a9c55434c205d14

Observation 71817efc-3217-43ed-ad6e-3f631e95fc82 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.520364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.520364Z digest=sha256:238e7135d8afa3707f7b8f9f1675f8296aeddee886aa0ed4122827dc450c9ffa

Observation 3a903595-c68d-4c9e-9d41-4151345d5212 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.844726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.531169Z digest=sha256:6fbcc8f80dc6f2b7d21fc46d6d0dcea95ad4bce1c8e81948be23dbedaefb4f0f

Observation 7de2259f-8e1f-466d-8cfb-4bea5307059a · outbound

This paper cites InstructionBench: An Instructional Video Understanding Benchmark.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times InstructionBench: An Instructional Video Understanding Benchmark

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.540817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.540817Z digest=sha256:4ebccc3e6afdd0734f64a4d5846b2c097181f7ff951249d91ce8949a09526f0d

Observation 73c04c96-407e-490a-a3fd-f5f0974c377e · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:58:09.813619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.551390Z digest=sha256:2bd6f54afc3cfb2269ab71eb54578b914f2186ff75cf458a42b1c3469ae20f75

Observation f87a7a24-3391-4e34-ad8c-22d14662f196 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.559129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.559129Z digest=sha256:6fcd974df8c9e75c3d734d835cf691a13e5f506b806b41af2f53e7d7f5b2775e

Observation dfd73e6f-180f-468e-9b48-2121b184fcd2 · outbound

This paper cites Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.572193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.572193Z digest=sha256:6ea9d18f3295791147ce83512312e115a050ef9d57140d33ddb517d61ad69b2e

Observation 583eb29f-cf09-4dcb-a04d-1b18f0d70251 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.580265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.580265Z digest=sha256:1825e197d177e776d6603d357958ab0cf05b583eb937141594f93cbbf8f15be0

Observation 0f910687-a0df-43de-a1f3-02378c2032ba · outbound

This paper cites ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.586217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.586217Z digest=sha256:c93bf0279c975581a901845b5ecc5b30e4a2b111fc21525f4de63df75ed7d670

Observation 4a733a98-b3dc-49d6-b2e2-f4ad12fd0db0 · outbound

This paper cites Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.593276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.593276Z digest=sha256:589fbefeee6e6d92d5195ef61758c04ae49ecfe8bc891c8f05f8ca5eefc84c48

Observation a858a6da-6e70-4864-8619-3aaf6e95981d · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Sigmoid Loss for Language Image Pre-Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.598707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.598707Z digest=sha256:bccce8f5d22f490c4612b2ec86ea3b44fc1cc984462df403d21e08d889bc88ec

Observation 5762f209-5a99-4b11-9c1c-55cd14564bbe · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.604138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.604138Z digest=sha256:072d0b95a99eb76228286ba058add5a2403289c06494525be868f1122166ed0d

Observation 41957909-cf19-4920-ac01-d848ae17fb67 · outbound

This paper cites an unresolved cited work.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Unresolved cited work

Reference 58

Resolution
verified exact
doi, observed 2026-08-07T11:58:08.704124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T11:58:08.611110Z digest=sha256:18804e74ad36b382f666456d4721649c5e85eab78a00968f7437b7867e357fcc

Observation 8896bb95-5376-4b02-a7d7-7c0f7a5e5aa9 · outbound

This paper cites Video Question Answering: Datasets, Algorithms and Challenges.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times Video Question Answering: Datasets, Algorithms and Challenges

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.626602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.626602Z digest=sha256:a9e432c3ac93e98fd88104689f3662a55c465f690d48f0a603bc65e37227bc37

Observation f4ff0b5f-be66-46ea-94b5-cd49718bdd99 · outbound

This paper cites online" 'onlinestring :=.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times online" 'onlinestring :=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.634847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.634847Z digest=sha256:d722438180d5ae40882c7306e93c75c465721e97885feeec92256b53e452d85e

Observation 7ee018d2-2192-4ea8-be87-1c20cc788438 · outbound

This paper cites write newline.

Deep Temporal Reasoning in Video Language Models: A Cross-Linguistic Evaluation of Action Duration and Completion through Perfect Times write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:08.643991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:08.643991Z digest=sha256:3235490ee65a887bef5d8e7fba0cafb6c50db0c4af2c133801f46015e5111747

Pith citing papers

No inbound Pith citation observations are available.