Pith. sign in

Paper Citation Record · LEDGER

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

As of 6 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 3 inbound Pith citation observations for arXiv:2512.03963.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.03963 v3

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T02:18:21.718091Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:17:27.916261Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

72 of 72 outbound references displayed

  • verified exact35
  • verified fuzzy36
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9e924d67-296f-4bc0-97c4-95c6e387773f · outbound

This paper cites GPT-4 Technical Report.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.240459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:853a8be05eca42ef1019c95c8422e49fcee279f0ef5657ee35e693240a261c63

Observation ca102fe4-0294-4bc1-9fbe-3246bf0dcebe · outbound

This paper cites Localizing mo- ments in video with natural language.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Localizing mo- ments in video with natural language

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.929711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:78a39f4e00086af01c465fc129ad503ce538570eba677cb89dd73aa91ec7ac85

Observation c2da755f-03bf-48c4-8594-4bb41a80e636 · outbound

This paper cites Qwen2.5-VL Technical Report.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.236564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:6014e825063183f1f00187945802253e121e071d0b02b524f5cdc33e50420a08

Observation 2d14289e-c51c-4387-99ec-71096cad233b · outbound

This paper cites UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.266369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:592ba033ce961b39db83fadd9ce7aec22b7ce4909e492b99f223f9f10a981b6e

Observation 23aaa48a-bfdf-4790-98c1-42b278245487 · outbound

This paper cites Dense events grounding in video.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Dense events grounding in video

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.927612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:f222ee71a4f69a41cef4dbfc9002df2916fa4b41ca7f1fb7a47d96ceb34c55ea

Observation c6fec625-18a3-4d76-9307-8443d9984044 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Activitynet: A large-scale video benchmark for human activity understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.925448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:fc1c44209f5363e1a2cd470a9c571dc5bffbca869971d3686f0e98f11239bb70

Observation 6a65651b-f8ad-4a84-80c0-e377032b98ea · outbound

This paper cites Flashvtg: Feature layering and adaptive score handling network for video temporal grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Flashvtg: Feature layering and adaptive score handling network for video temporal grounding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.923441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:10e4f7eac2cb615ecd4587644b1c7dfb9637d3ff6f8a6b322589618fc8e9b410

Observation cd5c16b1-0bb6-4f72-b4ca-8e53382e30c7 · outbound

This paper cites Scaling rl to long videos.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Scaling rl to long videos

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.378019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:ded3fe7948d4eaf22120a7e4237f65766479988f789edb85f42f87b770756767

Observation 42e182a2-3342-41fe-9686-0d1cf0474f6d · outbound

This paper cites VisRL: Intention-Driven Visual Perception via Reinforced Reasoning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.374154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:594c3510ca4424201d249302f5a86a54c43318fa0f32426f88a35b5ea2d2e3f5

Observation 1b69e8a0-90a9-4dce-8cd0-419b334bec29 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.279116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:e8b1ce625b2f46615e3cc0f27174fe2716d375e4d0b014ce5158b7aa6189c7a6

Observation ebccfd11-141e-4040-af8c-0d4fabdf6973 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.921377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:5e89ccc1abecac5dd623472b15654e3094de018f1103925e39f1ae1b15bed99e

Observation 68b7ba79-714e-4aeb-92c4-ee1a6901d77f · outbound

This paper cites Tall: Temporal activity localization via language query.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Tall: Temporal activity localization via language query

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.919212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:d78f1820410e1fe1745f89ad4c9bade0c9fccd2fae8a97e99478e8b218a308b0

Observation ecea151b-2de8-4a09-91a9-e993788542b7 · outbound

This paper cites TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:17:07.011358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:34795c87c010cb778234987407a2e8f17a9a3cda6c7d753c84dd3e1fdf8b321b

Observation 520d7b46-2d7f-4254-a6a2-34761e1c5191 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.257349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:3fb0fc69f23d02f2b40a0fd7adce9f5be067aaebf6245012d5f41208d84fe107

Observation 169e689b-ea81-47b2-ae35-2c795d170f3c · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.385855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:173ffe63641b57f236a997e2ad331628839bf05ce1b35527a3c000afd35b7465

Observation c545c391-4de8-4bf9-9af6-5814e52614dd · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Vtimellm: Empower llm to grasp video moments

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.917145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:a53a68d93febe4ac8710be7299ad34e6eed31c35f84c73a6f213b9180d240214

Observation a01b7b66-7d29-4ad0-94ec-e21dfb402f7b · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.289328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:8ec63248c284cabfe5f1890504d69f33679d3698003c749e1ec2321c020be71e

Observation 9807d542-83f1-4aec-9b6c-a81ac8618563 · outbound

This paper cites Online Video Understanding: OVBench and VideoChat-Online.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Online Video Understanding: OVBench and VideoChat-Online

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.274928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:2e043d6f8f957e09af309bbb341c6dc1a1da50f480d86176530d3472c65d187b

Observation 4623874d-a6f0-4582-b3d5-c36b164e0407 · outbound

This paper cites in the wild.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning in the wild

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.914937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:ef4f63c47271f5c1e10da91605f4478bac93546e44a651c8198883fb3bccbb92

Observation f2974a46-ed9a-4530-a305-09ee1de41461 · outbound

This paper cites OpenAI o1 System Card.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning OpenAI o1 System Card

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.365201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:f8b229173bf9aeb0d1e9d341c693ac03f6bccf2664372457eaf7b4e0043a2f6c

Observation a1bfd47c-1aea-4da4-a8d9-9ffb80f8ef83 · outbound

This paper cites Knowing where to focus: Event-aware transformer for video grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Knowing where to focus: Event-aware transformer for video grounding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.912742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:6b0cd959f24ac1d23eac4b4231a60cb04e18e189381cbcd4fb2dafc610eca7fd

Observation de5e88f1-43a6-4c72-9611-8fbe04d39583 · outbound

This paper cites Dense-captioning events in videos.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Dense-captioning events in videos

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.910580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:41410347fac32d73ef9d793d3ba2ad4d0f264db64a80b33426cf9526bef40f01

Observation 6c2f1939-f828-48b5-941a-67ca5bde11a3 · outbound

This paper cites Detecting mo- ments and highlights in videos via natural language queries.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Detecting mo- ments and highlights in videos via natural language queries

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.907868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:aa3e1a3a9b2da1410c4c3d1b5b47ba3e539ae655f828623264197bb5fc92a876

Observation c12dd039-c2f9-4ce8-818e-28bf6ffea39e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.294965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:c51f96c49d2a7312c506663b1feeddf379de055e60bed59658b58f2ca6f01d2f

Observation b52f970c-bf0a-4400-9c52-e9db42d95dfe · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Unmasked teacher: Towards training-efficient video foundation models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.905506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:d63ddd90e9af286d4b883dd0598f903e861e235e08900cb1155ed1b3ae5b6b2f

Observation 01f84cc9-2158-465c-85cd-b87dd00ccf88 · outbound

This paper cites Mo- mentdiff: Generative video moment retrieval from random to real.Advances in neural information processing systems, 36:65948–65966.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Mo- mentdiff: Generative video moment retrieval from random to real.Advances in neural information processing systems, 36:65948–65966

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.902868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:0b802025cb2d6caf19145c71d6a9ea4cc5ea4a58374240d5a60a7d604ec9f294

Observation 272ecf85-534d-471f-9d58-5b4347c00ece · outbound

This paper cites VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.299074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:5e9fe5d11ed61a01d0b122d63afe1e09bf6d71e0892179e0125e729409188842

Observation 93ac2896-e900-4888-a00a-4a25f5a5458d · outbound

This paper cites Ground- inggpt: Language enhanced multi-modal grounding model.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Ground- inggpt: Language enhanced multi-modal grounding model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.900629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:5b937f9fbd118512d1357959514d4895fee1e9c8eac08ad17aedfe729565ddf3

Observation d7c17fae-3323-4fa8-bc27-438b4c46233a · outbound

This paper cites Video-llava: Learning united visual repre- sentation by alignment before projection.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Video-llava: Learning united visual repre- sentation by alignment before projection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.841392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:605bfd0ad6812b48e00f69dcbddc28ac33a6463276b8f02d7b9d61f8901ddf2b

Observation e56712fc-fcc8-4a34-96fa-f3204fd3b467 · outbound

This paper cites Univtg: Towards unified video- language temporal grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Univtg: Towards unified video- language temporal grounding

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.898220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:5e633b4c6ba0c7fc66ae517093ac527295c0ada792f6275ee4702f91bb8f2a1e

Observation 681f3c37-9bc1-4d10-bcc4-050b33aa5877 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.896063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:e266162e836885a7c8845731f75e7026589b33a821ce3558058e26a3244a941b

Observation b8c9c9b0-1b9b-45fe-920c-371565cd023a · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.893610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:e4de657e3d8ad85a3a6637ca64570af8bcc5d83bd9c333c0e870b17a2aceb463

Observation 672cd8b6-56a2-4a32-a61b-78fd726cba21 · outbound

This paper cites r 2-tuning: Ef- ficient image-to-video transfer learning for video temporal grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning r 2-tuning: Ef- ficient image-to-video transfer learning for video temporal grounding

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.890767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:e8aeba6bc5131e52af1fa2a6524abe3183b10f84448b183cdb177f63899f47fe

Observation 28b801d3-bcaa-4514-b818-0c7cdf0e3eab · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TempCompass: Do Video LLMs Really Understand Videos?

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:17.144047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:cf0250507627372b249b2f6d0689b9c1007df471b490c7fa5c47b7d8673c7338

Observation 442ec364-5fa3-4d06-918a-45a016735ab8 · outbound

This paper cites an unresolved cited work.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-17T02:18:52.888371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:5bc49c38daf49611839118609a76c455ffdd9cab19ed5256f58bdd2239174118

Observation 4f370744-bb59-4f16-9eb1-72d9333d5167 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.360675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:8f2cfe97846ec78a527a10e371c205eb937168b22d7520e6feb1a6dd984217c1

Observation a810bf84-4b8e-4662-a534-5e19cae1e3c7 · outbound

This paper cites MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.382073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:9d30587675751fad167b4401194a08740df412a77c1acd07a8b80be6903a231a

Observation ce32cea4-58d5-429f-acab-add86dd88f39 · outbound

This paper cites Correlation-guided query-dependency calibration in video representation learning for temporal grounding.CoRR.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Correlation-guided query-dependency calibration in video representation learning for temporal grounding.CoRR

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.885789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:9db31673875e6cbc5af5b1c2b7dac3780ab12484ccc1ae90dac5d544f0844cdd

Observation a2f9eeaa-84bd-46d6-ac96-24c143f27db4 · outbound

This paper cites Query-dependent video representa- tion for moment retrieval and highlight detection.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Query-dependent video representa- tion for moment retrieval and highlight detection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.883621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:f26fbb39906afd2aacd92a262c2a81be22ff34fd7ba4fedcc760124906a6bde3

Observation 9dc16043-078a-4a3c-8d20-3d1239b663b5 · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.244560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:9374455e402645942a978f42040b0fe6ad58c5095c522979a344bc934e421c45

Observation 1a5daa42-d2f4-4fb3-af9d-0cf2f6326335 · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Sys- tems, 36:42748–42761.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Sys- tems, 36:42748–42761

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.881302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:035cec66aa651376b497db5c27387e088296a7f4e79a28b384d2bcb7e0c7fe39

Observation 11d2dfa6-32f8-47de-a608-34136f45a500 · outbound

This paper cites Chatvtg: Video temporal grounding via chat with video dialogue large language models.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Chatvtg: Video temporal grounding via chat with video dialogue large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.878949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:d27b77da2dc80f7b2e744121d1b0207cee51fd67a908f3564688c56d1eea1879

Observation 59f2e683-7596-40d4-8c8c-769b76484e69 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Learning transferable visual models from natural language supervi- sion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.876556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:2c43f6fa85039112351d2f9b0958382b86b57b207ef10d8b74ed8aa4de7b4313

Observation 6dbe3818-759c-4abe-ae7f-eb186a401164 · outbound

This paper cites Timechat: A time-sensitive multimodal large lan- guage model for long video understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Timechat: A time-sensitive multimodal large lan- guage model for long video understanding

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.873515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:476b440d0c49b9020beb366e785968dba69340a0fb039e724aa5e7eb4e26a567

Observation 4f603076-e551-4285-867f-d759b3d35f27 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.248653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:e9edb4dabf4922f8a0e38477c057dd700c426cdb2ba21615c33123d2a3cdd838

Observation 6d828343-faca-4796-998e-a3172586f1a0 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.253010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:dc77a50794d706a7f6b0d9be1b5f13d0cc9d44963d34ff5be016363bae58a962

Observation e4818a9e-30d2-4592-9db8-3dd509e5649c · outbound

This paper cites End-to-end dense video grounding via parallel regression.Computer Vi- sion and Image Understanding, 242:103980.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning End-to-end dense video grounding via parallel regression.Computer Vi- sion and Image Understanding, 242:103980

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.870909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:9aa9b0228355f4adfa504c60ff9d5ae708091943ff55ca80007b932ae33fe6f1

Observation 1459d225-e80c-4a46-bab7-4e39c11f7548 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Moviechat: From dense token to sparse memory for long video understanding

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.868466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:49ece88654a78c50072fdfd1f28b4c0c731b08c8858da7afd90590d097100385

Observation 7e7cee7c-7670-4ba0-9e21-e339cd5aec3f · outbound

This paper cites Tr- detr: Task-reciprocal transformer for joint moment retrieval and highlight detection.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Tr- detr: Task-reciprocal transformer for joint moment retrieval and highlight detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.865957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:c4f452bb56225f70442d139a2a064399f6fafb40fb2398151824d488f087ecd7

Observation f233cbca-53fe-4135-bf8b-a2ad4156c04e · outbound

This paper cites Hierarchical semantic correspondence net- works for video paragraph grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Hierarchical semantic correspondence net- works for video paragraph grounding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.863440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:1e966153d2425fd325f5235811e214af3b090d530fcd9b516063b2de769ff329

Observation 15dd6e88-c57e-4c4c-837a-008464009a0f · outbound

This paper cites Hierarchical semantic correspondence net- works for video paragraph grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Hierarchical semantic correspondence net- works for video paragraph grounding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.860479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:2f2f31b0b88e15768c4e8ae30216aae75485615d8750be9b412e4d9fd9ad2a19

Observation dc84d14d-1295-446c-b300-38fb62d65e0f · outbound

This paper cites Tspo: Temporal sampling policy optimization for long- form video language understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Tspo: Temporal sampling policy optimization for long- form video language understanding

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.356084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:897c573be588a64819093dc08fa088bae759708f9f91c1112cf7b8ecdfc9ef23

Observation 862cdbb2-b580-43cb-b6df-242a8ec5a795 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.302760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:6e632b773b7984c6770f8463b29071a6b6ff46d3ae1746cd99640026364499e7

Observation c6333379-37db-4fd4-85b4-7502732d183e · outbound

This paper cites Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Videorft: Incentivizing video reasoning capability in mllms via reinforced fine-tuning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.351204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:e9aee1a05b1d11e6ec27b453e0b607d90cb683af4772931dcf8298ce7d42ffe0

Observation a8c5d2f2-9c0e-49b3-b0e0-b42e0379d4fd · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.306944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:ed178a38e5eb605c593b43b36ea1287a37d3f74fe91f6310404e09f5253ed0b8

Observation f8404e5c-5d49-4305-b484-126752c65f39 · outbound

This paper cites Internvideo2: Scaling foundation models for mul- timodal video understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Internvideo2: Scaling foundation models for mul- timodal video understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.857516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:d94055bc9674cbbc3174a933a4267bf7b7d92ca9a89cf1bb2feedeaf3755eb65

Observation c2633b38-8492-4fae-b0b6-62c18d264932 · outbound

This paper cites HawkEye: Training Video-Text LLMs for Grounding Text in Videos.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning HawkEye: Training Video-Text LLMs for Grounding Text in Videos

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.343793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:ef68af01eeefb4a91e41e863fc24c8da2ed6bf66aabff98f1fd4b3914ea33548

Observation 5de38293-abb8-4a48-8d1d-41bbe9413f47 · outbound

This paper cites Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.339400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:9f0fb8b0e436fb354c3684d48cc72b84396522ca515efbe38e2274417cf8ffed

Observation fb327704-a237-408e-84ee-7a14cf52e717 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:40:06.763282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:420a834656cecfaee396a3dfec486e021cc1940c74f8fa77003e68c3b2134438

Observation 1a187403-cbb5-4c2f-8eba-367d2645778b · outbound

This paper cites Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Visionary-r1: Mitigating shortcuts in visual reasoning with reinforcement learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.284171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:26e1682eca57b93f4282a79f1e46bf9886e74dc2b9c03b442cef684e29b21c7c

Observation 6a123d00-0d84-4dba-92cd-0ef2c475cd52 · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Can i trust your answer? visually grounded video question answering

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.854622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:b8c05b64c36b6f583f71fd42dbc5348c39aa1fc9ae9155a84ad3980856504b0b

Observation 1b539996-056c-4907-9e4c-50a5e54ba3c3 · outbound

This paper cites Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Bridging the gap: A unified video comprehension framework for mo- ment retrieval and highlight detection

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.851888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:bea6aac4dcc57de6794427827dd0314e2fc9a489415ad7313e186fe737c11592

Observation a30d6193-0b77-4a33-8a28-fdcf02fa13f9 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.335434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:17180dd1f74966cc5b694b30567255c30c0367a78670275257bde61036ddf69c

Observation dbb608f1-cb63-4ed3-ad90-f84c4bacb0c0 · outbound

This paper cites Videochat-r1.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Videochat-r1

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.314177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:aa80b2e24d018a1ededcda91ba4a645301c090ae9c7dbc8000a9447de0886942

Observation 067c528d-c020-458e-8a3a-3237b3e1d735 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.328163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:b7f2beca1ddcdefafed174267c62b57b9a3d27d730b922aab2f600af7e68a6a1

Observation e0fb155f-e300-4ee5-a867-e80b8a875f35 · outbound

This paper cites TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.310729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:53dfe4330b588939e128058aa0a63fdae90b818926a217ba6ff7bdf0d7506ede

Observation 3b27a866-be10-47e8-90f6-c954341edf38 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.324847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:d3f0f6eb4426f41c215cb89220fc59117306fcc39839a597767bab20a8226ce1

Observation be23abb0-5110-4e32-841d-bccc4487b546 · outbound

This paper cites Sc-captioner: Improving image captioning with self- correction by reinforcement learning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Sc-captioner: Improving image captioning with self- correction by reinforcement learning

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.849268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:9014cbd7ce0703720c4f9b627973e5b9e430d74076ee29e347623187d8a8a816

Observation 7e940b54-6378-4f98-b9df-28dc33c57448 · outbound

This paper cites TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:18:52.321509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:6838ab2a6b2cdbf998a4f96fc0c1b099339c474860e220a443f2df0b905dd587

Observation 6ff04040-ff6f-49d2-9d95-6d3bd7fc1f47 · outbound

This paper cites Hacs: Human action clips and segments dataset for recognition and temporal localization.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Hacs: Human action clips and segments dataset for recognition and temporal localization

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.846707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:bf36c83da3f10abf21cb5279f52e204fe1fad9c006684fc9d5021718589d70b2

Observation b5e099b8-c0de-49f0-bddb-704e8764274b · outbound

This paper cites Group Sequence Policy Optimization.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Group Sequence Policy Optimization

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:18:52.317842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:53d9c43c2735aaf26e39b4719091f02105ac952b4f915bdd48de21bd4e1b778d

Observation 72669cb1-ea7a-4f24-8e64-5a80eb90dc0f · outbound

This paper cites Rethinking the video sampling and reasoning strategies for temporal sentence grounding.

TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning Rethinking the video sampling and reasoning strategies for temporal sentence grounding

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T02:18:52.843915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T02:18:21.718091Z digest=sha256:672cc9d8304b44851013dc36ec2ce87e5f8d7efba39f08fe079fc76261afe279

Pith citing papers

Observation 38fb59fc-0b9b-4067-9d2c-c447183f7ddc · inbound

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks cites this paper.

Multi-Task GRPO: Reliable LLM Reasoning Across Tasks TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T04:17:27.916261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:17:27.916261Z digest=sha256:d8b4790d3dc04b5f6e492975ffd8918d8114e687ed1b29d8a45fba57cd77bec7

Observation ce54db9b-5438-4540-9a1b-cbffd719bea6 · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:256c76515bd62a4aa1c7f2f74565d82363c2c2e14ac7442c136fee0ee13f365c

Observation 31e83b57-04d5-45c2-93b1-b4bffa41c296 · inbound

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs cites this paper.

TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs TempR1: Improving Temporal Understanding of MLLMs via Temporal-Aware Multi-Task Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:16.823032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:04:16.823032Z digest=sha256:b7d4f7f107eb94d03ec27f2f5fcfb5458236a195ca4a56676a9b4396550ae8a1