Pith. sign in

Paper Citation Record · LEDGER

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 3 inbound Pith citation observations for arXiv:2505.15447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15447 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:59.938701Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:28:56.047745Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T05:56:07.964144Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf12f22b-1d92-4016-afcf-f6ebf19bc729 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.425325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.425325Z digest=sha256:ff22a59c6935173a52b9e183897e7f46013958cdcf58250bc188e605737b5c18

Observation 27ac58e4-903c-438f-8d0d-c130fca152aa · outbound

This paper cites Qwen2.5-VL Technical Report.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.475604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.475604Z digest=sha256:169c0f7cf0a1a631a8e5accc6d1ff6c78aaec6bc6f0f402224f7aa89cae0f466

Observation 342666bf-c2fd-4629-ad13-108355f1159c · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472– 19495, 2024.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472– 19495, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.552758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.552758Z digest=sha256:cd6808d1429e836866fa3950fa5f77a47d52a7227a748cc0f26cb6880a6aab16

Observation 6bd6c3fb-fa54-444c-8940-686612390029 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101, 2024.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.612041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.612041Z digest=sha256:6ef324c81a902f286bab83d26331c3dd7e3d4e7b365b7d089fb9572cd7446285

Observation 7df3f253-5ae7-484a-8a85-caa50585f1e0 · outbound

This paper cites Supervising strong learners by amplifying weak experts.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Supervising strong learners by amplifying weak experts

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.700519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.700519Z digest=sha256:6496241b185ac3232becae882d493f22432130328b87b1868f361e0d763affdf

Observation 4c265e10-9dcf-4610-9ba4-825985f4bb4d · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.775593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.775593Z digest=sha256:35f6920889ad0a2918bc28ab98a866b1f6154b5b5d0e6c6147da4e3beac3da86

Observation 71f78dde-a786-407f-a9c0-4cac870c7a37 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.873824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.873824Z digest=sha256:2ffe2228eedfcbf2966e03fda011a327f62572bb6dedb763bf04aef8fdaa53b6

Observation dd2ede75-5c5b-4b78-bb51-b4595d376349 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:56.957053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:56.957053Z digest=sha256:fbde16ee10719a7f0a081de94432fcb165d71c21883cc0ee549840a63c6ff6ab

Observation f85709bc-2a0e-4bd8-bbe6-daaf4d5c108b · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.028711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.028711Z digest=sha256:9777a8990f0473ae84e9429554e405e09c3424129913b14267ed6d5c29c9892b

Observation db2ab293-089b-4639-b534-f9246a05f5a8 · outbound

This paper cites CoS: Chain-of-Shot Prompting for Long Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning CoS: Chain-of-Shot Prompting for Long Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.091933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.091933Z digest=sha256:67896943dc6147c09f364371cfae1e76bad4eb109b035ae47acba1111e6ca0f9

Observation 0e53acd5-2aaa-4d8f-9a02-327391ec0747 · outbound

This paper cites M-LLM Based Video Frame Selection for Efficient Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning M-LLM Based Video Frame Selection for Efficient Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.174134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.174134Z digest=sha256:f0aa65c7ae92ac6971a87bc7f9b249723ed97834ad58baf2dfb023c986a06a23

Observation 113e3bc5-eadf-4c37-92bb-29c545e23c57 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.242364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.242364Z digest=sha256:6de3babc241388995ea94f1580048a9e2aedef59318b4809836817746ef4e58a

Observation 0cf1332b-7be7-4177-b5e0-de0665d5f8ad · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.315079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.315079Z digest=sha256:5497df8745be9d837058cfdcf1d149d8c7104dd68ae6ff8332c0ae69b3d706b3

Observation 2e2e35f4-2831-4741-985a-97dad4e1fcc7 · outbound

This paper cites ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning ReMax: A Simple, Effective, and Efficient Reinforcement Learning Method for Aligning Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.409103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.409103Z digest=sha256:c49174fbd390dfdd0bf76ca3e145c604b7924e4711bafed57c41185cd9ee297f

Observation da9d7c4a-a09d-46f1-aee2-ea29148cb717 · outbound

This paper cites KeyVideoLLM: Towards Large-scale Video Keyframe Selection.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning KeyVideoLLM: Towards Large-scale Video Keyframe Selection

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.485440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.485440Z digest=sha256:f7676d643884a9c4e54df0f122801ed46d2edd399bef14f768e68f5f5fb5ff97

Observation 720ecb11-afd2-4c2e-a1bb-ad1569845d9f · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.562364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.562364Z digest=sha256:9b8a564acb6487495ca9eb463d76db6eaa46274c3dc5b550295310d05e6ebdc4

Observation 399c84bd-a020-40a0-a430-6bdb27af0921 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.634142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.634142Z digest=sha256:4c2c52e61f2fa102b5b74b09eaa28c918cad9449ee22500828ece35c08534298

Observation 8335557e-cb75-467b-b018-6d8712a4a5b0 · outbound

This paper cites Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.694251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.694251Z digest=sha256:a0c74699be88e359e25bd79551ac934d8167d21c9f5ad9af6402af545350f1db

Observation 3a699b2b-43ba-4426-a63a-3c27efa0be05 · outbound

This paper cites Hello gpt-4o.https://openai.com/index/hello-gpt-4o/, 2024.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Hello gpt-4o.https://openai.com/index/hello-gpt-4o/, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.763628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.763628Z digest=sha256:fc184b7f2582e0ff28358c8f0a5c9c293589e406b22fc33c1faffe23a69d3e37

Observation c66a802e-73d2-4f05-8583-4eb95c48834c · outbound

This paper cites Ai 2027.https://ai-2027.com/, 2025.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Ai 2027.https://ai-2027.com/, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:00.704398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:20:57.848572Z digest=sha256:f86b1bfb8cbe1fb8fb7556ed4490c4bcc4ca7f5fd8a90819922bbc88e0b9962a

Observation 74c24c73-18f8-4c77-a611-7f3807c6e079 · outbound

This paper cites Introducing openai o3 and o4-mini.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Introducing openai o3 and o4-mini

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:57.917067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:57.917067Z digest=sha256:d8f568522371cab4084d606a303e100ab7e0e04c22768babe79b2500265cb940

Observation fbfa3356-28f4-4ebe-b8a7-afdc691d7104 · outbound

This paper cites Learning transferable visual models from natural language supervision.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Learning transferable visual models from natural language supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.009144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.009144Z digest=sha256:ed38fcf37bd7938829c8131a9e45d4ada047744b5c7dee147cc7fe5d36c7bda4

Observation 436ddb25-59ae-4f29-aed5-494b4d635f43 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.095416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.095416Z digest=sha256:2f1659a154889ee957b361f1bf54d11d4a1c184511c7e87fa2bef11dc86ff913

Observation d1e8c877-96e5-4be1-834f-8bdbff7da2c3 · outbound

This paper cites Trust region policy optimization.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Trust region policy optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.248304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.248304Z digest=sha256:0ec2cd136e95d37391c07d7b9d59fe75bdbd8e662165a75af0cac328f8990698

Observation 9115f05f-a2f0-4b8b-9f4e-6fe819ccecad · outbound

This paper cites Proximal Policy Optimization Algorithms.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.362338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.362338Z digest=sha256:ed5e919f8621298e322b4dd0b56f79a798c261ba90884568967ecf255f69a1c9

Observation 87adb3c3-3276-4d74-b086-2120528e9b70 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.456192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.456192Z digest=sha256:af4f3307dc842148671b3fc76267b8c4d5841559f3185f51e8a82d9281c0b691

Observation 8d9c289e-2d5a-4082-a8ec-fcc72a4a1192 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.584507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.584507Z digest=sha256:351dbd8fc4df8f7708851a8d2ae324fff921b829c970dd3a7d58c6a8189331d1

Observation dbe22634-eae3-4339-8748-7cdf90877cd4 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Moviechat: From dense token to sparse memory for long video understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.676634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.676634Z digest=sha256:1f7ebdc49e8fca684a4ce8e56f28287cb069d164e160c6c379d68e2234e198e2

Observation a5fc3f40-2601-4cb8-a4ee-395a76b206cf · outbound

This paper cites Adaptive Keyframe Sampling for Long Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Adaptive Keyframe Sampling for Long Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.751103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.751103Z digest=sha256:8f465f3efdbe45d2d37a89e63ec23d32e7d54e007e74ad823f43ec3e3eb95152

Observation 65c61772-c52c-4bbc-9c30-2dd5b0872121 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.826451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.826451Z digest=sha256:a260118e6a07fb06c00886331f034330e689b13f883e62b49fe8ad503013fc5d

Observation be03e616-7c36-4067-ba95-416a5bc8a359 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LVBench: An Extreme Long Video Understanding Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:58.975786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:58.975786Z digest=sha256:eeb423124a30ce4306d39dd4d523be5a254efdc68fc91dfd352365f0e150df0e

Observation 8ffb3ee0-50c0-4ea5-affe-48f2ee8e725d · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.072257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.072257Z digest=sha256:bc3c3f83b1efb752bfdd913e1584d3b020161d4c08d4b4620f2d4d5a13201b27

Observation 1445af28-cabc-47dc-8c25-d4d3be3d296f · outbound

This paper cites LongViTU: Instruction Tuning for Long-Form Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LongViTU: Instruction Tuning for Long-Form Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.162625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.162625Z digest=sha256:e985c32dc3fa8eb720451bcb34eb99e73df3abae2d2b01ab6058cad61b076a6a

Observation 4e3c51c5-4ff5-45b9-bf7d-a5636c50b639 · outbound

This paper cites Number it: Temporal Grounding Videos like Flipping Manga.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Number it: Temporal Grounding Videos like Flipping Manga

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.262348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.262348Z digest=sha256:898642a91840392870a3d56d674bc02475c8f02d2fa8ff0dac99e94a93a88c40

Observation 7a6ad6ee-18f4-468e-92b1-152590b8d05f · outbound

This paper cites Frame-Voyager: Learning to Query Frames for Video Large Language Models.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Frame-Voyager: Learning to Query Frames for Video Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.359347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.359347Z digest=sha256:17a86d11a2548235371e388d72d2173d12242dda8baee09e5af9a8214b6d351f

Observation 8961aedc-0025-46cf-a0aa-5d0cca75b54d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.427397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.427397Z digest=sha256:717cb5f1d16b25cf8ef129b5172154359e485c579a71788e224c48e9013d3528

Observation 686d382c-f69d-4325-9780-7c48d5248f4d · outbound

This paper cites Long Context Transfer from Language to Vision.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Long Context Transfer from Language to Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.552615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.552615Z digest=sha256:60236eaa030bac02af70b8ff030e3119b2a205e1d74bd300041faf28baae0ebb

Observation 75ba5317-b5ea-4f21-b12a-517e7916add4 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.633738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.633738Z digest=sha256:d37195a82a41cb584944881847f8ab5baf42cdccbc7ecd7ad434c2aab156458a

Observation 2855bbcd-de4e-45a1-b7e1-4e80a065abd8 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning MLVU: Benchmarking Multi-task Long Video Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:20:59.712696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:20:59.712696Z digest=sha256:25737fe6836eee246220274fc08a36a1e540fefd19554320e6d821b1588a4c34

Observation 28a12429-787f-48e2-805e-b199c7f94349 · outbound

This paper cites - Check if the occurrence time is mentioned.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning - Check if the occurrence time is mentioned

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:00.510642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:20:59.765758Z digest=sha256:96225f3385e49dfd74c8c1c0c0c42453e0f80146dbf86c08bf5769840d8059c6

Observation 4940706e-4e41-4fd1-92a3-f8c3efbb7516 · outbound

This paper cites an unresolved cited work.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:21:00.388794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:20:59.845090Z digest=sha256:eb8b2b1e2e65cfe6c020afe8e4e0887f3d4eeb4d4ca7fa6b7c978701a994ff50

Observation 3571a2f4-bc9c-4d10-a20a-9d55c9d3ba98 · outbound

This paper cites mRNA" and.

ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning mRNA" and

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:21:00.261706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:20:59.938701Z digest=sha256:5bf3d822964bf03d46994bdd320a5678f441cc1672d812de3706cb74c27d771e

Pith citing papers

Observation 8527e4c0-d110-4e6d-818e-d6e4fa602931 · inbound

Phi-Ground Tech Report: Advancing Perception in GUI Grounding cites this paper.

Phi-Ground Tech Report: Advancing Perception in GUI Grounding ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T10:28:56.047745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:28:56.047745Z digest=sha256:37d0dfdebd31bca45dffe4eaa10b2e11c3471f0c8c2070151186548db02caaec

Observation 125702e3-0b83-442c-b5ec-d8ac8b39a4d9 · inbound

Swift Sampling: Selecting Temporal Surprises via Taylor Series cites this paper.

Swift Sampling: Selecting Temporal Surprises via Taylor Series ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-22T05:56:07.967892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T05:55:23.479344Z digest=sha256:3d5acac276395524f7a85c14c448a6a8561aeca304fc0940263237eb2fe0e895

Observation 3638ec97-6a7f-47c4-81fd-d04f151aeebf · inbound

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors cites this paper.

Efficient Frame Selection for Long Videos at Test Time with Attention-Based MLLM Selectors ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T22:38:43.510039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T22:38:43.510039Z digest=sha256:7df0609b737eeb80839db9cfcf21ca28f1b33c61a0e5a0ef695d73bb560b95e4