Pith. sign in

Paper Citation Record · LEDGER

Kwai Keye-VL 1.5 Technical Report

As of 14 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 31 inbound Pith citation observations for arXiv:2509.01563.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01563 v3

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:28:32.039808Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T20:43:52.925722Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 96155cd7-0c6c-4024-82aa-e5c64c3c9243 · outbound

This paper cites The Llama 3 Herd of Models.

Kwai Keye-VL 1.5 Technical Report The Llama 3 Herd of Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:24.911132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:24.911132Z digest=sha256:4a611fd08dd4ff097e5807bfbbd94079e7181fe7b3a557121fff70ac15a0ce8a

Observation 54891770-a499-44b3-9f06-783853d26b34 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Kwai Keye-VL 1.5 Technical Report Emu3: Next-Token Prediction is All You Need

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:25.288692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:25.288692Z digest=sha256:82fe4bd7da94fbc7ea939d5d0fc22b7c42e3a2eb8119020266d3cef0b0dcfc12

Observation d0ad010b-53b9-48b2-bc88-156d9c6f71aa · outbound

This paper cites an unresolved cited work.

Kwai Keye-VL 1.5 Technical Report Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:25.600786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:25.600786Z digest=sha256:1009a3ef2f6fd7c4128cc58eb9228435eff7e265815bf517856ac86f43f90e4b

Observation dd942ec4-2f9a-403f-8681-a1c98494e426 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Kwai Keye-VL 1.5 Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:25.798334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:25.798334Z digest=sha256:52916d5f9803050ee3f1834054fecdb30516aae45dcf8dbbf3fc999e7def0c34

Observation 8fb37b12-9d52-4f8e-bb8a-47334319bc4f · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Kwai Keye-VL 1.5 Technical Report Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:25.940572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:25.940572Z digest=sha256:4a57aa99456c5ac1eab0c461977b70fb83c3fa1ca5e4f0043242e120add86cf4

Observation 542b9c3b-94ea-4e45-8c62-8379f0fe64b0 · outbound

This paper cites Kimi-VL Technical Report.

Kwai Keye-VL 1.5 Technical Report Kimi-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:26.070670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:26.070670Z digest=sha256:0702967a08d870d0a900eb65075cbb9ae160adb95519f7cba5365d981e4fd465

Observation c2adc5cd-ff45-4554-9c9a-f5ecbf6f407d · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Kwai Keye-VL 1.5 Technical Report VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:26.227633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:26.227633Z digest=sha256:c1de79b792339420db4f02d8248598988228d3072c96b08e5dfde5f9c658d3c8

Observation 9af9e927-2ea1-4b3b-b5c2-312c6ead1120 · outbound

This paper cites RAIN: Your Language Models Can Align Themselves without Finetuning.

Kwai Keye-VL 1.5 Technical Report RAIN: Your Language Models Can Align Themselves without Finetuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:26.384522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:26.384522Z digest=sha256:bf114f4d3892758dbecd5f9b50aaf028040d89875e90ea88c99a0d7ce514f378

Observation 4a4d3bcf-c754-407e-90e5-db416ee141c3 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Kwai Keye-VL 1.5 Technical Report Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:26.680272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:26.680272Z digest=sha256:e273839f90c9eed669ecdecdf2bd80a132b6a1bfeb8e9815e4696fd6112f3928

Observation 7ca2ec95-3991-4fe6-8abb-0740c4141272 · outbound

This paper cites DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World.

Kwai Keye-VL 1.5 Technical Report DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:26.818295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:26.818295Z digest=sha256:1532d18702a7debc4c39ee32886cd8c58136039847ee1144c911976331af16e5

Observation 5d978ab4-d259-45df-8b6f-e2b7bfb53809 · outbound

This paper cites MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning.

Kwai Keye-VL 1.5 Technical Report MLLM-Selector: Necessity and Diversity-driven High-Value Data Selection for Enhanced Visual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:26.968895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:26.968895Z digest=sha256:86e711968733847b3374aa05082f5d9e77280bf307bbc80cc7ae5491853a69ad

Observation 87bf0350-663b-4f0c-9efb-884c8ddf2137 · outbound

This paper cites OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning.

Kwai Keye-VL 1.5 Technical Report OpenThinkIMG: Learning to Think with Images via Visual Tool Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:27.117812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:27.117812Z digest=sha256:03c3f8bec265abd29d5faa253b6f4641c12a2f6ff9af9d6f645a2b7bed597b92

Observation b3f71dca-5c40-40b1-b45d-d5f9da73f9e4 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Kwai Keye-VL 1.5 Technical Report Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:27.256289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:27.256289Z digest=sha256:0876b897e6d05bc987bd97b6bd87f3270f384f9cdde65b47176e2e6750ae67fc

Observation 0d1a1c19-d362-4e8a-ad93-90968ce2a127 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Kwai Keye-VL 1.5 Technical Report Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:27.379864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:27.379864Z digest=sha256:b1717195a8302def138041fc11604843e542d503e4e391f292e6a71ecc4aab85

Observation 495ff408-8c7d-4691-8774-7c4a7e2be882 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.

Kwai Keye-VL 1.5 Technical Report Video-rag: Visually-aligned retrieval-augmented long video comprehension

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:27.513986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:27.513986Z digest=sha256:4fcb4cb16d078abed2d3829090f6d98c10427e2741a849bf27afbd2a40ca93e9

Observation a12813f4-2c92-43ff-8589-3e99cc7e1263 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Kwai Keye-VL 1.5 Technical Report MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:27.854438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:27.854438Z digest=sha256:4111f25d1ade49bd7058ec38789890bba98f14cd2158fd078844c1d88ef6a364

Observation 57b1d7c1-320c-41fe-a077-90b401a1b58d · outbound

This paper cites Public Domain 12M: A Highly Aesthetic Image-Text Dataset with Novel Governance Mechanisms.

Kwai Keye-VL 1.5 Technical Report Public Domain 12M: A Highly Aesthetic Image-Text Dataset with Novel Governance Mechanisms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:28.020884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:28.020884Z digest=sha256:3f591326bf0a653b5ca9e6b66b06408ccf4de1a11b79ffcf035b8126400cd77a

Observation 7ba892cb-0d2e-4f90-8022-f7fd4d9b9606 · outbound

This paper cites Microsoft coco: Common objects in context.

Kwai Keye-VL 1.5 Technical Report Microsoft coco: Common objects in context

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:28:33.756086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T12:28:28.159724Z digest=sha256:7ea7f023a39bb2fe7db798f9343c37ea8ead3c5cd4643521bff936288caa51a5

Observation 5672ce2c-f641-4552-ba1f-88334ca51faf · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

Kwai Keye-VL 1.5 Technical Report Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:28.452005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:28.452005Z digest=sha256:98393d7ff663f40f19c741ebb8a0aeda6ea391636c9c8c6301ebab145bd6a38c

Observation 93cc01e1-b660-4105-bb4e-3443d681fd57 · outbound

This paper cites ReferItGame: Referring to objects in photographs of natural scenes.

Kwai Keye-VL 1.5 Technical Report ReferItGame: Referring to objects in photographs of natural scenes

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:28:33.428166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T12:28:28.609733Z digest=sha256:8ae439477317834b0581b264b63bfd171c7c06a7d6c6c6c05efa34df05cfd301

Observation 751aa989-2eba-4ed7-8995-5076ca6ec675 · outbound

This paper cites doi: 10.3115/v1/D14-1086.

Kwai Keye-VL 1.5 Technical Report doi: 10.3115/v1/D14-1086

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:28.751688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:28.751688Z digest=sha256:65af888ec2cad5f17f9eae5ee4cf5df867bfb71626d7e7e85b7e58a2328e5243

Observation f696242a-dd65-4674-8dfc-64c84568f856 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Kwai Keye-VL 1.5 Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:29.027544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:29.027544Z digest=sha256:ee4979b1744fd385e28e83384cec4d170d7013a44f5a9ae94453e2dd682a5757

Observation 80ed4ea1-fdac-4d1f-bec0-38616d15cd1c · outbound

This paper cites TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action.

Kwai Keye-VL 1.5 Technical Report TEMPURA: Temporal Event Masked Prediction and Understanding for Reasoning in Action

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:29.153701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:29.153701Z digest=sha256:de12fc43e6cb0edd671f4a1dbf6b35aea2a0212a96bcea6195e85ce897d4f187

Observation 97dacf39-386f-42ba-af0c-d2a9fb255764 · outbound

This paper cites TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types.

Kwai Keye-VL 1.5 Technical Report TaskGalaxy: Scaling Multi-modal Instruction Fine-tuning with Tens of Thousands Vision Task Types

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:29.290584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:29.290584Z digest=sha256:4151c629a8567d0a366e0c0cb7784c349ecd79301386c44465bf2f85b83fd9bd

Observation 6fde1ffd-faba-4cd4-8962-df6e18b2316f · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Kwai Keye-VL 1.5 Technical Report Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:29.437795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:29.437795Z digest=sha256:92a316a7570bce1b659cd476da1f2cccd9180bc44cc347549eacb6c5701be56e

Observation e2b8380a-fb52-4083-9e45-b8b32077d12d · outbound

This paper cites Chujie Zheng, Shixuan Liu, Mingze Li, Xiong-Hui Chen, Bowen Yu, Chang Gao, Kai Dang, Yuqiong Liu, Rui Men, An Yang, Jingren Zhou, and Junyang Lin.

Kwai Keye-VL 1.5 Technical Report Chujie Zheng, Shixuan Liu, Mingze Li, Xiong-Hui Chen, Bowen Yu, Chang Gao, Kai Dang, Yuqiong Liu, Rui Men, An Yang, Jingren Zhou, and Junyang Lin

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:29.577683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:29.577683Z digest=sha256:d4bf467722838e3bc41f137d7f37fbcf2ea3cb1e533d5f2443d2e1c5c742b1ca

Observation 31760e6e-13f5-4ec2-b006-15d6d3dab6df · outbound

This paper cites Group Sequence Policy Optimization.

Kwai Keye-VL 1.5 Technical Report Group Sequence Policy Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:29.805435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:29.805435Z digest=sha256:74658ff71840e1451daabe9b94b879495129e4927f091f17a937847c67f48324

Observation eae75902-24bb-42eb-aeec-b53e992d336c · outbound

This paper cites A diagram is worth a dozen images.

Kwai Keye-VL 1.5 Technical Report A diagram is worth a dozen images

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:28:33.150695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-05T12:28:30.019944Z digest=sha256:7f93800a111efd6fc99dfa24057efc93294f4b32c2917f10aba00dd2b3189d6f

Observation 13a4d9b5-34ab-4d86-b96c-42772aa0706f · outbound

This paper cites ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models.

Kwai Keye-VL 1.5 Technical Report ZeroBench: An Impossible Visual Benchmark for Contemporary Large Multimodal Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:30.182788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:30.182788Z digest=sha256:ee09170499e22fa6db7fb998f1069629761ac456afd7ddeda7db055bdeb3cfeb

Observation 650cbf96-a3be-4115-b2a4-b1a6341e933c · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

Kwai Keye-VL 1.5 Technical Report VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:30.391028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:30.391028Z digest=sha256:d71ec210c5b60565c338a40526452a03d5599df3e762cbc22092e2dfd839729c

Observation 50d5d9c7-9092-4a04-8155-352b8c3cffe7 · outbound

This paper cites SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models.

Kwai Keye-VL 1.5 Technical Report SimpleVQA: Multimodal Factuality Evaluation for Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:30.556672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:30.556672Z digest=sha256:11ba2e568d3a422d0d7c5282f72dd4900dd8c3396a654da04bc4400022f6233a

Observation 58847d44-65cc-47b0-83c6-4c87767cd765 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

Kwai Keye-VL 1.5 Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:30.702457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:30.702457Z digest=sha256:d6d0afe7e0ae1f854acfeea2122d10b09325d897d138281118c8b66cbd1bba42

Observation a420801f-3fd4-4250-95f7-555073a1599a · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Kwai Keye-VL 1.5 Technical Report MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:30.929404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:30.929404Z digest=sha256:fb6f65b15f885d049f2f092cfa2825b3f174d060ef7bc6b142cfb6bfe96f52f8

Observation d3495f12-f1a7-4bd8-bcb5-7dea6e2e3846 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Kwai Keye-VL 1.5 Technical Report OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:31.083905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:31.083905Z digest=sha256:da8c6f2bdc3f5d355b47e9f77481cff23dd38efbb2311bb573accb541947baab

Observation eb1c9eea-f679-4bcb-9227-7c8d10cc2ad9 · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

Kwai Keye-VL 1.5 Technical Report We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:31.248879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:31.248879Z digest=sha256:c4d7580bfb8947e39f31977369a0bcbc6915db0b735162b8ddc7fbb993d90b85

Observation f331c613-383a-4f77-8aab-df3e5cebe440 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Kwai Keye-VL 1.5 Technical Report LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:31.473440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:31.473440Z digest=sha256:46227bcff31fc12c3e075d22b40a81e787087c3eefd850f98337e40ee0ba2cbe

Observation 2a76cfdb-d8cb-4be2-b25b-a98ce65010b4 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

Kwai Keye-VL 1.5 Technical Report DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:31.683925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:31.683925Z digest=sha256:20cab30338a52ccd003490e981f673c974729efc87556d6a6379747c52c5d6eb

Observation c366044e-c563-4ae0-bc80-11aa6d5e7e6e · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Kwai Keye-VL 1.5 Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:31.824685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:31.824685Z digest=sha256:982546363a77cd68787151707663c9a9e270351a1c187450e709f7b2cf782381

Observation 83f477d4-a0ee-4e46-b838-1ef3770b3e50 · outbound

This paper cites MiMo-VL Technical Report.

Kwai Keye-VL 1.5 Technical Report MiMo-VL Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:32.039808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:32.039808Z digest=sha256:9df691844cdf894d601fa1fd0a40227e194058fd063d314b9739f28bf5c2be23

Observation 67a1da38-4632-40f3-b0b5-40118ad57f57 · outbound

This paper cites Silent Data Corruptions at Scale.

Kwai Keye-VL 1.5 Technical Report Silent Data Corruptions at Scale

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:28.302233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:28.302233Z digest=sha256:35dd925e3606c71784351ad7e3c94e683eddd8b52660ea3f541d0f0795bb4be2

Observation 6f5049e9-a1c7-4e0d-8a87-9c585e5be49e · outbound

This paper cites Toloka Visual Question Answering Benchmark.

Kwai Keye-VL 1.5 Technical Report Toloka Visual Question Answering Benchmark

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:28.888144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:28.888144Z digest=sha256:4de0549c89fbf6f0f26b7933509c6268fe197f41b3d09bbe0d596f8bec0a8269

Observation af8abafb-9638-4e7b-b7ad-0ef1fd8de52c · outbound

This paper cites Seed1.5-VL Technical Report.

Kwai Keye-VL 1.5 Technical Report Seed1.5-VL Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:26.527163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:26.527163Z digest=sha256:3dec2f95cf3f701aaf78a8a6f2c6c78787b014039300268965e9f812858c66c0

Observation f2e98c2b-beef-44a4-a370-49cc3bf06313 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Kwai Keye-VL 1.5 Technical Report Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:25.063652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:25.063652Z digest=sha256:e00ad3130fc4b097b561741653365d5a53bcb02eff68f43f2b1342fb6bc8e52b

Observation 29e69062-42d1-44e4-90bb-5dbe94c70ed8 · outbound

This paper cites Qwen3 Technical Report.

Kwai Keye-VL 1.5 Technical Report Qwen3 Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:25.449452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:25.449452Z digest=sha256:1a0dcadda84154db4255844d359d0b694786c3c14b67ce43d27854097e1e94c6

Pith citing papers

Observation 76ef0aff-0b87-426a-ba9d-0f3b9bf419c5 · inbound

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities cites this paper.

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities Kwai Keye-VL 1.5 Technical Report

Reference 267

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:16:04.811876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T22:16:04.386706Z digest=sha256:defe4ca318bfc52053e3ec1a4f8f7ac52795282bd9c5045ba75e832c414de8a7

Observation 55bb90cd-8b5d-4e9f-b10a-f6b3d1cfef1f · inbound

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters cites this paper.

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters Kwai Keye-VL 1.5 Technical Report

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T14:16:50.869222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:16:50.869222Z digest=sha256:2178a622a1e60b7d483e0c4daf6724a2105a3b80a09a032461b21b9d6dc1eb17

Observation f384bfee-91f9-47ca-a7db-8196c89d3bf8 · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning Kwai Keye-VL 1.5 Technical Report

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.834586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:7b957275fbde09a3544e72c7ec4b554bd796e99fe344bbefd66f09388eeca592

Observation e452badf-0ea2-40e1-ba99-80313943525a · inbound

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding cites this paper.

Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding Kwai Keye-VL 1.5 Technical Report

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:21:29.819348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T04:21:29.526008Z digest=sha256:91a710eaa01fce71dfb09bdbad79276b49f9ff09776701565e239238ec6629f7

Observation 64fa5fcd-4120-429e-9e3b-143122edb5e3 · inbound

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models cites this paper.

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models Kwai Keye-VL 1.5 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T03:37:50.225051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:37:50.225051Z digest=sha256:a58efd5ec4b59eb9f4c71adaee1f690d357e8d93d771e43f3aafadd9acc2bfd7

Observation 6175d828-74e7-46f4-b61b-64de26e54434 · inbound

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding cites this paper.

Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding Kwai Keye-VL 1.5 Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:51.308252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:47:32.778695Z digest=sha256:7bcec59965e7f31077b1e7017f15170c0bb3c722a9afddae722153ac14ca2653

Observation b131c060-de44-4390-8e58-9b6d8e67149d · inbound

ESOM: Efficiently Understanding Streaming Video Anomalies with Open-world Dynamic Definitions cites this paper.

ESOM: Efficiently Understanding Streaming Video Anomalies with Open-world Dynamic Definitions Kwai Keye-VL 1.5 Technical Report

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:51:03.556398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:55:32.945721Z digest=sha256:d0fe80379c2c81a6efed4248e3740b23d1388d7afbadb0603525247aaf6f715b

Observation 3305d732-3469-4c84-8933-0d60cff64d7a · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Kwai Keye-VL 1.5 Technical Report

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:41:04.304928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:0a7b3c17f46d1d3a20b846b4e95e4cbe9886feb41045665b6cb38db1062c6777

Observation fb2752e8-e3f4-4f53-aa8e-34eb48ed1454 · inbound

Visual Preference Optimization with Rubric Rewards cites this paper.

Visual Preference Optimization with Rubric Rewards Kwai Keye-VL 1.5 Technical Report

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:00.979167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:45:52.980881Z digest=sha256:5e8ff8ae12c7a452046ad45db94e5dfd7c4e759b587f8cb0b2024bfd1bca0583

Observation 41e5262b-1c59-46c4-a7c8-f586ef0ccd47 · inbound

Building a Precise Video Language with Human-AI Oversight cites this paper.

Building a Precise Video Language with Human-AI Oversight Kwai Keye-VL 1.5 Technical Report

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:04.503586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T00:37:31.858728Z digest=sha256:550b62e24266ecf6e498c9e5ae73ce823cff94b4ba22c0fb0936cf68276a8c3a

Observation 3e9fdb62-85af-4771-8bce-e8c01f4fac48 · inbound

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration cites this paper.

Scaling Video Understanding via Compact Latent Multi-Agent Collaboration Kwai Keye-VL 1.5 Technical Report

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:21:09.428567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-09T20:11:11.410051Z digest=sha256:54a293dbdae5caaca7ed9e3d80353f0dffc7fb17d62aa9b5cd3690df9ce7d39a

Observation e1be7009-35e8-44c4-8976-5aa3fe49df56 · inbound

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs cites this paper.

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs Kwai Keye-VL 1.5 Technical Report

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:52.779669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T03:57:57.791351Z digest=sha256:287e3f414948ded380ca378db10c6dac15dcc9a24c3a7bb6b073b1095972f1a1

Observation 8e4e4b6a-0863-407c-bf3c-7af9e46783e8 · inbound

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation cites this paper.

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation Kwai Keye-VL 1.5 Technical Report

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:51:28.432591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T03:53:46.809307Z digest=sha256:e31d296501b3574d98ffa4811223daab59579a63933bfcbe122da91c22eeed96

Observation fa9b594c-bbee-4731-b213-66cc62c908f1 · inbound

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation cites this paper.

SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation Kwai Keye-VL 1.5 Technical Report

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:53:02.218280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-14T21:52:08.637283Z digest=sha256:692111e9f63ff65d220eeaa8307551f32471a3a3efa34e9fdd508a0553bb9c78

Observation 30606730-05cd-4942-b98e-cc2c166a7426 · inbound

Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning cites this paper.

Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning Kwai Keye-VL 1.5 Technical Report

Reference 1

Resolution
malformed identifier
arxiv_id, observed 2026-06-29T22:54:01.479148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T22:44:59.860225Z digest=sha256:38ee9ba1fc91c8aa001fb1c74f0e56911bc21495ec779494c369cf54c445195e

Observation 4d38d9c6-94fa-423b-aac2-bcab6e55ed0d · inbound

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker cites this paper.

Towards Open-World Referring Expression Comprehension: A Benchmark with Training-free Multi-task Consistency Checker Kwai Keye-VL 1.5 Technical Report

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:24:02.133497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T23:15:02.455554Z digest=sha256:f1916be28d48b7cd1a09845a130fce770227a8a952f484938b1aea149f425b46

Observation 5b48ea6b-4a0b-444a-a5ff-4f6ef7bdb83c · inbound

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence cites this paper.

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence Kwai Keye-VL 1.5 Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:13:59.514435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T22:12:05.365596Z digest=sha256:320be27e4e1db0f563e201aed221d8233fa68811a9727fd0b61ec46330c7dad1

Observation 2fda1bbc-b5ac-491f-a4a9-9cf40d4ef74c · inbound

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams cites this paper.

IPIBench: Evaluating Interactive Proactive Intelligence of MLLMs under Continuous Streams Kwai Keye-VL 1.5 Technical Report

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:33:51.038683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T18:24:57.881644Z digest=sha256:02fe346a84edf1a22a4a88e984331e32f59a34b4429f78bd27677e9afbd27abc

Observation 4e4d2b64-a9b1-4fb5-ae5e-bf1688a568fb · inbound

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding cites this paper.

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Kwai Keye-VL 1.5 Technical Report

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T17:53:46.753893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T17:52:59.346637Z digest=sha256:9995278d63dc329ce2fdc2bc49510c6253e547e925abc5136f66bfe786d4a541

Observation b2f52786-ada2-4b3a-8b12-92cb7ab9a5ab · inbound

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events cites this paper.

Moment-Video: Diagnosing Temporal Fidelity of Video MLLMs on Momentary Visual Events Kwai Keye-VL 1.5 Technical Report

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:20.836897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T14:50:02.159411Z digest=sha256:d9397a7e097b5316af21a3ad5ad30cba82f982d3b59110000e1faf8971f38034

Observation 32456377-9743-448f-aad7-5c34295d1fde · inbound

AdaCodec: A Predictive Visual Code for Video MLLMs cites this paper.

AdaCodec: A Predictive Visual Code for Video MLLMs Kwai Keye-VL 1.5 Technical Report

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-28T15:22:19.505199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T15:20:48.248576Z digest=sha256:6611c577d57f1273f60bdc693fdc22c1b186cd357c05f12d8ea418e9cd8d0371

Observation 2d3b4783-77b5-4c33-8d97-49711e774cd4 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Kwai Keye-VL 1.5 Technical Report

Reference 216

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.720322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:90f8bc74e6ced2cf9eb558ba1858f10aa44abc38ceebde165389744ed9b41d24

Observation 9a79df8c-72fd-42e8-8823-f3234311aa4f · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report Kwai Keye-VL 1.5 Technical Report

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:27:37.037810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:b7a80aad7d381ae2921a51aa0f4358080014c72944e487133f8b3ed2e579fff0

Observation 3520b587-a9e9-4f67-800a-215826e6ea14 · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Kwai Keye-VL 1.5 Technical Report

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.945834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:9448fa58194565927df7b7aa1c49ffcd1c052d7fdb12e924a1c495411341219d

Observation ef9a48c3-39ad-43a9-b2df-7f1179c7e568 · inbound

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering cites this paper.

ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering Kwai Keye-VL 1.5 Technical Report

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:57.603559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T00:24:20.208132Z digest=sha256:91afd13b83267eec8e44221807f09ecbad0a7c737143a653db3b0316800b7918

Observation e260f3e2-c21d-4845-9da8-2f498738f837 · inbound

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs cites this paper.

MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs Kwai Keye-VL 1.5 Technical Report

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:34:19.457791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T06:25:38.593423Z digest=sha256:1a6c7e481fa00e6db8852c16b4bd08bbc0b1450b272cd2449060b28db33dfbe5

Observation 32b845a2-5f39-41b7-b819-87ae2d8710fc · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Kwai Keye-VL 1.5 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.600830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.600830Z digest=sha256:140197414fad13b5f7b76ee2be43402a6876474825392d3c7914bf1caadfe1d8

Observation e7d57d85-9cea-4001-bb59-f7aac00a37b4 · inbound

RefCaptioner: Multi-Reference Image-Grounded Video Captioning cites this paper.

RefCaptioner: Multi-Reference Image-Grounded Video Captioning Kwai Keye-VL 1.5 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T05:08:20.137192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T05:08:20.137192Z digest=sha256:1005471cd2c976ebd83ea8b0e92d4f08687f7ee9b98743f47971e9cbad9a1140

Observation 24489f74-013c-4643-89e9-eff675abf1c1 · inbound

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards cites this paper.

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards Kwai Keye-VL 1.5 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:40.944996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:46:40.944996Z digest=sha256:e11540f7c7ec4e7f695cd488247a16c94cbaa0f7ded1b38a134dd8edff1f4143

Observation efbfd439-c2f4-4a1b-bc92-e35173578d3f · inbound

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards cites this paper.

DocPO: Advancing Document Policy Optimization via Tailored Step-Aware Rewards Kwai Keye-VL 1.5 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:56.101858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:59:56.101858Z digest=sha256:57bc688114f59e7bde033b1d30ea96716ed91c3d2529569f6314cfc1f1c5b38f

Observation 0af4d824-63b7-4b5b-8709-ca23b5f8e758 · inbound

InSight-doc: Agentic Visual Perception for Long-Document Understanding cites this paper.

InSight-doc: Agentic Visual Perception for Long-Document Understanding Kwai Keye-VL 1.5 Technical Report

Reference 129

Resolution
unresolved
no resolver link, observed 2026-08-12T20:43:52.925722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T20:43:52.925722Z digest=sha256:68d9c44088136eed8dd382e014252d9ebdf2c28d67cf6cfcaaa89a41ed627c8b