Pith. sign in

Paper Citation Record · LEDGER

Seed1.5-VL Technical Report

As of 15 August 2026, this Paper Citation Record lists 100 of 208 outbound references and 100 inbound Pith citation observations for arXiv:2505.07062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07062 v1

Coverage vector

measured 100 of 208 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T05:26:04.960844Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 100 of 197 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:57:17.357726Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T16:15:06.198778Z

Reference resolution

100 of 208 outbound references displayed

  • verified exact51
  • verified fuzzy45
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation afcfd011-6c93-41fd-a90d-904ea8bc7c11 · outbound

This paper cites an unresolved cited work.

Seed1.5-VL Technical Report Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-11T05:26:06.278271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:cec0e457a9e0a520be35dda8220eb6d1e59772e062e140151d5aba33033299d5

Observation 2fac0f76-f434-4805-9089-0cddead8d2fa · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Seed1.5-VL Technical Report Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.205044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:bf1e4c1b3b12b2e262df9fe76912b23cb5e06ac0c2a73c10f722702b6ba1a654

Observation e43bdde6-fc62-4d79-83a5-a18769811fe9 · outbound

This paper cites Countgd: Multi-modal open-world counting.Advances in Neural Information Processing Systems, 37:48810–48837.

Seed1.5-VL Technical Report Countgd: Multi-modal open-world counting.Advances in Neural Information Processing Systems, 37:48810–48837

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.288023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:45247925d98d26afeeec5230780c26b5278089d96a9fb0d518df240c5dbce0de

Observation a1c198b8-abb0-44c3-b207-214fe00f3088 · outbound

This paper cites Understanding Alignment in Multimodal LLMs: A Comprehensive Study.

Seed1.5-VL Technical Report Understanding Alignment in Multimodal LLMs: A Comprehensive Study

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.032367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:ffadcfa5c3f402fce089e5c3c7c15d109801ad1e822f649a17e4a94c9ea40402

Observation f1bf4ed3-812b-4a0f-b4e2-a11f97f052f0 · outbound

This paper cites Claude 3.7 sonnet system card.

Seed1.5-VL Technical Report Claude 3.7 sonnet system card

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.294589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:dfbe8d570dc70ff6331118cf599e84781ee622c3c7afb0081d757c5082d63756

Observation 5d7fb1ab-4250-4767-bd01-d397437cc2cb · outbound

This paper cites Claude’s extended thinking.

Seed1.5-VL Technical Report Claude’s extended thinking

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.299128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:ad5312636a940be7fce56511bdb95ff704c67e8390f4e85304ad669220c3cfdb

Observation d46de2ad-e269-45ff-a58b-31e1eeb1baac · outbound

This paper cites Qwen2.5-VL Technical Report.

Seed1.5-VL Technical Report Qwen2.5-VL Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.047883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:9dfa0a8833f0d5df56d386c8dae87a251083f02fdf545198c3eaa8568c7fc799

Observation 1ca5c67f-9af3-438f-8585-b5e9c1637ef1 · outbound

This paper cites Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models.

Seed1.5-VL Technical Report Objectnet: A large-scale bias-controlled dataset for pushing the limits of object recognition models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.303629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:4edf5401343cd6f7369c473154db51897a769f9404546bef11bb6d60f65f4cdd

Observation ae0a26b0-86c9-44c1-bcf6-1e8a0a4070b9 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

Seed1.5-VL Technical Report ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:47:09.105767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:bc098af5d54e5de307b8b5765c6569e312ef8a73079b6c4a95cad83abde38967

Observation d0cb9a49-493f-49f8-a2a3-5ea434b3113c · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Seed1.5-VL Technical Report PaliGemma: A versatile 3B VLM for transfer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.987351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:44097d4d2f3e6893db3c6d3b00429ca1e3f8140375ab48d4c2802d42ad4b65bb

Observation e7ed0343-1903-4b5c-b731-67e105619faf · outbound

This paper cites Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale.

Seed1.5-VL Technical Report Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.586427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:95bb480c323ec3e918609bf7a26e1d8420b3b52ec63872a0f9dc8419cfce34fd

Observation 4d7bc3d6-cca5-45e2-9860-084fb8cfc5d9 · outbound

This paper cites TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models.

Seed1.5-VL Technical Report TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.681583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:41fd654210223faae5392ce5c6f5c0991cb41597c60a4417ec6522c1c93a5d74

Observation b78cbcf5-abcd-4ff4-88d0-d514868036a1 · outbound

This paper cites FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion.

Seed1.5-VL Technical Report FLUX: Fast Software-based Communication Overlap On GPUs Through Kernel Fusion

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.905725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:0f56cfb38948ce924591e0a00056b8abe292f40c37b07375ea934cef9e1d3e4f

Observation 7ee4d3bb-01e7-41cf-86b4-87ff10d51410 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

Seed1.5-VL Technical Report MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.955218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:e75c3dd164e05fb22bfd28553fee94686be49d7e55f17e546854222b9dd01146

Observation 2caeec6f-6c80-40cc-9668-6d42f76dcdf6 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Seed1.5-VL Technical Report Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:41:44.612219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:bb16ede4d0939e461c85e4fee8d4c29a0e2cac77d95fdf578fa7f9d11a48ff7b

Observation 6dc2d73c-eed9-495b-9d1f-6eabae4a22a8 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Seed1.5-VL Technical Report Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.311222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:7700afe25dc3a09ac16eb41b7c3af1a9cf6d980c1166fb681bfd6d5dfa115258

Observation 52f4aac0-f2bf-4fc2-bb29-f3dc51fa2321 · outbound

This paper cites Yolo-world: Real-time open- vocabulary object detection.

Seed1.5-VL Technical Report Yolo-world: Real-time open- vocabulary object detection

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.318911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:027aaa7d631bd30e0b6017080d996e87e3b4ba94743c7e437ce10b302ca1ad30

Observation 44c2e2ee-1414-4644-a691-54fb50e9a9d3 · outbound

This paper cites PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns.

Seed1.5-VL Technical Report PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.081961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:16a9b50dae62733b1ef41c5e87ca9863894c7b4f1bfc008cfbcd600f52139dc2

Observation 769e071d-a57c-4a42-a863-a23d22964c86 · outbound

This paper cites Lost in Time: A New Temporal Benchmark for VideoLLMs.

Seed1.5-VL Technical Report Lost in Time: A New Temporal Benchmark for VideoLLMs

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:06.098537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:84627f8c0fcf208228779d24ba4cfff1fbe3f53ff59b7dae56387d217fdcda3c

Observation e299bb67-faa1-4f78-a3cf-f2fe4ceb9ec0 · outbound

This paper cites Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution.

Seed1.5-VL Technical Report Patch n’pack: Navit, a vision transformer for any aspect ratio and resolution

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.323904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:61c3b71b67dbfa6a841b8d260bb370fe14305ac42341162312b2bc0a0c836e82

Observation e4e88b71-6c17-4e55-9f22-d64811bd6c3d · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Seed1.5-VL Technical Report Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:55:12.888993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a7a6273ed266492a50fbb93cafd900f7534bbc7c8040ded54e34d69e0375304d

Observation 68302095-fe8b-43b9-aa82-07fa136a25dd · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Seed1.5-VL Technical Report Imagenet: A large-scale hierarchical image database

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.328500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:3a4017094f4289794c846e3a16881155667076035e105d608ba2a06186fe5bc3

Observation 5c81fc57-ac1e-4e04-81bb-c0024add02cb · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

Seed1.5-VL Technical Report Unveiling Encoder-Free Vision-Language Models

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:26:05.462877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:246f921048e370239821b1c0160dd43620d78d58bdfef291acba7856fc3cd633

Observation c32d8719-ac79-4347-9124-6935b7d63034 · outbound

This paper cites Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models.

Seed1.5-VL Technical Report Self-play with Execution Feedback: Improving Instruction-following Capabilities of Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.485132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:5cf75993c10b84525702dd002e1d8dd517a9241ee18908c96d9872698eb162d8

Observation 033ed50e-56d8-41f9-89f3-a6252a2938af · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Seed1.5-VL Technical Report An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.491357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:6bde13d44a6a454781561da670ec46c5b9016e8891dd64d04ac534a5283a2e6a

Observation 0744f927-0bff-469a-bd64-35c6682fd51c · outbound

This paper cites Counting out time: Class agnostic video repetition counting in the wild.

Seed1.5-VL Technical Report Counting out time: Class agnostic video repetition counting in the wild

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.336018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:cd667aa894064c7b568dfd064843983602a649d6256bac1d08dd10b68bfcf4c3

Observation 4c5f184d-9b44-4e23-a71a-19977cb6d7c2 · outbound

This paper cites Data Filtering Networks.

Seed1.5-VL Technical Report Data Filtering Networks

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.559184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:ebe49199d368724950029f4152643e85ea7637a5ac91d23d1bc36bd41a51779f

Observation 4a417498-d993-4e04-8995-825af5143ba1 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

Seed1.5-VL Technical Report Eva: Exploring the limits of masked visual representation learning at scale

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.341778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:1bbc6c784b49d3cfb5dac438cffc3a843e7790fb634221e889d8ed3f34735d35

Observation bfdebabc-c4ea-45fd-ad31-39e70d3dfba5 · outbound

This paper cites Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation.

Seed1.5-VL Technical Report Optimus: Accelerating Large-Scale Multi-Modal LLM Training by Bubble Exploitation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.608208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:43902e1d66535f9db4e9ba8496097a5d0db345a77fc8a4fcc9870175ec13bde0

Observation 1ef4e04e-0b5b-4f35-b2db-c16925e41dce · outbound

This paper cites Helix: A vision-language-action model for generalist humanoid control.https://www.figure.ai/ news/helix.

Seed1.5-VL Technical Report Helix: A vision-language-action model for generalist humanoid control.https://www.figure.ai/ news/helix

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.348494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:657596bdaa41fea93b5544bf27d61938df49c081153050520f5bab7426950834

Observation 53ca0743-daac-4203-b767-93aff87a3a4a · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Seed1.5-VL Technical Report Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.774438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:756792efd116cf93431e663429e851e9a5c5651d8b6efd4ad37f2d094489fd8c

Observation 6516d9d2-cdca-414d-a6cc-52f0cb4445ce · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Seed1.5-VL Technical Report Blink: Multimodal large language models can see but not perceive

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.353238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:65820df1d40d45b5f0afd410755e1d026fa5b2e5285fd1281817118dfa37a42b

Observation a70836d8-b460-4680-a901-79595d8f3d12 · outbound

This paper cites Tall: Temporal activity localization via language query.

Seed1.5-VL Technical Report Tall: Temporal activity localization via language query

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.357984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a1697f590436179e5db1b071f3122d1f49dd50e89f3e3a7ac67551d9cd228bd8

Observation 05887b6e-1f85-4531-b72f-50449df61832 · outbound

This paper cites Experiment with gemini 2.0 flash native image generation.https://developers.googleblog.com/en/ experiment-with-gemini-20-flash-native-image-generation.

Seed1.5-VL Technical Report Experiment with gemini 2.0 flash native image generation.https://developers.googleblog.com/en/ experiment-with-gemini-20-flash-native-image-generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.367407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c94ca38450032f52fcd578e655804882de91ca613916ef81049981f27b629bfc

Observation fc8c93b0-3905-4e7a-85c6-46a92376c5ed · outbound

This paper cites Saliency-guided detr for moment retrieval and highlight detection.

Seed1.5-VL Technical Report Saliency-guided detr for moment retrieval and highlight detection

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.981905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:7a55ae3eccc513dd92c44859701bd1ef27231c04b1c749165b46f67a5740465a

Observation 30db478d-cbaf-4879-aa3b-5b22634ce670 · outbound

This paper cites The Llama 3 Herd of Models.

Seed1.5-VL Technical Report The Llama 3 Herd of Models

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.015925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:d32e9982fc5c739a0ae34b9b30e1823fbc1f37709bd4a7ffd0f5ce7ca0ee763c

Observation 49f84c85-9b74-4a42-87e5-1e7d4ceceb8a · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Seed1.5-VL Technical Report Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.372253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:e89b8cc2dbc8fd8eb14df3054b9380cda8fe88f879fdf1c619bf017a0c58b6d5

Observation 9ffc2f68-d28e-4814-b09c-179010b8042f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Seed1.5-VL Technical Report DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.042133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:bd33c82583fbcde8c29d4391ec1050e332d3efb6944890419f5cc526ac38966c

Observation 203ec044-cc71-4792-877f-e90610d595e9 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Seed1.5-VL Technical Report Lvis: A dataset for large vocabulary instance segmentation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.376984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:5d24c24822e8e1adc5f09f0ef8de47d2c675568cd88d40f65520fd2d732fba3f

Observation d5ebee7c-323c-46ff-a9db-0e45abb7a7d2 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Seed1.5-VL Technical Report OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:38:21.105975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a12109d3a0c9d932022ce817452b9ecc3ae888c800267008d901ea7abcc0ccf3

Observation 37746a3e-98f4-4307-b361-d4434033f04f · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

Seed1.5-VL Technical Report WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:43:35.479641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a370b4ce2b6689d7fdaaf67bdc5ecfa2b399ef637cce65f37c6cc1a43962479a

Observation 1b50b982-6127-4f1c-a8e5-3f58838874f7 · outbound

This paper cites The many faces of robustness: A critical analysis of out-of-distribution generalization.

Seed1.5-VL Technical Report The many faces of robustness: A critical analysis of out-of-distribution generalization

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.381706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c40163c5458b97827ac5dcc74071951f6240fa09e706b5e1b948e85bcf3939e4

Observation ff96f097-bebd-4349-a2aa-1db2fa5e4b6b · outbound

This paper cites Natural adversarial examples.

Seed1.5-VL Technical Report Natural adversarial examples

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.391576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:08e6ff45bf353e8382c1b1ebab720d7891914e2c2d7fabafcab16d4f54d308b6

Observation 0be140a3-ab89-4a5b-b008-ddf7dce85e9c · outbound

This paper cites Scaling Laws for Autoregressive Generative Modeling.

Seed1.5-VL Technical Report Scaling Laws for Autoregressive Generative Modeling

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:49:43.923372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:eefdbd0eb25edb60a1cfc21c52c64f9d4cf6dd3beca2edb146b622e7af56629b

Observation 6b7e3d2b-dca4-42e4-88c0-aad33465ab40 · outbound

This paper cites Training Compute-Optimal Large Language Models.

Seed1.5-VL Technical Report Training Compute-Optimal Large Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.128471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:19626ed3e1a457c1c219cb3f18d55fae1968043d52efe1c62d860fa07f67b2db

Observation 48407459-62b0-4714-a90d-1bdd31dc1aaa · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Seed1.5-VL Technical Report The Curious Case of Neural Text Degeneration

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:18:24.042352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:0b5d2af6d7b4c48aad3e5875b135f1a3796f184910cb510582f9510fd8873266

Observation c389374c-5eab-4d3e-9c2a-c10924b44ed9 · outbound

This paper cites MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models.

Seed1.5-VL Technical Report MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:48:11.661227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:9ee0748adf60f5f01843ce9befc1a79b3b007d691f86b3622e83124054fdfe23

Observation 8b90fa82-393e-438c-9a3a-e2da6583651c · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

Seed1.5-VL Technical Report Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:32:41.480822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:10238ce9eb1ca0575122f1f49cf044c9023ef556fc8277de02563f9ec606cdab

Observation d872c490-ca22-425b-928f-5311fd07972b · outbound

This paper cites Gpipe: Efficient training of giant neural networks using pipeline parallelism.

Seed1.5-VL Technical Report Gpipe: Efficient training of giant neural networks using pipeline parallelism

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.396470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c4c209712a8b678e3fe0b00ab3557828992b167e799e379a980a9cf3e9e841fa

Observation f1005a7c-390d-42be-bba1-17b0d6d91301 · outbound

This paper cites Online Video Understanding: OVBench and VideoChat-Online.

Seed1.5-VL Technical Report Online Video Understanding: OVBench and VideoChat-Online

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.215670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a04aa06dd3e6de447652dcd242f5df028e2ad9e717bbe2173add425049ed5d1a

Observation 6182c567-ca84-41f4-a445-04c7b292a1db · outbound

This paper cites Classification done right for vision-language pre-training.

Seed1.5-VL Technical Report Classification done right for vision-language pre-training

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.400636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:1994a84083f4406ec05c5f40205a4dd9eb3404e53f3c37381d605aba3ebdf4e0

Observation 602425a6-015e-4768-9605-dabf9c43e2e4 · outbound

This paper cites D., 2007, @doi [Computing in Science and Engineering] 10.1109/MCSE.2007.55 , 9, 90.

Seed1.5-VL Technical Report D., 2007, @doi [Computing in Science and Engineering] 10.1109/MCSE.2007.55 , 9, 90

Reference 53

Resolution
verified exact
doi, observed 2026-05-11T05:26:05.347076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T08:08:11.947291+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:2a454796b3df51fecdc45924f623b30663bbd5145603d94f2ef9941314a6d39c

Observation 4b45629c-a7b9-4085-a3e9-431dd431efec · outbound

This paper cites GPT-4o System Card.

Seed1.5-VL Technical Report GPT-4o System Card

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.379349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c8866b3c4cf1a71ea5b7c2a84c4029cc2f990e38152f50ef5d8412e2af1135ad

Observation b67af7ca-89d7-4017-8dcd-a9aed805bb63 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Seed1.5-VL Technical Report $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.406172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:7f97d7218fcb2eff4b93f8bf7dbb3112f272389718af6593565721187b9e5a42

Observation b05d473d-0521-41c9-a644-524f69447186 · outbound

This paper cites OpenAI o1 System Card.

Seed1.5-VL Technical Report OpenAI o1 System Card

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.415512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:9c661b542cd06050a18ac53460b6461abdca8c0bf5c3f2afaf22908a2ad6bfe2

Observation 309dea63-518c-4c51-af93-3783303ff936 · outbound

This paper cites In 21st USENIX Symposium on NetworkedSystems Design and Implementation (NSDI 24), pages 745–760.

Seed1.5-VL Technical Report In 21st USENIX Symposium on NetworkedSystems Design and Implementation (NSDI 24), pages 745–760

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.404958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:7f2338fb905093eb6373844316597b25b06694ed34be35f327bff361cb78d4b3

Observation 3f3c3296-3673-4a32-a746-7c6ae81b6f56 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

Seed1.5-VL Technical Report FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.448721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:808e568e9c1d3783cc19b776c9bd7de6495fe9042bf19647c62966c8186b5ec5

Observation 7f19c731-9841-4029-9af4-962cfd0fea24 · outbound

This paper cites Scaling Laws for Neural Language Models.

Seed1.5-VL Technical Report Scaling Laws for Neural Language Models

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.456643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:ee8a2ee18dcef92b15f955e9437d1d440c83a3fcfd31354fd7bee09446e4878c

Observation 8ff76b86-eaa7-46f9-9c10-6a321a390258 · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Seed1.5-VL Technical Report Referitgame: Referring to objects in photographs of natural scenes

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.415359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:5205de4da3c7d9f65094631b34b5b6accbf14605fea93593aff857e7bf451d03

Observation 2ba3d0f2-ebc1-4318-91ba-aa660858ecc6 · outbound

This paper cites A diagram is worth a dozen images.

Seed1.5-VL Technical Report A diagram is worth a dozen images

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.423351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:4498a56dd6b6a4892d7a9cf5d64d5d41c91c9cd7fa07ed7a929507a53e05ae1a

Observation 890d3d42-54dc-4fa6-8e10-aebe70c29ca4 · outbound

This paper cites Ocr-free document understanding transformer.

Seed1.5-VL Technical Report Ocr-free document understanding transformer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.429262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:6f442e3810e4125355adc4a55c0f9460c33da2451252838bc261fe51e23fa6ff

Observation 1c6a1699-9822-4bb0-adc9-e031f6a7bbfe · outbound

This paper cites Openvla: An open-source vision-language-action model.

Seed1.5-VL Technical Report Openvla: An open-source vision-language-action model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.439687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:65cb87b881c69d0751a06f374fdf2411c1e0d9ad9e21814cd18b9f4ea9d4603e

Observation 12256451-28ef-4dea-9daa-58621d2d0fbd · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Seed1.5-VL Technical Report Adam: A Method for Stochastic Optimization

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.552726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:b8497bf89a36fdeadcfe1a5ed4fe346d53cdd964d8eee59793ef92e2ac301db0

Observation 87a025c5-d2a8-4348-9836-ddf276b9acf2 · outbound

This paper cites Reducing activation recomputation in large transformer models.Proceedings of Machine Learning and Systems, 5:341–353.

Seed1.5-VL Technical Report Reducing activation recomputation in large transformer models.Proceedings of Machine Learning and Systems, 5:341–353

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.445032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:110b79e1aa0e00e99638d76ceaca095f3220bbe5d22a8de3b698e87bdcd1918d

Observation c012fa6c-48c5-4d5d-b938-ebd8d6ed12d8 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.IJCV.

Seed1.5-VL Technical Report The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.IJCV

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.451376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:7a13c6f1933b54a0955424b3dedbfcc95b270ea650d46b40a268d3b9ff681ec0

Observation 02d3c154-8b2e-4bc1-b46c-4f627f3fe759 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Seed1.5-VL Technical Report Gonzalez, Hao Zhang, and Ion Stoica

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.460726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:894f5169aed024399af31ec8de99409cdc91eadf160c3cf7ec58c29273a9b55c

Observation fd045068-0ad5-4337-b55a-975fd7ff9702 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Seed1.5-VL Technical Report Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T15:03:08.022500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:94c7ef4accd87182e42e14e5e9f8cfcf33996ad890610781b8e9c17161d7a8a1

Observation fe7e3a20-3c4b-4e94-8988-52459e76e5dc · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Seed1.5-VL Technical Report Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:26:05.664367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:005ddea7880871d2e5af3f3075f14542863c60c11e7c9b0a250037939971c971

Observation 7cd21718-a876-4dc6-abd9-b77cb7dd2a94 · outbound

This paper cites Echarts: a declarative framework for rapid construction of web-based visualization.Visual Informatics, 2(2):136–146.

Seed1.5-VL Technical Report Echarts: a declarative framework for rapid construction of web-based visualization.Visual Informatics, 2(2):136–146

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.465664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:2d838b1c838e383f640d9107bcedd6bab3f35709b615a0cb1d80f783f794de40

Observation e66a46c0-3c91-4517-ac37-8e63025eff47 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Seed1.5-VL Technical Report LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:01:54.585341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a0666effe07fbdf46f804fa3fe58724293c75b78503fdb96746903539be42056

Observation 33b9e398-5f45-423e-92dd-6c3272960c36 · outbound

This paper cites ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use.

Seed1.5-VL Technical Report ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.723122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:6a21d66c17ac5b2d1662b094bc9930fd5e9e9c87a624d703559f55c8b761cf1e

Observation 675eeab5-8790-4c16-bf3d-5c39d5151d88 · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Seed1.5-VL Technical Report Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.470167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c7a528e1c65814c409b0cf8a80a3f1479bdee34ed14bd0c603105b03ddf557d8

Observation 6cccf783-ee74-4d1e-b9d0-eac6fdb1ab84 · outbound

This paper cites OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?.

Seed1.5-VL Technical Report OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.803350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:37f96ab61d1f15c45d1f2a4e1d1846b3d5addf6119396c110380f25a6180b548

Observation 6d210c9c-db51-4ece-b83e-e8e59d775169 · outbound

This paper cites The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models.

Seed1.5-VL Technical Report The Devil Is in the Details: Tackling Unimodal Spurious Correlations for Generalizable Multimodal Reward Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.816865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:a1c189418d074c2c23dbb1212af0458efa4e459a4abc864f68818b1265cdb08b

Observation a7739048-e734-4807-b2d6-c769a53a1e94 · outbound

This paper cites StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding.

Seed1.5-VL Technical Report StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.849502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:2355bd859a93c597b0a284c3487dde02b324b975673cd4c5aed3a8dc3d825d57

Observation c963f759-fa33-4d40-926c-2d7b7c60b7df · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Seed1.5-VL Technical Report Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:28:28.418803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:53cb9d07dd254f3db3fde2d4de35395c8cb3baf02a0562a41b36105e0a7708c2

Observation e125210c-4fab-4ab3-9c5a-50a03211a77f · outbound

This paper cites Visual instruction tuning.

Seed1.5-VL Technical Report Visual instruction tuning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.476449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:4e505f862309183738bad34ec10f1624083b5913ca2c42181a521831977f3a66

Observation 45680f7d-05e7-4bf0-8c52-88abf138915e · outbound

This paper cites VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?.

Seed1.5-VL Technical Report VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.946519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:0ad2528b6bb761ceb1294bbe2f62d8d48c4be077d764d4ea7136b25d98c2163e

Observation 40f7a881-bd38-4256-b550-fd9872191a45 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Seed1.5-VL Technical Report Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.485651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:e8238d8c63cd48519376369919714de15f7f1b30871ecdd2299f972489a85198

Observation ddb64011-8cbc-4818-b680-416ee764cb66 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Seed1.5-VL Technical Report Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.493387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c3354c146c6e2df07210bd5e1a754cb347103846605e40ebdf02fd0b5ca58327

Observation dee4ab59-fd2c-489e-8f14-73b3a8735974 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos?.

Seed1.5-VL Technical Report TempCompass: Do Video LLMs Really Understand Videos?

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:46:17.144047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:d90aee4d6711baa3b0ccf6787192295287f4d4dc4f8305c900baefa1dd9c5ccf

Observation 7cd549ae-5d9d-4e8e-af05-5d325c973b62 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102.

Seed1.5-VL Technical Report Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.498097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:46befba7a29681deeda9bff051ca71e071b2127230e9dc301a814d59dad107ac

Observation a34b8a8b-4669-4658-85b2-d5557e30d5d1 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Seed1.5-VL Technical Report MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:05.989375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:918cee6f6edad375f12372918b8da912e4a90be4ac46ee4910aa26f8fa4dc2e9

Observation 5b7c0424-0ad2-48dc-ae2b-ae49ede7034b · outbound

This paper cites Ursa: Under- standing and verifying chain-of-thought reasoning in multi- modal mathematics.

Seed1.5-VL Technical Report Ursa: Under- standing and verifying chain-of-thought reasoning in multi- modal mathematics

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.001472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:64b54f01e945ce85134677f4e70235d432b189b563ce7e63a71e7c2ee98d080d

Observation 4d6b2d37-f2b3-4221-b7d1-cfc6bc8689b0 · outbound

This paper cites Generative Reward Models.

Seed1.5-VL Technical Report Generative Reward Models

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.009302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:366cdd757dcdd9db9a7b5c2eedd9c7c71b22b5d7907b6b01aad8b27c9fce3a93

Observation fdb0da2f-6ebf-452b-a2f8-a44d1a594cf2 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.Advancesin Neural Information Processing Systems, 36:46212–46244.

Seed1.5-VL Technical Report Egoschema: A diagnostic benchmark for very long-form video language understanding.Advancesin Neural Information Processing Systems, 36:46212–46244

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.502713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:28bffe1161d2e19247349dbb0056017dc6784b26acfb94de0b4f07bb0220bcc7

Observation 7ddd73e2-591e-4fce-99f0-8b25006ba738 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Seed1.5-VL Technical Report ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:13:07.396442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:d65ead866f395e025feb3536de8bee1bc1fb24156cdd55c2665dcc46a5f2b493

Observation 89a6b06a-1091-4261-a221-92a35397ba27 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Seed1.5-VL Technical Report Docvqa: A dataset for vqa on document images

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.507185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:f22d0de7241095da787129e08720ce7b21ab37c2e26e598730d4c36cc39dbe71

Observation ba0a29d2-3a8c-4cfd-886f-cf1f7e7539d8 · outbound

This paper cites Info- graphicvqa.

Seed1.5-VL Technical Report Info- graphicvqa

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.511470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:3402bc5a49d991323eab821152044aa019478b1d1179a854c1850a95dd467fb6

Observation fe0f7ed6-1c69-4b57-8df3-d082f58edab1 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation.

Seed1.5-VL Technical Report The llama 4 herd: The beginning of a new era of natively multimodal ai innovation

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.515776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:f516faf2ebc8058fe82dfde86d60e9c432f962096d2cfe2f44019913f458a3a7

Observation a476d431-18ca-4b33-b46d-eb513dae0cf3 · outbound

This paper cites Modeling context between objects for referring expression understanding.

Seed1.5-VL Technical Report Modeling context between objects for referring expression understanding

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.519936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:441750973e4988ea9d65f0a90ee9dd3e9af3d126ff51ea0f07f0b1cd46170298

Observation 127ec65b-80c7-4015-9ea1-2867c375c79c · outbound

This paper cites Memory-efficient pipeline- parallel dnn training.

Seed1.5-VL Technical Report Memory-efficient pipeline- parallel dnn training

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.524415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:3db1b8e359d9f59fe97df569c8774a6a66c166c811727df0475593576ce9f5fa

Observation 4b8b888f-1bae-438a-8669-c29a28dd0630 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

Seed1.5-VL Technical Report Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.528739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:e9d921b16c643fb693813f5c9f87ae19a5ea7c3f9e5dadc6359e4726e2e3e086

Observation 59e7d375-3180-4bdb-a701-d888ef651391 · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

Seed1.5-VL Technical Report Indoor segmentation and support inference from rgbd images

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.533141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:5d565f8714a178dba74a4cf35df019a1219a3cd31b1b593d8bf42fbca59ec8d3

Observation 9b4338fa-51ad-456e-b352-b7e501089b62 · outbound

This paper cites Gpt-4v(ision) system card.

Seed1.5-VL Technical Report Gpt-4v(ision) system card

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.537573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:d1a1df3fcf3ffabb05f3a1686b5c238326456fa880dd15a932a270618c017ce0

Observation 1d8d762f-4c4a-457e-8604-ae5f53943323 · outbound

This paper cites Addendum to gpt-4o system card: 4o image generation.

Seed1.5-VL Technical Report Addendum to gpt-4o system card: 4o image generation

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.541908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c160a30c3dc1fd055ba86d60d4969829371152b30279c00a645c6ec5d9bd5c4e

Observation 7b569b03-50fc-4213-9727-1736b58ec368 · outbound

This paper cites Operator.

Seed1.5-VL Technical Report Operator

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T05:26:06.545987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:542e5f0d6bfb56ee015adff8759e2ef3d43ba68e97769a48a0ed34b3290aa169

Observation 6f2c3e87-a354-4abf-bb94-4c044ab34a4e · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Seed1.5-VL Technical Report DINOv2: Learning Robust Visual Features without Supervision

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.146024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:510f3a4a32b55598e26adc3b7ca7e765e0661b791fc14f16aa61773a61ddc9fc

Observation cccf8905-622b-4a0b-8a49-6fe756b8475e · outbound

This paper cites Training language models to follow instructions with human feedback.

Seed1.5-VL Technical Report Training language models to follow instructions with human feedback

Reference 101

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:26:06.156357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:4e60825463ecf40544d00a752599369af13dcd348a895ad844f77c8a1f609328

Observation f4f514b0-5bae-4d7b-bb64-ab29ea276fb6 · outbound

This paper cites How predictable is language model benchmark performance?.

Seed1.5-VL Technical Report How predictable is language model benchmark performance?

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.167444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:2e4041cf436331001dfb8e18da456aa968db5c0b51655504068afc211b50176c

Pith citing papers

Observation b01406b5-88b8-4a52-989b-02fc4a1d11d9 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark Seed1.5-VL Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:55:30.168739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:c1137cddeaf3eb2c377e3974fa8321b6bed5e30d30ab47dc3ed6f3d1f789f250

Observation 53713d88-b350-4739-9bce-0668ae3d9b80 · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning Seed1.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:17.357726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:17.357726Z digest=sha256:02cba8d9e15089011944bbbec6429a3a3b504c9fdcd52a2523f155dda4276b98

Observation c6a16b61-2b4e-4409-bea9-8c62caff5779 · inbound

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs cites this paper.

MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs Seed1.5-VL Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:35:56.947387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:35:56.947387Z digest=sha256:a55db2fcc96f40d3c4aa46579faa7e5b11b3ef2aa9a50fd0f0ed4bfe03765e5a

Observation 4681f850-b3f1-48c3-9422-1a801a1c9c6a · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs Seed1.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:01.890332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:01.890332Z digest=sha256:d6203db31df4cf9a1572b062a205fc9df0e5110e90569033d09179519a49aea1

Observation 11484dac-1c3a-4667-965e-39c62904883d · inbound

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment cites this paper.

Generalizable LLM Learning of Graph Synthetic Data with Post-training Alignment Seed1.5-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:41.732462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:41.732462Z digest=sha256:4762c435e7f946d070fc118a790f17f4f1309ca0fafb29b97aeae008d0c9558f

Observation 45015151-b2e0-4234-9914-512f14f5603d · inbound

Affordance Benchmark for MLLMs cites this paper.

Affordance Benchmark for MLLMs Seed1.5-VL Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:53.010003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:53.010003Z digest=sha256:7bed18d24494d1568d24cfcb504e56f8f80631057fbc56681c9f0fdd3140bd3f

Observation b8dae6e2-b953-4ce0-8233-84253194770e · inbound

Native-Resolution Image Synthesis cites this paper.

Native-Resolution Image Synthesis Seed1.5-VL Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:27.691230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:27.691230Z digest=sha256:c6e6a42787be60a6b592e4347e34f48aaea87f9b8463d928b41952ba5a7ca306

Observation 9dafa5e9-0b8a-4fa2-9642-9d47c1c28142 · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning Seed1.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.663442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.663442Z digest=sha256:cdfcbd5c9c58903d0eb1ad44057ee7b1c8fa9919f52ba2a842088005e73232b2

Observation 158538fa-a423-4af1-88ff-2d6458423bc2 · inbound

SeedEdit 3.0: Fast and High-Quality Generative Image Editing cites this paper.

SeedEdit 3.0: Fast and High-Quality Generative Image Editing Seed1.5-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:06.482746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:06.482746Z digest=sha256:e8d6bb414cae9c8b48f530967442890a15a71a529b1f6d90b2535848d67ae5e4

Observation 2d5d455d-553e-4098-aa0f-ffc0607bc6b7 · inbound

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs cites this paper.

AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs Seed1.5-VL Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:05.608221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:05.608221Z digest=sha256:683974ba8a2c944ea0f1ac18c56aed1e3eb5ab9c68639b883056411bc08a59dd

Observation cad310a7-038b-468f-892f-0617b815aaeb · inbound

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning cites this paper.

MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning Seed1.5-VL Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:48.410947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:48.410947Z digest=sha256:e7b54466792e54b84fa9b85b4546bf59f2230d027445cffdd354c276ad73b608

Observation 09a0e717-34f9-4ede-8d0e-d02ee2bae432 · inbound

Event-Priori-Based Vision-Language Model for Efficient Visual Understanding cites this paper.

Event-Priori-Based Vision-Language Model for Efficient Visual Understanding Seed1.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:01.306179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:01.306179Z digest=sha256:fb7cd2b187b84f0e64171f5e343d446c6eea551680789090189c2e0f53215679

Observation b0e77c38-bd66-4c1c-afe5-753e6bd1b3bb · inbound

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers cites this paper.

DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers Seed1.5-VL Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:27:39.665617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:27:39.665617Z digest=sha256:770a4a0296bea04c6faa1ef22ab84df476592459857942b04aa14faef9e6e15d

Observation f3654494-ed02-4745-968a-0a9d0153c134 · inbound

Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification? cites this paper.

Breaking Bad Molecules: Are MLLMs Ready for Structure-Level Molecular Detoxification? Seed1.5-VL Technical Report

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:17:57.636188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:17:57.636188Z digest=sha256:c7e3b53d47c7bc59c1a2d90992b324e27e6a4166de5981137426b8cf033e2467

Observation aa453ad4-0ad9-4b46-8019-78d046e0a629 · inbound

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? cites this paper.

VFaith: Do Large Multimodal Models Really Reason on Seen Images Rather than Previous Memories? Seed1.5-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:32.364515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:07:32.364515Z digest=sha256:8e6ffe29daa80b03e3e749675e66dd5fd220775a7cbe7a9d48d751ebc7a8a5e8

Observation e29e291a-9035-460e-a735-6373514e41b5 · inbound

VGR: Visual Grounded Reasoning cites this paper.

VGR: Visual Grounded Reasoning Seed1.5-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:12:14.418167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-19T09:11:00.295700Z digest=sha256:cc6b3474911d99c739d70207b0bb4c833e97d1de4d30baab0b9489159472ff41

Observation ffa2a57e-5e3f-4d75-bc21-1f6fa110169a · inbound

Generalizing vision-language models to novel domains: A comprehensive survey cites this paper.

Generalizing vision-language models to novel domains: A comprehensive survey Seed1.5-VL Technical Report

Reference 299

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:07.103228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:21:07.103228Z digest=sha256:bdf07163f29bdeb7e2f2764e8eaee435e5ba1274c3abd81c7a764ff6130364d8

Observation f92bd1b7-0158-4d54-80f4-72785261939e · inbound

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning cites this paper.

MARBLE: A Hard Benchmark for Multimodal Spatial Reasoning and Planning Seed1.5-VL Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:35.866414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:35.866414Z digest=sha256:8dda4293bd9935a463cb5ba13366c552543ea6d811ae3bb45c4c75e82d8ee238

Observation f811175c-171c-4137-a96b-46026a58ac04 · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Seed1.5-VL Technical Report

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.663603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.663603Z digest=sha256:651cf696ba5e380ae1db5d7b6dc6339c32ad9ad2e57effc3459bcd81043a908b

Observation 31960e2f-eab4-4db4-a130-b6b293d6ee4e · inbound

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning cites this paper.

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning Seed1.5-VL Technical Report

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T04:48:26.355351Z digest=sha256:d1bcd1d3b5f86ed0479cabd1bd1328c1d9b9a78e8956a9fda33e94278ab24f6e

Observation c67bc056-83d2-4264-83f5-d7a24b974d02 · inbound

RoboBrain 2.0 Technical Report cites this paper.

RoboBrain 2.0 Technical Report Seed1.5-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:47:28.125248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:47:28.125248Z digest=sha256:17558524d46fe72949955ff92871b59651e94a5adfd364ca75d0d03fe036eae0

Observation 93cdf003-0ea1-4adb-ae41-92f3d9eff97f · inbound

GTA1: GUI Test-time Scaling Agent cites this paper.

GTA1: GUI Test-time Scaling Agent Seed1.5-VL Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:55:00.062838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T13:54:59.938216Z digest=sha256:57949ff3a15add1e790f0d91db16fd4eb1d9ee8162c37c386947a48c3f1249e9

Observation 595aabb9-7d61-48d2-a3f9-3f6748706de8 · inbound

Skywork-R1V3 Technical Report cites this paper.

Skywork-R1V3 Technical Report Seed1.5-VL Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:03.246738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:14:03.246738Z digest=sha256:8846edc6a706f3f570b55e9adaafd4732d4dbec500376c2df8e58f39715a5cea

Observation 5416e37a-3b9e-424a-b3df-ad990b0b24d5 · inbound

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization cites this paper.

CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization Seed1.5-VL Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:15.299773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:14:15.299773Z digest=sha256:ba75a1414175a7ada672a74d43e18a06a75d91735805aff8445e9a612337ac83

Observation a61bccae-b27c-4896-903c-de4f5ce04ff3 · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report Seed1.5-VL Technical Report

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-17T08:04:12.591241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:d6961046f4a024d676f75cee12178fcc425263865b988459c37e8d6541157202

Observation 0675370e-914a-4442-8602-b557c2ee1f04 · inbound

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback cites this paper.

Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback Seed1.5-VL Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:23:30.295036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:23:30.295036Z digest=sha256:d2fe27ed3aa99bc8997679711f8fe36ab35af89f56a2717d84c5ac8a2e1791aa

Observation 5674d609-ce4d-4451-9ee5-3abafcfda2f3 · inbound

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning cites this paper.

VL-Cogito: Progressive Curriculum Reinforcement Learning for Advanced Multimodal Reasoning Seed1.5-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T11:35:11.788450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T11:35:11.788450Z digest=sha256:18f135ac9ed5c9cfbce7616fea7f752fe448906c41aa4e407ef38d4b1821e118

Observation cbacbe79-cebe-4d20-9405-6e8e68f8c647 · inbound

Diffractive electroproduction of light vector particles: leading Fock-state contribution in the presence of significant higher Fock-state effects cites this paper.

Diffractive electroproduction of light vector particles: leading Fock-state contribution in the presence of significant higher Fock-state effects Seed1.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T05:25:57.361418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:25:57.361418Z digest=sha256:a084fd72630701dcbfdd039d43f8c7775465c52f8c9c1f3077f51863b018c917

Observation 80ad6a69-6d28-470a-a1e7-4609fbb6fd38 · inbound

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models cites this paper.

StructVRM: Aligning Multimodal Reasoning with Structured and Verifiable Reward Models Seed1.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:29:13.766583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:29:13.766583Z digest=sha256:f0975742a17199cc711f93e2342cc9af817d925fd41465a9ce669435e3fb6140

Observation 79f9783d-e128-445a-a1dc-9c3c90ce629d · inbound

An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal cites this paper.

An Explainable Machine Learning Framework for Railway Predictive Maintenance using Data Streams from the Metro Operator of Portugal Seed1.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T23:25:47.893839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:25:47.893839Z digest=sha256:56409f0543f6c4f17b5ad4fea816790c986a3f4dfd47a4e9abeee9c54875df9f

Observation baf71de0-f6fb-4fa1-b7ac-c21f72141e0a · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Seed1.5-VL Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.466198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.466198Z digest=sha256:579a386688efd13d7f1fa8936aadaa578f02837ecea020f6ee2ba390b2f081ce

Observation 00ceeeda-b9c8-4dc3-82ac-78f64201205f · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Seed1.5-VL Technical Report

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:397f79f47b779775ce06fe09b486a0143712503017d89f1d284e51465854928b

Observation 6a21393a-2d07-49a2-841b-36ec9af8bfe8 · inbound

SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control cites this paper.

SWIRL: A Staged Workflow for Interleaved Reinforcement Learning in Mobile GUI Control Seed1.5-VL Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:52.028713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:18:52.028713Z digest=sha256:88fc8e6a0c27e28668872fd539805d55c2ba88534a421256396d364624217821

Observation 08b242fd-5c6d-43ba-a046-fa0521f1fdcc · inbound

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning cites this paper.

R-4B: Incentivizing General-Purpose Auto-Thinking Capability in MLLMs via Bi-Mode Annealing and Reinforce Learning Seed1.5-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:21.036423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:42:21.036423Z digest=sha256:7f9f9b957189f7bed218694fc6791d1764219536ed5dd2ac6ff9c0194fceddc0

Observation 28ade166-9d20-4a20-a2c6-5e656d96b080 · inbound

UItron: Foundational GUI Agent with Advanced Perception and Planning cites this paper.

UItron: Foundational GUI Agent with Advanced Perception and Planning Seed1.5-VL Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:35.711374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:35.711374Z digest=sha256:034c5cc32ba37b8cc17dcbd3e573dfc09df9a6ed6e15edd54b59780812bdeab3

Observation cef2c61f-7ab2-4e5c-9639-9b34512f8f88 · inbound

Robix: A Unified Model for Robot Interaction, Reasoning and Planning cites this paper.

Robix: A Unified Model for Robot Interaction, Reasoning and Planning Seed1.5-VL Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T12:59:04.402669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:59:04.402669Z digest=sha256:6e24d8fb774cff48fb9678e85f237aa13ffe3f951853ef9462236d3209dfac7a

Observation af8abafb-9638-4e7b-b7ad-0ef1fd8de52c · inbound

Kwai Keye-VL 1.5 Technical Report cites this paper.

Kwai Keye-VL 1.5 Technical Report Seed1.5-VL Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:26.527163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:26.527163Z digest=sha256:3dec2f95cf3f701aaf78a8a6f2c6c78787b014039300268965e9f812858c66c0

Observation 9ae68a05-9103-4eaf-8053-81e0dc4ad9bc · inbound

Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning cites this paper.

Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning Seed1.5-VL Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T20:28:58.597765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T20:28:58.597765Z digest=sha256:f6010be84d2db051814b7481bca61ce997c2e0ebedc59cb8e4abe8b479e97821

Observation 9f6d2274-fe55-41f4-8389-1e9fe22a5061 · inbound

Seedream 4.0: Toward Next-generation Multimodal Image Generation cites this paper.

Seedream 4.0: Toward Next-generation Multimodal Image Generation Seed1.5-VL Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-12T16:39:00.920258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T16:39:00.828442Z digest=sha256:ebd1b3eae0659fd8e69f28db80b9ba46af74389a236ca7ab9ee0bab9b224bf51

Observation 4e95f10a-35a2-499a-9591-3c343fe030d8 · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing Seed1.5-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-17T13:25:31.987859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:1f87a440ad1d742fa0e433796a901f47ea007d005f96116a2b1aa347fbc6109d

Observation 0ccb065e-fc53-4243-8d35-1912f4ee181b · inbound

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training cites this paper.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Seed1.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:53:26.633881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:020f60b771cd68243d4820e12d22b0d42f5201beb99dac28d1d372f65472f707

Observation d2a72923-79c9-4ae0-8141-4550b7e18f87 · inbound

Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline cites this paper.

Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline Seed1.5-VL Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T13:38:36.451668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:38:36.451668Z digest=sha256:b355d3bcd21ae599e1ed24d192174bdbf7d0c1f49fdad1ae91036b9da43f0c5e

Observation c6b6c83a-6dca-4ec9-aed3-b060ce150775 · inbound

SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents cites this paper.

SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents Seed1.5-VL Technical Report

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T08:16:06.590003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T08:14:51.102085Z digest=sha256:188d49a87a97ec2e5f3a1785281d885a72b38efa588b665a049fd07175334ef2

Observation 7c10bb39-a4fc-4495-9a3f-ceb07cb31b96 · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Seed1.5-VL Technical Report

Reference 172

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:54.595641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:54.595641Z digest=sha256:e6e4ef7cba2a1b23f70f31cc16541e9808354898614aa2d0784336dfd316f650

Observation 9f752b9d-7723-4de2-aa7b-efa6ad70be72 · inbound

DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks cites this paper.

DSBench: A Comprehensive Benchmark for Evaluating External and In-Cabin Risks Seed1.5-VL Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T21:37:07.589391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:37:07.589391Z digest=sha256:f375f207e8bb9c4ec148ae8f9c6310f0a3cd07ec85a230bfbc967858f0ea4b81

Observation 63e80dff-649b-4887-be08-6a2e5abb1b77 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report Seed1.5-VL Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:42:05.738526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:523d890d2a1db95071ae4904f360c2ce2f7780b5322ab421b69b8b515b7d0bae

Observation 33745277-755c-4e19-8869-9eae62425e40 · inbound

Boosting Reasoning in Large Multimodal Models via Activation Replay cites this paper.

Boosting Reasoning in Large Multimodal Models via Activation Replay Seed1.5-VL Technical Report

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:09:04.026356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T05:05:48.682057Z digest=sha256:187068bb71d8469c94278586d19d6e679f770f6fd0764b6a423df801c2a7728f

Observation 15010b91-8e2e-4448-afe6-bf85f8207e74 · inbound

OneThinker: All-in-one Reasoning Model for Image and Video cites this paper.

OneThinker: All-in-one Reasoning Model for Image and Video Seed1.5-VL Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T02:11:26.581673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T02:09:39.820651Z digest=sha256:61e6bfcc30d406c3c32905e7e728469920ea9a4c0b9a4aef7b456e7686f166cb

Observation f9e5bdac-24fc-4247-8288-958bce8c4530 · inbound

Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning cites this paper.

Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning Seed1.5-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T17:30:50.912992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:30:50.912992Z digest=sha256:24051229e1b85a515931c252b7c25cbde8c276d7e7a5a84e7496db3d5fa4cdf0

Observation 49d30e0b-b72c-432e-bfeb-24d3a70d46b4 · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models Seed1.5-VL Technical Report

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:31:21.904826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:8bd49f545ee2d332319db2203a2d9473afd98f3b78e965d5d165474db94c247e

Observation 8ef4a5d2-12b8-40ac-9ecd-cd22d91fc50c · inbound

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning cites this paper.

Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning Seed1.5-VL Technical Report

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:44:15.952662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T16:43:11.995960Z digest=sha256:c99376a9573ba54915673f3626a74311ba0bb61d66644bf9d8e50f112d561392

Observation 00b8b6ef-d0be-4677-88e4-d814d681dd06 · inbound

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters cites this paper.

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters Seed1.5-VL Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T14:16:50.766906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:16:50.766906Z digest=sha256:da0a41f78f4133056065fbad14d0ac8da9a0ae4fb64e42d75446311e8bc6e542

Observation 3b153076-3b0f-4b79-9444-d69987996b3c · inbound

CountGD++: Generalized Prompting for Open-World Counting cites this paper.

CountGD++: Generalized Prompting for Open-World Counting Seed1.5-VL Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T13:45:45.266385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:45:45.266385Z digest=sha256:51e10efed316c1302a01a4032d871c78f184e4a5c2bb7c8b1936a6f577c48ba0

Observation bed1ebcf-69f7-4e96-9e5a-9f763a375257 · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data Seed1.5-VL Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:50.244864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:50.244864Z digest=sha256:b5530d331c2956cd1faff3acdb11324f3ce09b566d6516d87c65089d76dee622

Observation 0952cbb7-12ab-4624-af42-810838bc210f · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method Seed1.5-VL Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.145651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.145651Z digest=sha256:67683ac68addbb2eb6a0cf04146affcd69f2d9be207250d6b9f875543dd788ad

Observation 8e15eb89-a3f8-48b1-aeec-ab53b8a4d1ae · inbound

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning cites this paper.

VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning Seed1.5-VL Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:17:51.869714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T12:17:42.135851Z digest=sha256:a865fe5eea7d6eef4f019a85c91c4ff014684f97a60bb7f75521605884c99839

Observation 2ac7a027-5a9b-492d-9e3f-e67bf1daad3e · inbound

Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction cites this paper.

Agentic Reward Modeling: Verifying GUI Agent via Progressive Trajectory-Grounded Interaction Seed1.5-VL Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T06:01:38.768084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:01:38.768084Z digest=sha256:70c6021e0dc4cbcf04319deed2ac3e43245acdd5e228ceb349c494854e6801a0

Observation 77f142a9-63b0-4cb1-a101-7438cfd479a6 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence Seed1.5-VL Technical Report

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:7714569c8cabc523fe8810ca51bec5753031535a2119232ee53dd38326f5cb5a

Observation c9d91b91-63b2-4212-ba47-7455d8f22b9d · inbound

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning cites this paper.

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning Seed1.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:40:42.201090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T06:39:29.010937Z digest=sha256:ee7ea69e3bc66a2fb630a47b092f8d7fa6925e8ab51d042814948d3ffefe2da6

Observation f597c6e8-5eef-49d7-b298-43cfc185f295 · inbound

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization cites this paper.

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization Seed1.5-VL Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:10:43.152968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T07:09:46.254851Z digest=sha256:23d161ab27af0fa7622f08aacf00943d7d1a7f6f818cc7ab59ae1cdc715243cd

Observation d46b1046-30c1-4b3a-9694-0a43c78c9f30 · inbound

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs cites this paper.

MedXIAOHE: A Comprehensive Recipe for Building Medical MLLMs Seed1.5-VL Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:56:50.375022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T22:52:30.992054Z digest=sha256:9f96418bdfbd34c9b457882b611c8a7587eaeb5284f2a48ef1c974555274ae99

Observation 70193c6f-cc14-436f-8edf-b32e7c284204 · inbound

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments cites this paper.

JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments Seed1.5-VL Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:39.921774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:39.921774Z digest=sha256:99709afe3ebb2156beedf9808262c1fdb72191d54b1c0d9f66e95cb2f7f33fbc

Observation 25d4cdf8-7156-44a0-8371-ae0638bea619 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Seed1.5-VL Technical Report

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T13:05:09.941488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T13:04:30.544504Z digest=sha256:a5932705172cd7f4c61ee63ab4f7d783d25bf8740d74b4d56ffe0171255c03b2

Observation 9158f500-ae05-4265-8f33-ca70f58be589 · inbound

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies cites this paper.

PoseVLA: Universal Pose Pretraining for Generalizable Vision-Language-Action Policies Seed1.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T21:37:11.438697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:37:11.438697Z digest=sha256:2f013aac531b9ea0165af8e5555ac6285dd9cead08a892516238c262f0a1f521

Observation 335b71fd-9a91-4702-bd6d-6032358880bd · inbound

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation cites this paper.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation Seed1.5-VL Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:28.590628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:28.590628Z digest=sha256:09ec2b9d5c2f7ebeb53efd072836c1431b5f45f7aca8153fb0338a853eb3bd51

Observation e4f9eef8-ed59-4106-b25e-0304f4fb151c · inbound

CodePercept: Code-Grounded Visual STEM Perception for MLLMs cites this paper.

CodePercept: Code-Grounded Visual STEM Perception for MLLMs Seed1.5-VL Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T23:22:13.847876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:22:13.847876Z digest=sha256:8355cff86408f3552148ef3a5552134d9bada16030418051a49654da921a0e16

Observation c1cbe2ac-e98a-42db-8740-f94b9316b201 · inbound

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison cites this paper.

AD-Copilot: A Vision-Language Assistant for Industrial Anomaly Detection via Visual In-context Comparison Seed1.5-VL Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T11:55:33.279954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T11:54:18.587529Z digest=sha256:d9cf211e377a2d79dd39a153e16c675ba5e5a6fe04557f73c45bc0f45cc33dc3

Observation 705d5836-de75-4a73-a082-ccc5d2e8c2e7 · inbound

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP cites this paper.

Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP Seed1.5-VL Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T23:59:10.664837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:59:10.664837Z digest=sha256:f9fa194bf9213c26f807e0aa908b7ff96296e6cc2b6f87a011a178b840574f08

Observation 5d7687b2-011d-40ca-84c1-4979940d7d04 · inbound

Peel neighborhoods cites this paper.

Peel neighborhoods Seed1.5-VL Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T20:08:04.785209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:08:04.785209Z digest=sha256:502340be4707e3da660ba54ebba97faf67572f1120833ad9965e59cd5a3a2feb

Observation 20a792e4-84d9-46c4-892a-1ca8771ec2b4 · inbound

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios cites this paper.

TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios Seed1.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T17:05:13.857975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:05:13.857975Z digest=sha256:2ab70295501ef489968043013a0cdf1477a3dff4489f43be68f9ea5f422f679b

Observation 684d642b-a880-4338-a93b-9e3415a180a9 · inbound

InstructTable: Improving Table Structure Recognition Through Instructions cites this paper.

InstructTable: Improving Table Structure Recognition Through Instructions Seed1.5-VL Technical Report

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:58:16.051935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T20:53:57.029294Z digest=sha256:b67cefd851506649ceb4cd17dea5ead114a094b2791f2c28573b4d556f4f29ee

Observation ca1359cd-6476-4fd8-96b4-b15426c2db73 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Seed1.5-VL Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:feea97e5bb35c6909f10dcc98cf7ae9082f74c9f76d25add42f3b1661cbc8784

Observation 9f7f9bfe-e4d0-46cf-86c9-40be3fad9562 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Seed1.5-VL Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:6c8696c6d5b4c83642b9d1b27cfcbeca62e78407636fad4b7c4b4df8c67914ec

Observation 51bd245e-7b26-474a-af68-f51e37231b67 · inbound

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization cites this paper.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Seed1.5-VL Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:a6f537b69043f21356c28d1cd1ab32043342353f721c69165f5ccb6754c602c2

Observation 4729d643-ea08-4e50-874d-f574b59a4bc1 · inbound

OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence cites this paper.

OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence Seed1.5-VL Technical Report

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:20:56.579502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T17:41:31.544716Z digest=sha256:4ed74b058f791102174b25db999248aef5690cbaf8d176112dc6353ff6fbfc4a

Observation c15ddaba-36c1-4d49-9990-edd8358fce2e · inbound

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents cites this paper.

GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents Seed1.5-VL Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:46:07.941683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T17:57:36.038091Z digest=sha256:94991f8da3200e3681d7e1c8eedd50857ba3f6cc3032df9823d25f4f63fca9a6

Observation 1313a7be-ad69-4774-a5ac-00d1a6eb1146 · inbound

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding cites this paper.

Bridging Time and Space: Decoupled Spatio-Temporal Alignment for Video Grounding Seed1.5-VL Technical Report

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T18:38:16.204012Z digest=sha256:ebadc885a41bd4cb5320a7a0918d6ce6735fb0b852e18510389fb1a735757e1b

Observation f6083b10-d61c-459a-a5fa-761cc5435c5b · inbound

LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation cites this paper.

LAMP: Lift Image-Editing as General 3D Priors for Open-world Manipulation Seed1.5-VL Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:50:57.521289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:58:06.892531Z digest=sha256:cfd7dd58f88da97db400ab017cd83e783ffe88039aabdc26066939f89775570d

Observation ac4bdaa8-d253-4395-9cd6-0c806ee9d2ce · inbound

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration cites this paper.

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration Seed1.5-VL Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:41:01.158961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T17:02:05.265029Z digest=sha256:450e3cfc0890e4e15c6eb69c7180f49a529c13751f7be11c5155776f485d9e98

Observation 9911c8e4-65b5-4c44-888e-b94494bbab64 · inbound

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration cites this paper.

EpiAgent: An Agent-Centric System for Ancient Inscription Restoration Seed1.5-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T23:17:40.794140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:17:40.794140Z digest=sha256:6a514b6ba157d12a5bc4b9abb8a73ee945272e8f05ec46a148abf03e34ed1af9

Observation dc8eea66-cc36-4ae1-9e8e-f1e993ef632a · inbound

Towards Realistic 3D Emission Materials: Dataset, Baseline, and Evaluation for Emission Texture Generation cites this paper.

Towards Realistic 3D Emission Materials: Dataset, Baseline, and Evaluation for Emission Texture Generation Seed1.5-VL Technical Report

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:01:05.280745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:12:42.995562Z digest=sha256:660e02508a8aa5c208bfc95f1c3f863b1032345e00833b3c58102679b6bde3e9

Observation c343d3b9-ca6a-4f4e-acf0-5081ed2d6f6e · inbound

CLASP: Closed-loop Asynchronous Spatial Perception for Open-vocabulary Desktop Object Grasping cites this paper.

CLASP: Closed-loop Asynchronous Spatial Perception for Open-vocabulary Desktop Object Grasping Seed1.5-VL Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:20:59.855359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T16:42:02.271796Z digest=sha256:9cf9cab74b14dfc2d837b8730125912d66ba9c34974eb8a7149ab49db5e6742f

Observation bb5762b0-f8a7-4249-8fdd-0af5a468e403 · inbound

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs cites this paper.

POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs Seed1.5-VL Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:41:04.089096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:23:08.671342Z digest=sha256:2e5d819c5aa23905220674421575706322b872b681264f5c7714360493e63f7d

Observation cd114e61-0326-430a-b985-10dfaf86ebbb · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Seed1.5-VL Technical Report

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:11:04.630482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:9a77e878746be6b53d8f011e1f19fa6bd3dda6452d16de730c8a50e4de074aa1

Observation 1ad42240-2d80-4bc4-919a-06069a60de67 · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management Seed1.5-VL Technical Report

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T13:09:24.304696Z digest=sha256:4714d973023b10a855cfa5443d24591296b703d81695cb1c63ce441889207a5f

Observation 8b8c1bf1-2dd6-48d3-b435-21b9c24ecb7c · inbound

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management cites this paper.

POINTS-Seeker: An Open Recipe for Multimodal Search Agents with Visual Memory Management Seed1.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T16:18:20.921262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:18:20.921262Z digest=sha256:362bd2c8b53fe47f8102d9f736f06ea0ed1feacffa20826330929869ed997450

Observation 99aedbb3-c5f6-44aa-9dfd-c2bc4e0c655c · inbound

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding cites this paper.

UI-Zoomer: Uncertainty-Driven Adaptive Zoom-In for GUI Grounding Seed1.5-VL Technical Report

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T14:06:55.472857Z digest=sha256:e0c5776ecc76c7ca2a40e45db2d9780bdcb79e81c7635fa422223abe18ca3c52

Observation 30c6fd44-077b-4ec3-a935-a855b3328712 · inbound

Seedance 2.0: Advancing Video Generation for World Complexity cites this paper.

Seedance 2.0: Advancing Video Generation for World Complexity Seed1.5-VL Technical Report

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T13:34:36.248186Z digest=sha256:7f2b746b3275d5942b7d16cfc745e3c0af9d583e141774fc8de04c985c55b5c9

Observation 9fb25da2-9171-4cab-a7b5-87b2b715a864 · inbound

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow cites this paper.

Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow Seed1.5-VL Technical Report

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T09:02:09.075097Z digest=sha256:1f78a97eb6f829dbdd95dccd781875a7c7d6d796f1b128a5b529af42e46a07ed

Observation b4feee8a-4cc1-4113-89e0-628590d048e3 · inbound

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior cites this paper.

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior Seed1.5-VL Technical Report

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T07:21:45.633310Z digest=sha256:6626dca51320c3a44baab1f4b8ccf29b6d6a0732188ba6ab503e1872dee6251d

Observation 07575c65-6c00-4c56-98c5-3146de1aebdb · inbound

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning cites this paper.

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning Seed1.5-VL Technical Report

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T06:31:30.778309Z digest=sha256:f84108d6dd02ffd8ce69ca4bfac9a3062dbb157ae1ad0e37803884f3115fe95b

Observation 5f3ec7c3-edf9-44e2-9199-03e7af082812 · inbound

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation cites this paper.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Seed1.5-VL Technical Report

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T05:42:41.112158Z digest=sha256:a590cfcdb650a219ff3702452f09849fc52e22bba68047bc06d7f83529fe5bc0

Observation 1298f9e2-02e3-458a-af07-8c4ef8c99a56 · inbound

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation cites this paper.

Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation Seed1.5-VL Technical Report

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:25.937007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T00:54:23.508845Z digest=sha256:4cf1bd1eae0493e28f69e7d06ae90cdb9ccb15a463b7a70e9a8b28c06231ecad

Observation 3e3c3bf8-286b-42b4-a6d5-be387d7f8822 · inbound

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence cites this paper.

Wan-Image: Pushing the Boundaries of Generative Visual Intelligence Seed1.5-VL Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T13:11:04.111222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T02:16:03.854650Z digest=sha256:ea000ba3e563bda30d6428a428f96ff492fa7062a2a7df1d769d22b6dd977723

Observation 7deaca6e-06fe-4c94-bbc3-8133b4d2ba07 · inbound

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding cites this paper.

Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding Seed1.5-VL Technical Report

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:05.940196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-09T23:05:05.251150Z digest=sha256:129cbd2b663cc19d23ba24203c9635e2cd833c69005eba18f5f750fab667d82f

Observation 706b69ac-33c7-4e0e-86ba-b603988810b7 · inbound

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model cites this paper.

dWorldEval: Scalable Robotic Policy Evaluation via Discrete Diffusion World Model Seed1.5-VL Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T19:31:09.545993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T11:45:18.081248Z digest=sha256:5de17c463e31d3ad378b886c45c1dfcde3d1e0034b2bd59bb3f654e76c723e38

Observation a9231b2d-39e1-4c77-9296-b322c8ebba8f · inbound

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs cites this paper.

SMoES: Soft Modality-Guided Expert Specialization in MoE-VLMs Seed1.5-VL Technical Report

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:41:13.678645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T04:41:52.098355Z digest=sha256:1de389d98af237fe9a72ba06a34ed652859db45f3a238d9735bad1d610e73097

Observation 22df99a6-c8d7-4bcb-8dbc-02cb5f53f4ec · inbound

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection cites this paper.

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection Seed1.5-VL Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:36:17.378085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-08T04:46:16.497585Z digest=sha256:f84b24a74841a2f09070926fc90483f4c752863181a37d72ed74db8ce9ecae66

Observation 7fb67bc8-972f-4fa8-a465-cab0cb1f9ec4 · inbound

Benchmarking and Improving GUI Agents in High-Dynamic Environments cites this paper.

Benchmarking and Improving GUI Agents in High-Dynamic Environments Seed1.5-VL Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:31:13.993200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T16:55:29.332273Z digest=sha256:5d81e7b22fbafa37cb4e1ebf59250c179d2821852af823e399e93ccd0bb4f3fa

Observation 7d8bdbba-e9e2-4c21-af96-42e6b75f2e53 · inbound

Benchmarking and Improving GUI Agents in High-Dynamic Environments cites this paper.

Benchmarking and Improving GUI Agents in High-Dynamic Environments Seed1.5-VL Technical Report

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T00:54:49.351703Z digest=sha256:4697049b5636d1296250e1f523a9d31ad80397113be64a7bef86d7388419bc5f