Pith. sign in

Paper Citation Record · LEDGER

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

As of 17 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 15 inbound Pith citation observations for arXiv:2506.05302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05302 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:28:14.907661Z

measured 104 of 104 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:51:24.717294Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:27:15.782237Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy38
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5e6c0141-9f53-4a7a-9fb3-8161ec7a3f0d · outbound

This paper cites Mc-llava: Multi-concept personalized vision-language model.arXiv preprint arXiv:2411.11706, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Mc-llava: Multi-concept personalized vision-language model.arXiv preprint arXiv:2411.11706, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:07.834487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:07.834487Z digest=sha256:e4b0bb1ac6ec75be30cbec8609a5d9d6bf1d6314624bb9e916e6199586ec614a

Observation 9caa0b70-e2a3-4d5d-a811-fa791f80c4c2 · outbound

This paper cites Qwen2.5-VL Technical Report.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:07.986779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:07.986779Z digest=sha256:166ded9338bc7d35f1f6de9ffb88a8f4e1820254c10583078726f9faf8751511

Observation b44229d7-89fb-458d-934d-5bdf41176fcd · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.106718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.106718Z digest=sha256:556328c950a1d98e71c96337e32ac6183bc20962c26cb5e076b9ae405ee78d4f

Observation 78c355f2-42bb-4a4e-886c-ad97de098a34 · outbound

This paper cites Abductive commonsense reasoning, 2020.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Abductive commonsense reasoning, 2020

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.211037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.211037Z digest=sha256:715043d78614b2ad2cdd23802cd1d883c46e7fbaf22c4c463dc2f9ec188c2397

Observation 53317ec5-f5f6-4ddc-b81d-c892bb967dcc · outbound

This paper cites Graph cuts in vision and graphics: Theories and applications.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Graph cuts in vision and graphics: Theories and applications

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.339094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.339094Z digest=sha256:af3032b4af258a05a4e4931df5f998049d43c65266a4740428e734774addabf9

Observation b074da60-bfa4-4d8e-9662-7c15f588d927 · outbound

This paper cites Scene-text oriented referring expression comprehension.IEEE Transactions on Multimedia, 25:7208–7221, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Scene-text oriented referring expression comprehension.IEEE Transactions on Multimedia, 25:7208–7221, 2023

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.465002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.465002Z digest=sha256:032d4d9a191c3e08cb14f72085ff2ddaddaa95d117b75de809a854d2b6a47e57

Observation b8625921-7c21-4aa7-9a75-856e626b190f · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Activitynet: A large-scale video benchmark for human activity understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.647907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.647907Z digest=sha256:1f954058f12c898392274ec1ce0142b01578cb3fcb6791d62763de84efe9b399

Observation 6d1cb01a-b711-45b9-835b-124c53b2aa96 · outbound

This paper cites Vip-llava: Making large multimodal models understand arbitrary visual prompts.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Vip-llava: Making large multimodal models understand arbitrary visual prompts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.783473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.783473Z digest=sha256:d1fd37f6287dc42eda336ca2981eb5439293946652d229340b8be7645de141a5

Observation 9dac110d-8043-47ee-ae47-2bbdb3a948b2 · outbound

This paper cites Active contours without edges.IEEE Transactions on image processing, 10(2):266–277, 2001.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Active contours without edges.IEEE Transactions on image processing, 10(2):266–277, 2001

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:08.952935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:08.952935Z digest=sha256:1a93a34675771a5abdb8eb1695316fcc1c093e042cfb1f66b999bddaaae12bb5

Observation 8a0a7662-5adb-4947-9e25-a5fee112ded4 · outbound

This paper cites an unresolved cited work.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:09.089611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:09.089611Z digest=sha256:a76bafd015f6dfdc2de16a1f48dd03a7deffd12bb173f5274a95d2a0d62b1140

Observation b362ae58-7c9a-44a4-9249-8f628c09b56a · outbound

This paper cites Videollm-online: Online video large language model for streaming video.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Videollm-online: Online video large language model for streaming video

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:09.157074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:09.157074Z digest=sha256:e765fac3e6d72be2c68f3e733ec71fa89c63981f78ea03b65e85d268cc44c345

Observation 3d57f293-dc7e-400f-8b59-ed0e79f1b019 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:09.209098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:09.209098Z digest=sha256:f64a0f98e3f55a6db22bd2e51e11bd8f11c530f4bdea029b502e4cd0329bc4af

Observation 917015bf-08bc-43be-87e3-d8b9bc2d7aa5 · outbound

This paper cites Segment and Track Anything.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment and Track Anything

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:09.287312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:09.287312Z digest=sha256:c4d9037e55a250a5ccf1e6943a21bb7a6c614dd345cc029771c08b8004c67af1

Observation 000a4c48-a6f7-46bc-8812-11e99f4309c4 · outbound

This paper cites Total-text: A comprehensive dataset for scene text detection and recognition, 2017.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Total-text: A comprehensive dataset for scene text detection and recognition, 2017

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:22.167513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:09.366044Z digest=sha256:ebda7c335db4773a3d79e0ceab0679b13d0f9dd0689c7fef16e53671f8314f4d

Observation 9e559b13-e1bc-4755-b972-952d73d0b449 · outbound

This paper cites V ocabulary-free image classification, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos V ocabulary-free image classification, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:21.982352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:09.461655Z digest=sha256:518c54860c92dc7c9009886859141b7da6f87bc43033510dc28033c625e359bc

Observation 6bccd78c-1e66-4279-9cda-1f1ba90c107c · outbound

This paper cites Online action detection.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Online action detection

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:21.636264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:09.531677Z digest=sha256:a033e0fed74a7cb405ad7fd478aff5876893d25d01090b995630d31506d683ec

Observation 90bd38f4-2b46-4711-8c71-cd433798ce8c · outbound

This paper cites Mevis: A large-scale benchmark for video segmentation with motion expressions, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Mevis: A large-scale benchmark for video segmentation with motion expressions, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:21.160819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:09.622757Z digest=sha256:4dfda8b332aeee85ddb78ecaf54a7d9c665cc2963e1c333a73b9c83ec96d6529

Observation ec455c19-e304-4fd8-a4fd-c922bb6e03b0 · outbound

This paper cites Actor and action video segmentation from a sentence.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Actor and action video segmentation from a sentence

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:21.019625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:09.710833Z digest=sha256:3207d34869a701a557415ae9093ace0b2ad9a7e94e4b7ed0310f5e1cf0afddd7

Observation 23fc4dd0-9b5d-4320-828c-a68112c74745 · outbound

This paper cites Icdar2017 robust reading challenge on coco-text.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar2017 robust reading challenge on coco-text

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:20.857060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:09.778550Z digest=sha256:dc04c03d1963e7cf31e0833897c0df510dc27870e2a6d7b5f46f7af888a07474

Observation 93ee6ab1-f238-48e7-bb87-02d288870d6f · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:09.831040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:09.831040Z digest=sha256:49a057bb25f9d78e49129e5d4ca82e4d20668436afa123dbc2a575a62251710f

Observation 0fccdb72-31f9-4b8f-8922-eae20a984627 · outbound

This paper cites Regiongpt: Towards region understanding vision language model, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Regiongpt: Towards region understanding vision language model, 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:20.698363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:09.953283Z digest=sha256:f3d1b6a33e57931bfd2c2fae84148f4c36bff5b01a65ae2bbca4639fb4819d58

Observation 8e817e83-8877-43fe-bf03-842cf459c6cb · outbound

This paper cites TRACE: Temporal Grounding Video LLM via Causal Event Modeling.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos TRACE: Temporal Grounding Video LLM via Causal Event Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.068680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.068680Z digest=sha256:6a585fecd2a733db95b1a5b63d91cce46b53b3a86c97e1f3829c54bbb19cffba

Observation b1a95b82-b009-4ffe-87f4-a71ce956a7b2 · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation, 2019.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Lvis: A dataset for large vocabulary instance segmentation, 2019

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.160396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.160396Z digest=sha256:d6524d37b56497d523bb1d977c6bedbbebd1b2600baac19a5859b145002b2818

Observation bf17f7ba-2721-403b-921c-fcb3e4933269 · outbound

This paper cites Synthetic data for text localisation in natural images, 2016.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Synthetic data for text localisation in natural images, 2016

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:20.526290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:10.245238Z digest=sha256:aa6050de75d00eb3b1d9c59c426060fe33cfe6b98abc6dac21df16a216ff695d

Observation 9f6b5a53-9742-454a-9744-3d2693faf534 · outbound

This paper cites Omni-rgpt: Unifying image and video region-level understanding via token marks, 2025.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Omni-rgpt: Unifying image and video region-level understanding via token marks, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:20.368507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:10.341914Z digest=sha256:c76fa8389d1b561d2e94320b4847e6287a75f3d290b63a8ca9b126c3b6362593

Observation 329c1ee0-6d7c-4793-b219-11a137c499d1 · outbound

This paper cites Segment and caption anything.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment and caption anything

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:20.174249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:10.442313Z digest=sha256:beb242cd7cbc6c9075888435fdc2e2c437f015168ed76d796210b6933a11ecc5

Observation 8f5ee2b8-10b5-4d15-9577-c1b75af42a49 · outbound

This paper cites GPT-4o System Card.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos GPT-4o System Card

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.544501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.544501Z digest=sha256:0f2104d5d22b6f2fb9c60124f7e2ffa80f4d61d790cd06345231b050091082b6

Observation da1fde41-df61-47a4-b776-fe83df5a39a0 · outbound

This paper cites Visual prompt tuning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Visual prompt tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.660851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.660851Z digest=sha256:604d2daf3db1fb546db1b5d75e3dc9307818c8655b41ba35421bbbae2370c0df

Observation 7e4fafc4-c40b-4f2c-abba-fae1b630d798 · outbound

This paper cites ChatRex: Taming Multimodal LLM for Joint Perception and Understanding.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos ChatRex: Taming Multimodal LLM for Joint Perception and Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.735473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.735473Z digest=sha256:c83b26753a513635408a5befbdf7d5a7d3902556dcb72d287dbe863be136f7a5

Observation 0320ae6a-736d-466a-9137-320cf8b0beb0 · outbound

This paper cites Icdar 2015 competition on robust reading.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar 2015 competition on robust reading

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.977543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:10.822007Z digest=sha256:7798ee9b5a557e699a1fce4df538d4a8cb7bdc56c925304df8c2e13a5507e1f0

Observation 79fca956-3d36-4d13-8c15-fc1fce4ea5a4 · outbound

This paper cites Icdar 2013 robust reading competition.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar 2013 robust reading competition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.819618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:10.902744Z digest=sha256:ff7fc3b54defbf711a147bec8afd9dee69cab2f59d2d06c7ebc33c4d3c6059da

Observation 75fff697-834b-493d-9c0e-ebd32d1da19c · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Referitgame: Referring to objects in photographs of natural scenes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:10.990032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:10.990032Z digest=sha256:04f62fb73bd31835ce6394a4f2e460ff4419a0dacdc37a4e28061ae557bc494f

Observation 7dd22ac1-8d1d-418a-872d-1c1d175764ca · outbound

This paper cites Segment anything in high quality.Advances in Neural Information Processing Systems, 36:29914–29934, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment anything in high quality.Advances in Neural Information Processing Systems, 36:29914–29934, 2023

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.093403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.093403Z digest=sha256:507c6e696578823a5d14fc3b1d6b16e8b93c5b41e121e721bed0e6a9748a9ef9

Observation 9594d68a-989f-41aa-835f-5303e38fd1d0 · outbound

This paper cites Segment Anything.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Segment Anything

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.178747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.178747Z digest=sha256:b321095e193c65f60f7010d8a1a6f8c53ca98aa1bb0dad324ba7821c372e1828

Observation 02a13c13-2da8-45f8-bc92-16ded190fae2 · outbound

This paper cites Openimages: A public dataset for large-scale multi-label and multi-class image classification.Dataset available from https://github.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Openimages: A public dataset for large-scale multi-label and multi-class image classification.Dataset available from https://github

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.690256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:11.254087Z digest=sha256:23493d24d007b66c9ce6ede1d9b18a5f89a240eb651939abd96c6ceeac626f25

Observation dbc8b219-ba86-4eeb-a622-83fad787060b · outbound

This paper cites Shamma, Michael S.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Shamma, Michael S

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.558320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:11.338922Z digest=sha256:4d45573da679c37f71838cf9acf1e74a402b9b325baf572651ffe4f2087185c3

Observation 2343811e-b4e8-4161-99c3-e1eed22d929b · outbound

This paper cites Beyond mot: Semantic multi-object tracking.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Beyond mot: Semantic multi-object tracking

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.413523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:11.411250Z digest=sha256:81a3a32ecb52364da39c148500705268c975d371abf4f5c28c0317b53a03859c

Observation 3de15e83-db5f-4f6d-8e47-8de78f50c331 · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Describe Anything: Detailed Localized Image and Video Captioning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.485753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.485753Z digest=sha256:050113b642aadff44dd2c500a1c565c6aa8d7fb4ac1621afb5ea6ea89584b9c5

Observation a5acc589-816e-4212-9a41-05c03a5b85cc · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Rouge: A package for automatic evaluation of summaries

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.573027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.573027Z digest=sha256:90ebb598cea8d5c077f2fdbfa077d1a51f38b2c123f133cd4b1e4f06531b5e77

Observation 77c18e82-2491-4fb3-a229-e0f3415129a5 · outbound

This paper cites Lawrence Zitnick, and Piotr Dollár.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Lawrence Zitnick, and Piotr Dollár

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.641278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.641278Z digest=sha256:be95b0555f9729b6486516740a68c639c1fc73d2ae8699c4d9c24ef6e961d498

Observation 734693f4-cdbe-494d-a851-2c3d8c130a24 · outbound

This paper cites Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.715662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.715662Z digest=sha256:06194cfc043e03cb0d1ade1895d43cae551ea79ba7057d79b47a078cfe4ee15c

Observation af9be36d-cf66-4a61-ae61-4bb17b033d4f · outbound

This paper cites DeepSeek-V3 Technical Report.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos DeepSeek-V3 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.759586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.759586Z digest=sha256:c7a1e68bb3da22ff1237a5bc1d0e05835be572b228ddee36b420cd38ef144a13

Observation 9b80ee9d-3b88-48a8-aa56-1bae4df04a53 · outbound

This paper cites Gres: Generalized referring expression segmentation.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Gres: Generalized referring expression segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.842353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.842353Z digest=sha256:36965ce126b103a06d48205b56815583df39ccfe93ceb783b6702916d37b7795

Observation 9626591a-cf6b-4f0e-85b6-962987fec048 · outbound

This paper cites SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos SPHINX-X: Scaling Data and Parameters for a Family of Multi-modal Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.917393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.917393Z digest=sha256:2b9e6846b78b86a205a84d9c31b4e93aa220fa97e41fddd53d85b2b819a7c8c7

Observation d60893a0-c1ee-47ba-8b40-24254752769b · outbound

This paper cites Improved baselines with visual instruction tuning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Improved baselines with visual instruction tuning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.028467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.028467Z digest=sha256:1bff00bbaf104df9e718cb0497e7a11b59df8af5a1eb4f09b633f8bc2037aec2

Observation 2f5a9006-e068-4a88-9059-5b2f4069a2bb · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.090612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.090612Z digest=sha256:39567c75f7a6699e1615255b2964b4386eec50a4f1327fb980dd2ec4241da85c

Observation 40857708-b9b1-43ac-8f63-68cc2bbdbcff · outbound

This paper cites The 2017 davis challenge on video object segmentation, 2018.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos The 2017 davis challenge on video object segmentation, 2018

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.218820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:12.161483Z digest=sha256:788f04761d318364268c76aac5dc9b27da4cc77d7a69b30731243d25047a35e4

Observation dca6c900-6108-4828-b8be-4fda60a8bde5 · outbound

This paper cites Artemis: Towards referential understanding in complex videos.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Artemis: Towards referential understanding in complex videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:19.045839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:12.252891Z digest=sha256:d58f59957fdcd4df2834183728a966119bd0d326fc3891e8dfe31357589daa33

Observation 0411d034-bbf6-4397-8a7b-2527829345fb · outbound

This paper cites Artemis: Towards referential understanding in complex videos, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Artemis: Towards referential understanding in complex videos, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.841719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:12.318885Z digest=sha256:9bd7723703dab1bcd7a9f69a1d801a304ce7ed69cbbbc6e86ad9d0f2dc860a94

Observation 615de475-4912-4b52-843a-076e6b94c72f · outbound

This paper cites Paco: Parts and attributes of common objects, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Paco: Parts and attributes of common objects, 2023

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.669927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:12.398651Z digest=sha256:18eb3ef5535e2a77dcf7ab57b547f774ead871b4c80a118eb1dd69975b0c81ee

Observation 3b52654c-8e87-4765-9961-b934d46bc038 · outbound

This paper cites Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Anwer, Erix Xing, Ming-Hsuan Yang, and Fahad S

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.438893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.438893Z digest=sha256:b41a6fef97187fae22d9a20c5abfca639aa093c6830b08d9b88d1faecd8c6937

Observation 80503a7d-ce6a-4a73-abf5-43bb8717d469 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos SAM 2: Segment Anything in Images and Videos

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.549714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.549714Z digest=sha256:4fd47daecbce9af78dac57b851d10d48cea8679f94e71f5171c0fb0bf0c2edc5

Observation 77d3f82b-7679-4289-a451-10f39f150192 · outbound

This paper cites Sam 2: Segment anything in images and videos, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Sam 2: Segment anything in images and videos, 2024

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.604590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.604590Z digest=sha256:77fa6ce53c9ee515cff4412b45caf02f481eb45742c76af7762709cac7ab6a4a

Observation dc7a87ef-9de9-4a0f-a155-21f4c3590af6 · outbound

This paper cites Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.695638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.695638Z digest=sha256:ce03f31fa54441349d22258eaac85ae26143382852b1f3afd47f29b6fd764dd2

Observation 53fdc3f9-85b0-41cd-9010-687b9512139e · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Objects365: A large-scale, high-quality dataset for object detection

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:12.752865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:12.752865Z digest=sha256:47dacb2b7bf2ddacd323fb799705c1a0d2ac0de8f485c38f4f0966aa32f080a2

Observation bd84cbc5-8476-4f71-9d32-a5ca9a7245e6 · outbound

This paper cites Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text, 2021.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text, 2021

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.501128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:12.830981Z digest=sha256:a74456c02e876b46ce817c11811fed55e2305851d9abdbcc40a8d53d41365fd6

Observation 003298e2-5e3e-4286-a1da-025dd32bb251 · outbound

This paper cites Icdar 2019 competition on large-scale street view text with partial labeling – rrc-lsvt, 2019.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Icdar 2019 competition on large-scale street view text with partial labeling – rrc-lsvt, 2019

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.357067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:12.899875Z digest=sha256:21eaedf7c94f2fe9f377bac50dd16be54899e22b2ba22621eea219e848faa402

Observation 97ae4dcd-4b1e-4f1b-90fd-60373548ef85 · outbound

This paper cites Human-centric spatio-temporal video grounding with visual transformers, 2021.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Human-centric spatio-temporal video grounding with visual transformers, 2021

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.211229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:12.974590Z digest=sha256:9b985099d68414396b8a0a85f5a5b067aac6aaba1e468f33ed6b4297ee5e0a3b

Observation f81eebbf-4a2b-4617-958f-4e2a0ea98be1 · outbound

This paper cites Human- centric spatio-temporal video grounding with visual transformers.IEEE Transactions on Circuits and Systems for Video Technology, 32(12):8238–8249, 2021.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Human- centric spatio-temporal video grounding with visual transformers.IEEE Transactions on Circuits and Systems for Video Technology, 32(12):8238–8249, 2021

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:18.042690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.075724Z digest=sha256:6056d145ce734e640499ad73ef51b0f391a4c3ae9aca44d590290568c3e9adb7

Observation a6c7bb8a-736d-4403-b5a3-847e4dbdc50a · outbound

This paper cites Cider: Consensus-based image description evaluation.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Cider: Consensus-based image description evaluation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.106006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.106006Z digest=sha256:e6d2114458f22dcdf7d905ddfd2715594525d45547b54e3ab87b857f2e9b0e40

Observation 6538d8ce-016e-4755-9409-05c67de2a623 · outbound

This paper cites Coco-text: Dataset and benchmark for text detection and recognition in natural images, 2016.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Coco-text: Dataset and benchmark for text detection and recognition in natural images, 2016

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.871637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.197387Z digest=sha256:27a9cc439b86d0a069ff872c980ba1767eb277fcf5860b5a18fc464cab3d1d34

Observation 2be0f5f2-0406-4073-b936-c8820e72116c · outbound

This paper cites Elysium: Exploring object-level perception in videos via mllm, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Elysium: Exploring object-level perception in videos via mllm, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.712210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.272446Z digest=sha256:a9ef8c915990cebce638c7397f09fd727294f3da67b3c2ec2c975be3a19a50b9

Observation 86242f7d-353d-4fce-9401-a217cb7e31d1 · outbound

This paper cites Elysium: Exploring object-level perception in videos via mllm.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Elysium: Exploring object-level perception in videos via mllm

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.544629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.376288Z digest=sha256:16990e4215b4c713a615aaa3947b7ba55ef3f732210d8467a6b7f6c73c3f10d0

Observation d27a146f-4d23-4b7a-8f8a-50063baddd19 · outbound

This paper cites Towards open-vocabulary video instance segmentation, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Towards open-vocabulary video instance segmentation, 2023

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.380342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.408456Z digest=sha256:e24becaa8016f47179e0668a01803bbcce16a78445dfac645819b63b3386acb7

Observation 24ea961e-be69-4b03-8336-d84a91f841d6 · outbound

This paper cites Git: A generative image-to-text transformer for vision and language, 2022.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Git: A generative image-to-text transformer for vision and language, 2022

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.430957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.430957Z digest=sha256:8bfa1e8dde579456a0788eb0afe9bdaac64ef96841207b55e7d9098822013052

Observation b4f51c45-574e-48be-87e2-b80093e76b7c · outbound

This paper cites V3det: Vast vocabulary visual detection dataset, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos V3det: Vast vocabulary visual detection dataset, 2023

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.243065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.449740Z digest=sha256:abbbf525a346e0b7c1e017aa6bae1e0fb435fd1a319481f8e5868316232b8d00

Observation 64243a39-bee3-4aca-b170-7f18f7ea9eea · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.474745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.474745Z digest=sha256:72f971704efd085f640bccdcc547d417f22a220f53cf0b69e6fcfd9aaaf1461f

Observation 7b4824b5-1247-46f0-9acb-6b6f537991e6 · outbound

This paper cites The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos The All-Seeing Project: Towards Panoptic Visual Recognition and Understanding of the Open World

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.496607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.496607Z digest=sha256:94da0f1610996675e66f54164d256260ec057e3ea6aed139e2a5055b8111a4c4

Observation ce98c32a-3bee-4a17-b336-b03bff860f17 · outbound

This paper cites Grit: A generative region-to-text transformer for object understanding.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Grit: A generative region-to-text transformer for object understanding

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:17.065248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.564740Z digest=sha256:acc02da06e642d6270c93cb17bb80cb5a2d27c78af0b6396ea4173050beda517

Observation feb50aa7-36ce-4526-9e51-e569d7c557cb · outbound

This paper cites Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.Advances in Neural Information Processing Systems, 37:69925–69975, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Visionllm v2: An end-to-end generalist multimodal large language model for hundreds of vision-language tasks.Advances in Neural Information Processing Systems, 37:69925–69975, 2024

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.586867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.586867Z digest=sha256:e7e3ad936cea5d9ce6b4abeea9e1ff32d0c7a5ffbd99c0d80c8a04a3b57f6fd2

Observation 4e4340d2-d0db-4de9-ba17-cf974d152d9f · outbound

This paper cites Youtube-vos: A large-scale video object segmentation benchmark, 2018.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Youtube-vos: A large-scale video object segmentation benchmark, 2018

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:16.863847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.605063Z digest=sha256:6e6370dda325addbd177793c537843a4fad07c5d9dfac11d496924950b209ab5

Observation 6306acd4-37f1-40c5-8523-43cfcc258b31 · outbound

This paper cites Qwen2.5 Technical Report.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Qwen2.5 Technical Report

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.643014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.643014Z digest=sha256:329c68a42694df7d9f0f2d611b8df601c2a3f0aa718f2261d3823d664a6144ed

Observation bba35fb9-1d5f-43a0-9e47-1b2befb8714e · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Vid2seq: Large-scale pretraining of a visual language model for dense video captioning, 2023

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:16.586598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.676752Z digest=sha256:a360fe7d87658fe65178112e4e2f575f8559ede2ca8efd499ba4385afcdbe97d

Observation 038aa238-4d5e-452e-897b-8f9fd415f96a · outbound

This paper cites Samurai: Adapting segment anything model for zero-shot visual tracking with motion-aware memory, 2024.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Samurai: Adapting segment anything model for zero-shot visual tracking with motion-aware memory, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:16.410525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.714771Z digest=sha256:8ed62317253cc98108e3c6858fde9332ec0aaf5c9791249528362a5e0f806b2c

Observation 6f24cf1c-e96e-498b-94ec-f38cf6b5d0a6 · outbound

This paper cites Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Set-of-mark prompting unleashes extraordinary visual grounding in gpt-4v, 2023

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.738804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.738804Z digest=sha256:bc34812f5cd97ec1a37996d1ecf6dc7824e1bce95409667ea64d3313dae7a913

Observation 7b581093-f66b-4a18-8762-233b527d2abf · outbound

This paper cites Detecting texts of arbitrary orientations in natural images.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Detecting texts of arbitrary orientations in natural images

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:16.269910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.788564Z digest=sha256:a1dbb29150c7a1d03ecf3c364c929c419d49232c2a45a3f183f718db3b490c0e

Observation 15915057-4c0d-42e2-9ead-e8bb545702d5 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.842739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.842739Z digest=sha256:ae139cc46ec66b1f30e5c1b6ff9ed5f618977dbb24631d78b311f9a18f78c5ab

Observation 54f1de86-fd8f-4e15-ac4e-e612520875a4 · outbound

This paper cites Merlin: Empowering multimodal llms with foresight minds.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Merlin: Empowering multimodal llms with foresight minds

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:13.904324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:13.904324Z digest=sha256:c660688224afa55922b9523433fc6329ab70bea276ac8d82207516f51b23f924

Observation 04667787-33fe-4d7a-b681-83be6aea276a · outbound

This paper cites Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos, 2025.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Sa2va: Marrying sam2 with llava for dense grounded understanding of images and videos, 2025

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:16.098427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:13.964588Z digest=sha256:bd6458a2e7f2db26d0658e32159f47546f4a8db5dc71811b176a84e85b979110

Observation 3532d1b8-250b-4955-bf83-c62da59b9e9e · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Osprey: Pixel understanding with visual instruction tuning

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.049510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.049510Z digest=sha256:ef2fb3f16be0389cf8097c4c21ab68403586f54cb95efbadb59ed4e26db48a41

Observation 60a624db-a6fa-43c6-aedc-ac269ea708c2 · outbound

This paper cites VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.117492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.117492Z digest=sha256:ef17dd36c624aba003330f824ef4c1d79b3749ed81f1a89597c9d17268509ef3

Observation 86547f57-af2f-49a4-aa87-4cd8e57a6fa5 · outbound

This paper cites Faster Segment Anything: Towards Lightweight SAM for Mobile Applications.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Faster Segment Anything: Towards Lightweight SAM for Mobile Applications

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.194914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.194914Z digest=sha256:d10ae9147beea62aba084784357624cdf29805ce28f860653c8434ae79058ab1

Observation ae642888-c3de-4cbe-aa24-3161505b0301 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.279027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.279027Z digest=sha256:1d5a260bfd57c773f76ff28ac6dcecd4827cf844df1ba97a1e8fdc9de11e5afa

Observation e377b8a5-dc62-43a2-b95e-5499da6488b8 · outbound

This paper cites Gpt4roi: Instruction tuning large language model on region-of-interest, 2025.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Gpt4roi: Instruction tuning large language model on region-of-interest, 2025

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:15.913568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:14.380692Z digest=sha256:38027f33fee68a49173481ece0dcd9d1649957d3d1fdf6e462c9fe7fd1cf49b0

Observation 614da7e3-d07b-487c-875c-d939bf00e4fa · outbound

This paper cites Where does it exist: Spatio-temporal video grounding for multi-form sentences.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Where does it exist: Spatio-temporal video grounding for multi-form sentences

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:15.744187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:14.500909Z digest=sha256:f22c514b9648487a4b4acdc6045e86aaf6dcd19dd82608545104aa94882c8525

Observation ccdc906f-94a4-4d76-bf34-91859924d5ba · outbound

This paper cites ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos ChatSpot: Bootstrapping Multimodal LLMs via Precise Referring Instruction Tuning

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.565949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.565949Z digest=sha256:8221643c83f1547ea738cfca29c5e04560f393a7d7baa525f7c52023facf3c1e

Observation 037adc13-83fb-4e76-a27e-08af8718a73b · outbound

This paper cites Fast Segment Anything.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Fast Segment Anything

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.704495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.704495Z digest=sha256:4a47c8e9f7ba395af7f53b2900ce95e6d891899064c202f5225730ce62e45b11

Observation 1b5a5567-5bb8-4335-a8ee-4977daff2dd2 · outbound

This paper cites Controlcap: Controllable region-level captioning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Controlcap: Controllable region-level captioning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:15.590718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:14.776221Z digest=sha256:814e7496a836cb40e8a29c4c15ea444c67e61bea908daf0295a49194cea024c9

Observation 103c9006-1579-4b31-b1f1-df970f5937e1 · outbound

This paper cites Streaming dense video captioning.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Streaming dense video captioning

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:28:15.393823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T10:28:14.907661Z digest=sha256:4d2b08496b16c1679ce75d1e52b5162188a07228836f529965e7e7e44a1bb979

Pith citing papers

Observation 1309521c-a60a-444f-ba0b-13861ed211f6 · inbound

Describe Anything Model for Visual Question Answering on Text-rich Images cites this paper.

Describe Anything Model for Visual Question Answering on Text-rich Images Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:24.717294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:24.717294Z digest=sha256:cd0ee7c4ea0b777f58a65c6c3fe258320175a0330eafb64e61444856a72750f9

Observation 54b3d678-10f0-4408-b6bb-06bdb6520ea5 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 293

Resolution
unresolved
no resolver link, observed 2026-08-05T20:29:11.675850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:29:11.675850Z digest=sha256:86dcc21965cc674d408110c24c27044da6e6119ac9b77b590d1cdce943d42c22

Observation 2b0f3896-8ba2-43c2-ac5a-aa020adf157d · inbound

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data cites this paper.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:16.722899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:16.722899Z digest=sha256:41f04ec4d5dc0c450465e2eee030db28c3f8c396c6acd159ac58dbd881888f65

Observation 5bdff3cf-8f46-4932-b4a7-efc5fd55bbdc · inbound

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension cites this paper.

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T18:15:05.829843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:15:05.829843Z digest=sha256:4d1b4d548eb1a21d78257d22ef0eae61fee9e8db4ded6c9af170e490fa0eebb5

Observation 64af0091-4638-4aaa-8598-57b8cbcfa627 · inbound

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition cites this paper.

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:21:23.379690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T00:20:58.483350Z digest=sha256:a2adfcc76a2a12a59fb66976aeb349fe136c3459205192af21a56b37bec3848c

Observation f78f9f04-29ac-4021-89e9-4f639e9d59ea · inbound

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling cites this paper.

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:41:19.224758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T22:39:32.955779Z digest=sha256:5833f81ac65f8d86284c117f69c2e93563a30bb8c8a8fde83e14a493e3932cab

Observation 59ec8268-c8fa-4276-9aee-c8c8bb744edd · inbound

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling cites this paper.

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:37:05.085568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:37:05.085568Z digest=sha256:a832b35e2bbaeb0c93ea6b27cb5992a793177d46ccf8b7300741a2292b835977

Observation 4d313a4e-0c65-4f54-b147-e665b3c3d17c · inbound

Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration cites this paper.

Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:37:37.229588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T08:33:49.841678Z digest=sha256:f8b162ad6d4fbeb5b4d444b4e78c9c5acf9896de7be728a0588e8793130ce0ad

Observation 728775e5-abdc-40d7-84b4-a2b3a515999e · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:45:49.048227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:0ed584726ff4d1de7db0e909134e947a74d82ad8db44f4401b352e7987a075fe

Observation 87f48f3a-8a30-4c2e-baec-3acd0e31e9bc · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:26f963181b0697329edd89dabbde80138ca0868d1950effb83213284619e0fb0

Observation c25c039a-fdd0-4fdc-a0c6-8609bd93d72d · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:06.525843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:9bf3ff44a1e35715d261a89415d7fa743c75bf60998bdf1fb8fef422ed5184f3

Observation fecadecc-bf3c-4c06-850e-e596473f3e66 · inbound

WOW-Seg: A Word-free Open World Segmentation Model cites this paper.

WOW-Seg: A Word-free Open World Segmentation Model Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:27:48.052668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T21:23:14.311122Z digest=sha256:b54d262427b9be43ca24e2fdccb818bce301e7c23b028c828fa40035eb3cf35e

Observation d1ce6cc2-ff2d-43fa-bcdf-2036aa9755d6 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.316928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:b9bbb37395db6ba1f46068e64d9470e730d53f7cc727401d4a0d6a876ee03dd7

Observation 694d2d02-f883-4042-9944-73eace34e136 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.783743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:d506e6749dbd2b4a90cf8ab9766854d465c937ab9f22f7fdecd8921e97423ae7

Observation f067d0b1-ed72-4548-b86f-4a796ce217ed · inbound

FRFDet: Efficient UAV Small Object Detection with Symmetric Sampling and Scalable Fusion cites this paper.

FRFDet: Efficient UAV Small Object Detection with Symmetric Sampling and Scalable Fusion Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-11T21:31:04.773453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T21:31:04.773453Z digest=sha256:22a1c33acc4cb4bc02a38c84497121a04df08e4f7621b6bc97c8ec661474cc2b