Pith. sign in

Paper Citation Record · LEDGER

Vision Generalist Model: A Survey

As of 7 August 2026, this Paper Citation Record lists 100 of 222 outbound references and 0 inbound Pith citation observations for arXiv:2506.09954.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09954 v1

Coverage vector

measured 100 of 222 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:43:49.218178Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 222 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 170492ca-7001-4dd5-b2c7-0adcd413ce29 · outbound

This paper cites GPT-4 Technical Report.

Vision Generalist Model: A Survey GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.161463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.161463Z digest=sha256:7ddd9838bf4fe582bdeca04277caff53e9ab129f9c92206a471aec5a146f37ac

Observation 11b6f81b-f26d-4862-8abb-66cd72e73c3e · outbound

This paper cites Deep Learning using Rectified Linear Units (ReLU).

Vision Generalist Model: A Survey Deep Learning using Rectified Linear Units (ReLU)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.301013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.301013Z digest=sha256:4f94c4a75f7e7526915960263a3c6f36288d7afd6d102cc7678cc88522055055

Observation fd7ab959-d6a4-4d1d-8a68-11faf61f2e16 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Vision Generalist Model: A Survey Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.492863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.492863Z digest=sha256:33c05c3c52b6abdbc37acd760a2f8ceb20aa38f3d7b8373b4839163b5d2950b7

Observation dda91768-a04a-4500-8903-48c57bc8f19f · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Vision Generalist Model: A Survey Bottom-up and top-down attention for image captioning and visual question answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.624205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.624205Z digest=sha256:091e76e878cf7a8c6dcae3ddc06cbd75f9e0415f5534c3f4404e91ee779f6b99

Observation b3433fa7-8647-45a5-a85b-412f0a3ab4e9 · outbound

This paper cites Neural module networks.

Vision Generalist Model: A Survey Neural module networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.808840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.808840Z digest=sha256:a5e6b8c32a28a6220e0601f2f223a0fd3f6f07b2bc99c1cfce0cde66526d617f

Observation d088b8bf-1a37-4879-abb3-8a7b17cf882a · outbound

This paper cites Vqa: Visual question answering.

Vision Generalist Model: A Survey Vqa: Visual question answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.944266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.944266Z digest=sha256:77a64ebb9075320b495431a90fcc46f0227c04882c647dc3e258a5445a2d4e47

Observation 8be7d67e-4201-4abc-acb6-e92aa4708ba6 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

Vision Generalist Model: A Survey MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.106252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.106252Z digest=sha256:23ce5a5477db5a65c19a00fdcd92834d8b98c62a9d45ee646d6c73c73770b967

Observation 92a4b703-3a1e-4f75-a2fa-9085b94f0e45 · outbound

This paper cites Foundational Models Defining a New Era in Vision: A Survey and Outlook.

Vision Generalist Model: A Survey Foundational Models Defining a New Era in Vision: A Survey and Outlook

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.268444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.268444Z digest=sha256:855e7f7946bb50eadd45e33aefe9bb14ae12d46f3ea4eb372e87ee49508ac824

Observation 4d794112-46c5-4cc3-99bb-02288db94b87 · outbound

This paper cites Layer Normalization.

Vision Generalist Model: A Survey Layer Normalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.429052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.429052Z digest=sha256:4f98d5b979e7dfe46cc5ad6a6104ffbd473405bf005e341adb584a4e66d71215

Observation da406ff8-ded6-4204-a7c5-c206e9a54b5d · outbound

This paper cites Multimae: Multi-modal multi-task masked autoencoders.

Vision Generalist Model: A Survey Multimae: Multi-modal multi-task masked autoencoders

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.515279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.515279Z digest=sha256:a941c427524dea7d5ffb6866af4af21ee8c6aa80e8c525008eec74dd760fd7ac

Observation eddf264e-d766-4325-b410-eab7a6996f9f · outbound

This paper cites Graph Perceiver IO: A General Architecture for Graph Structured Data.

Vision Generalist Model: A Survey Graph Perceiver IO: A General Architecture for Graph Structured Data

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:44:08.902937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:43:33.610330Z digest=sha256:0cff5b2936d3054d1c7f7190262dd3dd3c05c4966aa97f51f3e968b69d85dbf6

Observation 55253928-6d62-4585-8989-bc968d08efce · outbound

This paper cites Data2vec: A general framework for self-supervised learning in speech, vision and language.

Vision Generalist Model: A Survey Data2vec: A general framework for self-supervised learning in speech, vision and language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.753726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.753726Z digest=sha256:3f00215f2ad5d226152835440f902747a1ca7aa09ee43b455c05c57cc9c5a59c

Observation d1f3a525-27a8-4630-a34b-89abd0681f0e · outbound

This paper cites OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models.

Vision Generalist Model: A Survey OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.884437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.884437Z digest=sha256:3e5a14c9b12f9e7ef971eadf22fb22c4bf868fcdc5c4dc5c3f2a541cdfcb5e27

Observation bf4d225b-6fa5-4514-bf1a-065548e5d0fe · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Vision Generalist Model: A Survey Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.007655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.007655Z digest=sha256:6e9b22188ff7fdce8f2fe2a53e9154b9c9aaa02b118bf8cb794d70ff447af421

Observation 127adca0-bb3f-4319-8cbc-aea36c14eda3 · outbound

This paper cites Qwen2.5-VL Technical Report.

Vision Generalist Model: A Survey Qwen2.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.123786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.123786Z digest=sha256:485029f31da1a976f9cc60c1b77ede8ef39224c518260d0dc2c8e42edba5c5f8

Observation f0484fda-7151-4cb3-8b01-1c6adab20335 · outbound

This paper cites Sequential modeling enables scalable learning for large vision models.

Vision Generalist Model: A Survey Sequential modeling enables scalable learning for large vision models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.225106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.225106Z digest=sha256:fa9c166f2f4bfe4efec19517fec235f6c8ba9c3e7d82b5b5816937082ca3fbf0

Observation 81e0a385-e574-44b8-84f6-362c34c4d02a · outbound

This paper cites Visual prompting via image inpainting.

Vision Generalist Model: A Survey Visual prompting via image inpainting

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.375878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.375878Z digest=sha256:3dfbfffa06fbd6b9ac398467643081d45cf8793294e311eb704e668de8aeeb00

Observation 3fd5a6a7-ad46-4fb6-b29d-59a30bc2ac10 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

Vision Generalist Model: A Survey Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.481533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.481533Z digest=sha256:81a97b8fec223089c810c4c4b4d162828e41c45ee9a31fe12a38c13efaa8d911

Observation 59a1f3ea-9d07-4c95-bde1-319ffce3e17a · outbound

This paper cites Mult: An end-to-end multitask learning transformer.

Vision Generalist Model: A Survey Mult: An end-to-end multitask learning transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.595969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.595969Z digest=sha256:33b743ef627b412242ee6be48cb147b664f05015f2be626b8edd69c8ecd8b9f4

Observation d3e906be-d1bc-428e-bcac-c705dee63da9 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Vision Generalist Model: A Survey RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.702680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.702680Z digest=sha256:064455482bac42d7603a5153a792921e22148790d3648e3a1783e215d338795c

Observation 3527e882-6dab-4a35-ae35-e46fa12c112e · outbound

This paper cites What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test.

Vision Generalist Model: A Survey What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:35.039485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:35.039485Z digest=sha256:b774cd2dd9efa25432a1c64424cea41c0fc2bfec7b23341638b129a0e99d5f09

Observation f8c778f6-75ea-4d8d-9587-6f59cba1f1bb · outbound

This paper cites HiP: Hierarchical Perceiver.

Vision Generalist Model: A Survey HiP: Hierarchical Perceiver

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:37.724569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:37.724569Z digest=sha256:1f3fa3cb8713765e39caaecc14ed818e9141f95233d2778a7770afb5cbaf7bcd

Observation d3b70e7e-b3d8-45c1-8cf4-919570648fb2 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Vision Generalist Model: A Survey Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:37.902765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:37.902765Z digest=sha256:08bfa6b50cea304f24d9284f2034595e20e567159395e434695a11cc64948d4b

Observation ef691314-189e-4fde-9261-15bab6d652eb · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Vision Generalist Model: A Survey MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:38.064012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:38.064012Z digest=sha256:8efc20e306b9cb22d7150c72c8b2bca2b2ef131d68730c5570f19bc80aa894c4

Observation 6aca9191-08e3-412e-818a-2e2fa2143fde · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Vision Generalist Model: A Survey Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:38.403191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:38.403191Z digest=sha256:1599748f5766da94b17f9e11a4cefbb7dad3637b5a8664cb6a130b0e164be103

Observation 68732d81-8f1b-4a22-abae-203b7895dc93 · outbound

This paper cites Autoformer: Searching transformers for visual recognition.

Vision Generalist Model: A Survey Autoformer: Searching transformers for visual recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:38.558566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:38.558566Z digest=sha256:1ee4bb90dda8daef9485dca60e376a1ec3126b51d0988a4c2bc972a472bf7d34

Observation 20d558c9-03d1-460a-bf15-f7bd0da26a2b · outbound

This paper cites A Generalist Framework for Panoptic Segmentation of Images and Videos.

Vision Generalist Model: A Survey A Generalist Framework for Panoptic Segmentation of Images and Videos

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:44:08.653600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T04:43:38.767467Z digest=sha256:3ea01f7a76065178272dbcc39e772cb718126d149daaad338dd58de216a26623

Observation b7cbd325-4c3d-404c-9a13-e4a87d1106bd · outbound

This paper cites A unified sequence interface for vision tasks.

Vision Generalist Model: A Survey A unified sequence interface for vision tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:38.906223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:38.906223Z digest=sha256:9692f702ed6b9065aa875811c4cfe4802bda6bd2e12a9a4ae131ff95b7af1020

Observation 5b8afdf7-8fad-4379-8551-339e586231ee · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Vision Generalist Model: A Survey PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.103950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.103950Z digest=sha256:f1d0adeda9da9e5eb162eee60485839de16339bdbaf116b95615ff6ea00ae6fc

Observation 6d47b4a3-6534-4612-9afc-9bfa39b894ce · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Vision Generalist Model: A Survey PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.348446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.348446Z digest=sha256:4bceed255471600d6e5b72ad46db9b751c731700836268ec415d44adef7ed09e

Observation 56a37139-f307-4f1d-b085-0b1804378a07 · outbound

This paper cites Learning a low-level vision generalist via visual task prompt.

Vision Generalist Model: A Survey Learning a low-level vision generalist via visual task prompt

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.481354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.481354Z digest=sha256:1951e7afe3623e2b71b641690cadc1f9cc24ee01fab042ddca54794c144fbc66

Observation 86253100-b9d1-466e-872a-67e9ae83c7dc · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Vision Generalist Model: A Survey Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.618344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.618344Z digest=sha256:f10ae9cc03d7511e4c9f8a897536330c9884111f8b18d0e9480d38f1f1bf95eb

Observation 3449b306-e90d-4667-8150-d155d9118250 · outbound

This paper cites Uniter: Universal image-text representation learning.

Vision Generalist Model: A Survey Uniter: Universal image-text representation learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.786113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.786113Z digest=sha256:4ea7491d0f5a0132cdaf5f1d104a5708ae13d96d594ebaaba742a98e9fa63567

Observation 808eb52b-9a43-4682-89a6-bbda36e0dd47 · outbound

This paper cites Uniter: Universal image-text representation learning.

Vision Generalist Model: A Survey Uniter: Universal image-text representation learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.930759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.930759Z digest=sha256:5206cc11af73fa04ac542ee6d28cdaa973654fbd2af5939e1320edada3400c24

Observation ef054bdd-1664-49f3-9945-9a196a359437 · outbound

This paper cites Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks.

Vision Generalist Model: A Survey Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.112973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.112973Z digest=sha256:fbf3c18dd4fcc54802616ac81dcab3d177df8480cde6786e8d25ffdf85a991ff

Observation 1254a781-eb71-4467-895a-e0596341041a · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.

Vision Generalist Model: A Survey Per-pixel classification is not all you need for semantic segmentation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.257715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.257715Z digest=sha256:7e057f772c4f1deb3877515ddcfc16038eae3ef90c9ad86f027ff27c3ed9c7c1

Observation 6f797c5f-c547-49b1-8f41-2270661b9308 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Vision Generalist Model: A Survey Masked-attention mask transformer for universal image segmentation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.435426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.435426Z digest=sha256:b28622c25d10fca6a078a436184a80e6371a158a5225f9a1da17027fa0b40898

Observation 6818b099-ca34-4e96-8c1c-65b719307360 · outbound

This paper cites Unifying vision-and-language tasks via text generation.

Vision Generalist Model: A Survey Unifying vision-and-language tasks via text generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.528315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.528315Z digest=sha256:3f0697d55a2c24cc5a0a64a1a6232877e7dbf1c6bf47b17f4d6083aeeb507ef1

Observation dcc5ce5f-76eb-4247-8dea-4353601196d6 · outbound

This paper cites Conditional Positional Encodings for Vision Transformers.

Vision Generalist Model: A Survey Conditional Positional Encodings for Vision Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.606136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.606136Z digest=sha256:0a5e236fb4cfd2d0c7d7e4579789554a362456106a713bfbb021ce03de521f4d

Observation 394a616e-3668-4058-adba-c3c6372cd344 · outbound

This paper cites Instance-aware semantic segmentation via multi-task network cascades.

Vision Generalist Model: A Survey Instance-aware semantic segmentation via multi-task network cascades

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.694473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.694473Z digest=sha256:444d42e40b5d4e3feccaeea94ec03d833cf0366532d4fa0c22f296a3f03e4ae1

Observation 5d7f965a-1865-44fc-978d-b74a856fa843 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Vision Generalist Model: A Survey BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.810289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.810289Z digest=sha256:bb1c88767e52acc5f928482736c1d2b35578770bfd41fa886564e56add582fb3

Observation 7c327d4b-604d-40e0-97fe-0b3d6755a824 · outbound

This paper cites Decoupling zero-shot semantic segmentation.

Vision Generalist Model: A Survey Decoupling zero-shot semantic segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.887551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.887551Z digest=sha256:b88f1d4f11d6f9de501dee6dbcd56b5121a2ef2144c448f84245a7c4012a6549

Observation f74f887e-624d-4bc1-b7b7-c877a16aa6c6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Vision Generalist Model: A Survey An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.947027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.947027Z digest=sha256:8b8b51f0e53d725f48b58d782cbb418dc216494f35a882a884d368d4328713cb

Observation 76496712-ffb0-4f33-a0c8-dcf07cc78315 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Vision Generalist Model: A Survey PaLM-E: An Embodied Multimodal Language Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.092048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.092048Z digest=sha256:b5970f436c1f6c4b9b049419f5d9d2022a69f5d38b62723fb4c6ec780484c9d4

Observation 27e29feb-4249-43a0-80c4-70078c272ae5 · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts.

Vision Generalist Model: A Survey Glam: Efficient scaling of language models with mixture-of-experts

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.218617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.218617Z digest=sha256:6715b114551722fb018c29fb849d397b75b170943dafef7718999cd0196bc133

Observation fd737cc9-e891-4ce8-8965-e2194fa52c25 · outbound

This paper cites Multi-modal alignment using representation codebook.

Vision Generalist Model: A Survey Multi-modal alignment using representation codebook

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.312439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.312439Z digest=sha256:1146f3314a900a17294e0853c1628b3f97b47eab4f41df538a5c0568c0272fa8

Observation 8a70b3b4-f3b5-4b38-aa5a-dbcb44eb9a9b · outbound

This paper cites Low resource dependency parsing: Cross-lingual parameter sharing in a neural network parser.

Vision Generalist Model: A Survey Low resource dependency parsing: Cross-lingual parameter sharing in a neural network parser

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.401628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.401628Z digest=sha256:c8b872bc357900ac6c8baead986b1ced79811fa1a34168330f05cdf45b37dee5

Observation 7f022ef0-546d-47ef-87af-fef343b19160 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Vision Generalist Model: A Survey Taming transformers for high-resolution image synthesis

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.537311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.537311Z digest=sha256:4472cd1f3fd4334305d13b18c7e68b62bb9cf99b3b395aebff6580fb18e564bd

Observation 3bf1de95-4651-418a-9a3e-43aa77d84d8c · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

Vision Generalist Model: A Survey Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.642922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.642922Z digest=sha256:7f805971569c2e6d8a45b80dc69077751bd8cc2ebf553dacb7b616839d888af7

Observation a1c4fa58-0474-4b07-93af-27137f993a10 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Vision Generalist Model: A Survey Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.769713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.769713Z digest=sha256:2d456649ae3074e3a80d3daf72c2e670bb732b026b94c1559264f9b3344b570d

Observation 31a5d409-1186-47a8-b152-dc606bfc46eb · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Vision Generalist Model: A Survey MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.874281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.874281Z digest=sha256:2cd34d656001bbe9647b2a59128d1f58508c4f1fbb51da1ac84e118ee834ee04

Observation 6bce6f40-7fc0-436d-b6ef-710168cd00f2 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Vision Generalist Model: A Survey Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.983897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.983897Z digest=sha256:83ef5fd5b34672b01131db2c77fd5918fc7af9b4c371d2c704308c834f30d8ad

Observation 1198782c-3978-4d7f-bd57-b5177a52c381 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

Vision Generalist Model: A Survey MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.097080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.097080Z digest=sha256:92ad6ecbbdfc8c3912fa2e109ca8768765997a94e5adc05c5d9ef37cf32ca45b

Observation 9aa2c5a8-29d4-4069-bd67-2e16e1ac11fb · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

Vision Generalist Model: A Survey Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.214314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.214314Z digest=sha256:d17c6295a426da758cede580de9ebbcf92e0d0cab13e6ca126a3ecab865dad62

Observation 9539f517-a2bb-476d-8466-6d9f13caac59 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Vision Generalist Model: A Survey LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.303030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.303030Z digest=sha256:0be1daab35894c54fabd1b866d335dac1d64abea92f71a94560530ae5f832b4f

Observation 4c99319a-658e-431f-a42e-278aeaaaed50 · outbound

This paper cites Mtl-nas: Task-agnostic neural architecture search towards general-purpose multi-task learning.

Vision Generalist Model: A Survey Mtl-nas: Task-agnostic neural architecture search towards general-purpose multi-task learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.405980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.405980Z digest=sha256:ae51ce93a70f5f514c8ca63e05195c88cfe1717f192a54bbd84fc057ee65d285

Observation afa323e0-db9f-45fe-819a-ef9ed949a51a · outbound

This paper cites Instructdiffusion: A generalist modeling interface for vision tasks.

Vision Generalist Model: A Survey Instructdiffusion: A generalist modeling interface for vision tasks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.538463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.538463Z digest=sha256:03dd16f756c73e072b425e6ea61ec1f31f47711e5df82a4f7de374fa51f547da

Observation a4ba17bf-996a-45b2-97b6-2119824461a4 · outbound

This paper cites Scaling open-vocabulary image segmentation with image-level labels.

Vision Generalist Model: A Survey Scaling open-vocabulary image segmentation with image-level labels

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.648552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.648552Z digest=sha256:ff5ddaa5f5ce7f1e6493e570daef8f7b82c4a19ff636d4c723815275a41cea1a

Observation 59e09b5c-2843-4fe1-8a00-3791ea115380 · outbound

This paper cites Omnivore: A single model for many visual modalities.

Vision Generalist Model: A Survey Omnivore: A single model for many visual modalities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.724865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.724865Z digest=sha256:6895f5350d211b6034f7435d0256c1432d761beee99adf17982aa009dfb980b1

Observation 64ad3603-6dd5-4182-8e21-9d7114ca8a36 · outbound

This paper cites Mixture of experts models.

Vision Generalist Model: A Survey Mixture of experts models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.783410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.783410Z digest=sha256:22d153d77b5898a7e64dfb85fc2549195726d5c253bd3d8f6ea8f4c8fd43ba9e

Observation 108fbebd-7adc-449c-843f-58344ee2372e · outbound

This paper cites 3d semantic segmentation with submanifold sparse convolutional networks.

Vision Generalist Model: A Survey 3d semantic segmentation with submanifold sparse convolutional networks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.847069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.847069Z digest=sha256:658ed60d0bfb0483b541768037639cd494526ebf6fbb40f230769e0b4c5454c1

Observation 6c8fcf7c-9e11-41e3-988f-e7aae0c1e61c · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Vision Generalist Model: A Survey Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.929374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.929374Z digest=sha256:e3542a0ec30e85b5cba2b965a8174f2508946f9172aca37a46d4053c602dc141

Observation 524f39e1-6351-4720-8dde-b722bf31fbc3 · outbound

This paper cites Deepfake video detection using recurrent neural networks.

Vision Generalist Model: A Survey Deepfake video detection using recurrent neural networks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.022561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.022561Z digest=sha256:0ab9a9bdfacf515fe8f55db427977aa0acd36e5794442d550b9140207dd006aa

Observation dfc7ad55-1a4b-46bd-85ba-0aff07b70425 · outbound

This paper cites Dynamic task prioritization for multitask learning.

Vision Generalist Model: A Survey Dynamic task prioritization for multitask learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.122801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.122801Z digest=sha256:b6a56a5e5beb01ffae86467605816c4466b434d2620b2353ef0fea65fe810626

Observation 94e049cb-e175-4c4b-84bd-d7ed481aa2a6 · outbound

This paper cites GRIT: General Robust Image Task Benchmark.

Vision Generalist Model: A Survey GRIT: General Robust Image Task Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.270853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.270853Z digest=sha256:38593bbce16840afcada5de9715b275dc942b0715a8d37dcf1af786b324dcdb6

Observation 5909fec8-b3bb-47b2-80de-f956ce27527a · outbound

This paper cites Onellm: One framework to align all modalities with language.

Vision Generalist Model: A Survey Onellm: One framework to align all modalities with language

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.392291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.392291Z digest=sha256:47b434c284d7352435c9230ef716731e326c6bf1a503b8a7189a70beb2392b45

Observation e867694b-33ca-4192-8dcf-810361659346 · outbound

This paper cites Deep residual learning for image recognition.

Vision Generalist Model: A Survey Deep residual learning for image recognition

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.490926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.490926Z digest=sha256:878b7607c13e4ea4a1dc1df8a6c99d185caea08f8611ce06f64aea3ca463529c

Observation bbeb827e-2db2-4ef1-863e-5886a6606636 · outbound

This paper cites Deep residual learning for image recognition.

Vision Generalist Model: A Survey Deep residual learning for image recognition

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.588569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.588569Z digest=sha256:3c6397044dd49d8294b4848f3c51b24a6abaf3bcf149c92637286c73c4ae5b90

Observation 7f6f8551-b070-484f-b446-4d752e3b9172 · outbound

This paper cites Mask r-cnn.

Vision Generalist Model: A Survey Mask r-cnn

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.728282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.728282Z digest=sha256:ac4e774c95a315a3809e1d559f8012bfa8043d2489ae7f6218e899a11fce2533

Observation 6184590f-6af9-4971-a072-c13aab04b41f · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Vision Generalist Model: A Survey Distilling the Knowledge in a Neural Network

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.821811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.821811Z digest=sha256:06a2332ca446a4f050d53adee4401ca551a0176b31a891f3a58244ef5b8cc717

Observation 832914b1-78db-4b26-9364-e372c3f3f8a0 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Vision Generalist Model: A Survey Classifier-Free Diffusion Guidance

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.966930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.966930Z digest=sha256:16f27a3ff213bd472fab5eae87aa3a1cc2a4e5ba9bb15d98eb82485dc6b82096

Observation 1e3ed28a-94cc-4a54-be49-322a99ba2435 · outbound

This paper cites Denoising diffusion probabilistic models.

Vision Generalist Model: A Survey Denoising diffusion probabilistic models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:44.158944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:44.158944Z digest=sha256:59cdd3a959a7e1be6bfa756a96eaaa7d172dda03c280cf7b72214e58b25b9e32

Observation 3c883db3-43b4-4db6-9b60-9ea54adb2b37 · outbound

This paper cites Long short-term memory.

Vision Generalist Model: A Survey Long short-term memory

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:44.383971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:44.383971Z digest=sha256:1488fc82bebe1258b67ab3c2aaca99d17c153fc9b1108b92ef59c45894fc982b

Observation 14bc7247-bcb2-4667-809a-57246b0a91cf · outbound

This paper cites Unit: Multimodal multitask learning with a unified transformer.

Vision Generalist Model: A Survey Unit: Multimodal multitask learning with a unified transformer

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:44.546997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:44.546997Z digest=sha256:5d4d56ebf64bf73d4da243812a0ddee83fb166f1536309d761faec4019d0f065

Observation f25d31fa-61a2-4121-a7cd-e61f9fc7b8b5 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Vision Generalist Model: A Survey An Embodied Generalist Agent in 3D World

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:44.771145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:44.771145Z digest=sha256:6794ab28a5fd978b14ee7647c9f581b376ab4d5361901befaf0602bc625e642a

Observation 109d6306-2545-47c5-899b-c9395d515211 · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

Vision Generalist Model: A Survey Language Is Not All You Need: Aligning Perception with Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:44.973382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:44.973382Z digest=sha256:5ff35c68e0b092f50a39292993da09e9336d371cc1e32b9a1dfe5b516a63bf16

Observation d1552308-4ba0-4b04-b8f0-3a44805f168c · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

Vision Generalist Model: A Survey Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:45.149991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:45.149991Z digest=sha256:7fd506836f5f807f7f45839ed74f181098f08f77cc31738e63dc12a6da15d576

Observation ff85f8d2-94f1-4363-9875-2957dcc4ffa8 · outbound

This paper cites Perceiver: General perception with iterative attention.

Vision Generalist Model: A Survey Perceiver: General perception with iterative attention

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:45.347606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:45.347606Z digest=sha256:c50ef39e08ffa1ca82eb18798f0ca26bc4294503feab468a9844a7c1947e9afd

Observation 06cfafc9-e956-4a90-a212-cbacda38b73a · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

Vision Generalist Model: A Survey Oneformer: One transformer to rule universal image segmentation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:45.503099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:45.503099Z digest=sha256:f428625f5426f7fc0fd64845742d83cb6338d8b16973a28c20572c9d22da9dcd

Observation 283d2273-548b-42f1-8710-1b6284021692 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Vision Generalist Model: A Survey VIMA: General Robot Manipulation with Multimodal Prompts

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:45.709626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:45.709626Z digest=sha256:596cae0d2331a4058ccbd20c0bfffc688635f04be33e350be5caa53e611a9220

Observation f4c42f47-a091-49cf-9006-eedc72a1839c · outbound

This paper cites Unified language-vision pretraining with dynamic discrete visual tokenization.

Vision Generalist Model: A Survey Unified language-vision pretraining with dynamic discrete visual tokenization

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:45.853890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:45.853890Z digest=sha256:35f8ef3f1bea35ff57706be3d57d6c6bb62998015683329378658c892ac6bbb4

Observation 19153742-5e8b-45c7-8dcf-9fa41894dff8 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions.

Vision Generalist Model: A Survey Deep visual-semantic alignments for generating image descriptions

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.072821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.072821Z digest=sha256:dda1147857afc364e8899fb9f04bab5a4e15a23303a2ae3a21ea8976bc6f5c02

Observation f62d91b2-1762-4e9d-94ad-f0c8621573a0 · outbound

This paper cites Multi-task learning using uncertainty to weigh losses for scene geometry and semantics.

Vision Generalist Model: A Survey Multi-task learning using uncertainty to weigh losses for scene geometry and semantics

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.234329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.234329Z digest=sha256:3a13e5579aad1fbcf1763d4c619154eebed111ac75f308ef1f5523e300a094ba

Observation e66a144c-082c-4c5c-abfb-8a2ef80199ae · outbound

This paper cites Segment anything.

Vision Generalist Model: A Survey Segment anything

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.394972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.394972Z digest=sha256:d80eb4f0acd763fd8537c029a1d528a94838eba90f89e0e2fb1fb8f6346824e3

Observation 56e22b71-6341-4b29-b253-d599cf106d49 · outbound

This paper cites Segment Anything.

Vision Generalist Model: A Survey Segment Anything

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.544041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.544041Z digest=sha256:76481a4730782f4791418fccbca68ae8887a8724fb65e58933552c51e8095526

Observation 3e65f32f-fda9-4c12-9da9-ba44a80e60e4 · outbound

This paper cites Uvim: A unified modeling approach for vision with learned guiding codes.

Vision Generalist Model: A Survey Uvim: A unified modeling approach for vision with learned guiding codes

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.806244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.806244Z digest=sha256:9825d2691cc025de70e677c8daf3a92ecc7175f63afd8d1db4993042fb460f8a

Observation 05aee257-12ce-42cb-a0d3-d062bb16a7c9 · outbound

This paper cites Tabddpm: Modelling tabular data with diffusion models.

Vision Generalist Model: A Survey Tabddpm: Modelling tabular data with diffusion models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.982104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.982104Z digest=sha256:481e875424020b8a89d82e0e2157d4d55c2ca8a4a62dba8c11eedb3019fe47ca

Observation 722487cb-9dd2-4d78-97f4-8292a40b8eb1 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Vision Generalist Model: A Survey Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.129273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.129273Z digest=sha256:67fea421702a8d7f84a25c1dcc1402adf78455d988a5570ab8f9f15df81366f0

Observation fcf16247-114b-4183-ac79-25779cbfa097 · outbound

This paper cites Imagenet classification with deep convolutional neural networks.

Vision Generalist Model: A Survey Imagenet classification with deep convolutional neural networks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.280227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.280227Z digest=sha256:4264dabba7ef51bc3e2c6d7d97a2bbe5431c6c3f8ee957d11bb05914a063113e

Observation 60f0e233-0d61-4ae9-846c-abe12bd7ed1f · outbound

This paper cites Babytalk: Understanding and generating simple image descriptions.

Vision Generalist Model: A Survey Babytalk: Understanding and generating simple image descriptions

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.423323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.423323Z digest=sha256:9b619a008ad4bad50a1425bbc399a62ac06f15bd90525654be76830c41d1f1f7

Observation 12beba6b-9608-4375-aee5-8cd89e2565f0 · outbound

This paper cites Base layers: Simplifying training of large, sparse models.

Vision Generalist Model: A Survey Base layers: Simplifying training of large, sparse models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.581848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.581848Z digest=sha256:f83d30aeb79d39cbaf507b8eb188334098ac8c495275a85fd21d10ecb9f32805

Observation d7a90b71-5da5-46df-af46-7e3b8af095e2 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Vision Generalist Model: A Survey SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.735480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.735480Z digest=sha256:a391a099f5d15fb74e418886da9f81ea417e3974f6215998da9be8d88ec4d536

Observation 594b128c-0f69-44c4-a94c-102c4f6df394 · outbound

This paper cites Multimodal Foundation Models: From Specialists to General-Purpose Assistants.

Vision Generalist Model: A Survey Multimodal Foundation Models: From Specialists to General-Purpose Assistants

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.872697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.872697Z digest=sha256:a8b515d51c9936ca967ba53126925c9d155e2a2a8b25685fe427dadc16cd013a

Observation 12500a93-70ca-4ca8-80c0-8f904f0bac2c · outbound

This paper cites Mask dino: Towards a unified transformer-based framework for object detection and segmentation.

Vision Generalist Model: A Survey Mask dino: Towards a unified transformer-based framework for object detection and segmentation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:48.106927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:48.106927Z digest=sha256:8c008a60dfe18bd496c30cd24276ebe54fc1a762c30d4d57cde8fa387548a269

Observation 1a796a68-0850-4bff-9da1-2123fcf53024 · outbound

This paper cites Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks.

Vision Generalist Model: A Survey Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:48.304969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:48.304969Z digest=sha256:9d501ce4f935ecab1183fa5ba15fa705c447620df733d068a3f3f173f72c1e64

Observation 42360ffd-43e2-4579-a9fe-22f6808caddc · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Vision Generalist Model: A Survey Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:48.536824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:48.536824Z digest=sha256:a7f63594793f66a315b1434811abf59a243ca3a5381b77f5dc925a2fa6df0c4d

Observation 37db749b-7cc5-4cda-a039-5ffb7e93afd4 · outbound

This paper cites Lavender: Unifying video-language understanding as masked language modeling.

Vision Generalist Model: A Survey Lavender: Unifying video-language understanding as masked language modeling

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:48.684658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:48.684658Z digest=sha256:a192aaf7aebbc3fd548963ea66726ac1afa1c8990fd4b7b11b1267022e4c7716

Observation 1327ed3c-d85d-480b-ac66-2fc4e89e18cf · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Vision Generalist Model: A Survey VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:48.858023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:48.858023Z digest=sha256:1e3363affd8b6fd576a3afbd8e100f675ce1e2ec93884e53f558e925d919c5fe

Observation b42aa6ee-2894-45bf-99dd-8907276e4d7f · outbound

This paper cites Grounded language-image pre-training.

Vision Generalist Model: A Survey Grounded language-image pre-training

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:49.054065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:49.054065Z digest=sha256:818cd6707f9ae91a4b8362102bf13a1b9c0b57988b80d4aecab4c7d0f383679f

Observation 779fed7b-b38a-4d86-8a5e-eb7c7f3a422e · outbound

This paper cites Omg-seg: Is one model good enough for all segmentation? In CVPR, 2024.

Vision Generalist Model: A Survey Omg-seg: Is one model good enough for all segmentation? In CVPR, 2024

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:49.218178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:49.218178Z digest=sha256:a8d0b33f32670f7984f5a897935a92b78300c527b237f35318bd49a9150081e1

Pith citing papers

No inbound Pith citation observations are available.