Pith. sign in

Paper Citation Record · LEDGER

Vision Generalist Model: A Survey

As of 21 August 2026, this Paper Citation Record lists 100 of 222 outbound references and 0 inbound Pith citation observations for arXiv:2506.09954.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09954 v1

Coverage vector

measured 100 of 222 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:43:49.218178Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 222 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved98
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 170492ca-7001-4dd5-b2c7-0adcd413ce29 · outbound

This paper cites GPT-4 Technical Report.

Vision Generalist Model: A Survey GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.161463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.161463Z digest=sha256:5cd73edab55bbcd1333188b60d1582dab7fe459ac343b19215f864c76695638b

Observation 11b6f81b-f26d-4862-8abb-66cd72e73c3e · outbound

This paper cites Deep Learning using Rectified Linear Units (ReLU).

Vision Generalist Model: A Survey Deep Learning using Rectified Linear Units (ReLU)

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.301013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.301013Z digest=sha256:4c662ea8e5ce6426eb3dbe0c2e6e01ba54cd19547a4f0f7b164abe780a7fa79d

Observation fd7ab959-d6a4-4d1d-8a68-11faf61f2e16 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Vision Generalist Model: A Survey Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.492863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.492863Z digest=sha256:89818a7f09d6ae02f4b1ce0f9928d411155862f66466c433190019ed5f6dd4d0

Observation dda91768-a04a-4500-8903-48c57bc8f19f · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Vision Generalist Model: A Survey Bottom-up and top-down attention for image captioning and visual question answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.624205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.624205Z digest=sha256:70f7a13a0553ff459e5abac0548facdd6dd94c3dacf2bd1cabe0db510f6af964

Observation b3433fa7-8647-45a5-a85b-412f0a3ab4e9 · outbound

This paper cites Neural module networks.

Vision Generalist Model: A Survey Neural module networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.808840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.808840Z digest=sha256:d5a466ad906e23e7d54df07d074aaea6d79fc1e1b3d0c15abd10d6519ca301c0

Observation d088b8bf-1a37-4879-abb3-8a7b17cf882a · outbound

This paper cites Vqa: Visual question answering.

Vision Generalist Model: A Survey Vqa: Visual question answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:32.944266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:32.944266Z digest=sha256:01553b29bad34c78ba6d890bbd0318bdd877f7aaa3d3172957ef2f6b6c5fbdfb

Observation 8be7d67e-4201-4abc-acb6-e92aa4708ba6 · outbound

This paper cites MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens.

Vision Generalist Model: A Survey MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.106252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.106252Z digest=sha256:3aa3f4ff4ae21e7ab4ccf6356f631ac0ae1f2049d863195c0f303234ab6a25a3

Observation 92a4b703-3a1e-4f75-a2fa-9085b94f0e45 · outbound

This paper cites Foundational Models Defining a New Era in Vision: A Survey and Outlook.

Vision Generalist Model: A Survey Foundational Models Defining a New Era in Vision: A Survey and Outlook

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.268444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.268444Z digest=sha256:dc8dc137892c43f8abc439becc79a3f0e38b4123aeedf0e39317fdf0a2992347

Observation 4d794112-46c5-4cc3-99bb-02288db94b87 · outbound

This paper cites Layer Normalization.

Vision Generalist Model: A Survey Layer Normalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.429052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.429052Z digest=sha256:43b6f69e8cdcdbd39a09742fb20b519c75288275d0cc9de04e706bf27ced35ba

Observation da406ff8-ded6-4204-a7c5-c206e9a54b5d · outbound

This paper cites Multimae: Multi-modal multi-task masked autoencoders.

Vision Generalist Model: A Survey Multimae: Multi-modal multi-task masked autoencoders

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.515279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.515279Z digest=sha256:8cb11ea4b19e824071151c2c445870a6b2bcccdc0db7b1948fa505ae021cf32c

Observation eddf264e-d766-4325-b410-eab7a6996f9f · outbound

This paper cites Graph Perceiver IO: A General Architecture for Graph Structured Data.

Vision Generalist Model: A Survey Graph Perceiver IO: A General Architecture for Graph Structured Data

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:44:08.902937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T04:43:33.610330Z digest=sha256:3ea28e5c771dd489eabf13df2a6c1a0eb8445e6aba48a7a378cd6cda3d71dcc0

Observation 55253928-6d62-4585-8989-bc968d08efce · outbound

This paper cites Data2vec: A general framework for self-supervised learning in speech, vision and language.

Vision Generalist Model: A Survey Data2vec: A general framework for self-supervised learning in speech, vision and language

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.753726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.753726Z digest=sha256:71dd378a07239a3f062feb385ac3934d8b3ec6abaf94a3e076394f7dc85c1db9

Observation d1f3a525-27a8-4630-a34b-89abd0681f0e · outbound

This paper cites OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models.

Vision Generalist Model: A Survey OFASys: A Multi-Modal Multi-Task Learning System for Building Generalist Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:33.884437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:33.884437Z digest=sha256:71e483cbc1f6fdb3479d51be4118f5126d259efb04929073bad1a2f9a72e852c

Observation bf4d225b-6fa5-4514-bf1a-065548e5d0fe · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Vision Generalist Model: A Survey Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.007655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.007655Z digest=sha256:a6aa5f3c42c4e30467889ce8eabe51760277da540312d62f480847f641248ce9

Observation 127adca0-bb3f-4319-8cbc-aea36c14eda3 · outbound

This paper cites Qwen2.5-VL Technical Report.

Vision Generalist Model: A Survey Qwen2.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.123786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.123786Z digest=sha256:01f804c90e76fbd144baafac0dddeced5a9d3eef2665c9865f0f223f5eaa9bac

Observation f0484fda-7151-4cb3-8b01-1c6adab20335 · outbound

This paper cites Sequential modeling enables scalable learning for large vision models.

Vision Generalist Model: A Survey Sequential modeling enables scalable learning for large vision models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.225106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.225106Z digest=sha256:4f377a625dd0a5dc5257e176c1dfc5c453c0bcf9616762957944b941d3321428

Observation 81e0a385-e574-44b8-84f6-362c34c4d02a · outbound

This paper cites Visual prompting via image inpainting.

Vision Generalist Model: A Survey Visual prompting via image inpainting

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.375878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.375878Z digest=sha256:002e625b95429af0128a520d97ce4e4303c3911dd1099a15d966173b9d23a02d

Observation 3fd5a6a7-ad46-4fb6-b29d-59a30bc2ac10 · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

Vision Generalist Model: A Survey Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.481533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.481533Z digest=sha256:f11636b467cb83b1470896676c97780da377a962626d3afb2e7fb64fb8468a58

Observation 59a1f3ea-9d07-4c95-bde1-319ffce3e17a · outbound

This paper cites Mult: An end-to-end multitask learning transformer.

Vision Generalist Model: A Survey Mult: An end-to-end multitask learning transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.595969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.595969Z digest=sha256:9b730568a2f2525487bcf22fc465bd3ad1d2e24f58c0331c15fe358ce7b2ba25

Observation d3e906be-d1bc-428e-bcac-c705dee63da9 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Vision Generalist Model: A Survey RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:34.702680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:34.702680Z digest=sha256:bfa81a1ae52f360bd2d7dd5fc83ca5edc7173a4d2ecd8dd50c8f1950860e02b9

Observation 3527e882-6dab-4a35-ae35-e46fa12c112e · outbound

This paper cites What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test.

Vision Generalist Model: A Survey What one intelligence test measures: a theoretical account of the processing in the raven progressive matrices test

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:35.039485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:35.039485Z digest=sha256:e68a52da5a679eb8c5cba8df98a93952c9593cea56d7a001f28110251f594421

Observation f8c778f6-75ea-4d8d-9587-6f59cba1f1bb · outbound

This paper cites HiP: Hierarchical Perceiver.

Vision Generalist Model: A Survey HiP: Hierarchical Perceiver

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:37.724569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:37.724569Z digest=sha256:63207c1b7fd3e5d1dd28378df8e94367233a65550e74c01e39548e835c5e2846

Observation d3b70e7e-b3d8-45c1-8cf4-919570648fb2 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

Vision Generalist Model: A Survey Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:37.902765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:37.902765Z digest=sha256:cb5fd86901f27e3452c72b66416bf65f7f8d60891a3a26563da9d4754eaaa9d5

Observation ef691314-189e-4fde-9261-15bab6d652eb · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Vision Generalist Model: A Survey MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:38.064012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:38.064012Z digest=sha256:f61e84bd752efae7ce0426b8f7893709680fdc9a24de748e676544c7104e26c7

Observation 6aca9191-08e3-412e-818a-2e2fa2143fde · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Vision Generalist Model: A Survey Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:38.403191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:38.403191Z digest=sha256:f95df6a26ef05095e6d25d0d25dc0de5a92d6b8b9e2df39b91513fd89546a184

Observation 68732d81-8f1b-4a22-abae-203b7895dc93 · outbound

This paper cites Autoformer: Searching transformers for visual recognition.

Vision Generalist Model: A Survey Autoformer: Searching transformers for visual recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:38.558566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:38.558566Z digest=sha256:9a31b4fdf28a6a4f19c9d3fb7d3794f6b4c73fec1b58c37ef31c230493652edb

Observation 20d558c9-03d1-460a-bf15-f7bd0da26a2b · outbound

This paper cites A Generalist Framework for Panoptic Segmentation of Images and Videos.

Vision Generalist Model: A Survey A Generalist Framework for Panoptic Segmentation of Images and Videos

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:44:08.653600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T04:43:38.767467Z digest=sha256:635b112db8908109759fd9a5da87f35217747df13e2bb73d719c7e76117764d9

Observation b7cbd325-4c3d-404c-9a13-e4a87d1106bd · outbound

This paper cites A unified sequence interface for vision tasks.

Vision Generalist Model: A Survey A unified sequence interface for vision tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:38.906223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:38.906223Z digest=sha256:3f13ef4de0be30d4fa2bec438aeae8f73cadf00d484c3e04a9c9cd0c671f7471

Observation 5b8afdf7-8fad-4379-8551-339e586231ee · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Vision Generalist Model: A Survey PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.103950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.103950Z digest=sha256:a4fa0b1e6ec0ba7acceb67d30fae7de103cb517cbe7ac2782ac617e6c9231ff5

Observation 6d47b4a3-6534-4612-9afc-9bfa39b894ce · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Vision Generalist Model: A Survey PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.348446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.348446Z digest=sha256:92e541e2e8c9f871211a52dfebecde76ad8f7584d941158fe84eb9d3a4949e9f

Observation 56a37139-f307-4f1d-b085-0b1804378a07 · outbound

This paper cites Learning a low-level vision generalist via visual task prompt.

Vision Generalist Model: A Survey Learning a low-level vision generalist via visual task prompt

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.481354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.481354Z digest=sha256:e87c25ff88fc3d7a43231c46b40d9a97dbb41b0f7fd111e453a6d4b6a5bc6de1

Observation 86253100-b9d1-466e-872a-67e9ae83c7dc · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Vision Generalist Model: A Survey Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.618344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.618344Z digest=sha256:9e1b2d99f4124294430b1d9230de89ea75f211ba69124c66cf997117c06f8d60

Observation 3449b306-e90d-4667-8150-d155d9118250 · outbound

This paper cites Uniter: Universal image-text representation learning.

Vision Generalist Model: A Survey Uniter: Universal image-text representation learning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.786113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.786113Z digest=sha256:2edf56bc48b21e469a3c03ea9b2f2b868ea7f55459a331f05c707edbb11daf07

Observation 808eb52b-9a43-4682-89a6-bbda36e0dd47 · outbound

This paper cites Uniter: Universal image-text representation learning.

Vision Generalist Model: A Survey Uniter: Universal image-text representation learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:39.930759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:39.930759Z digest=sha256:b1875585a3cf7e93053fd122c4b9fc41f529a2a73963f8afff4c7a12422bc9c4

Observation ef054bdd-1664-49f3-9945-9a196a359437 · outbound

This paper cites Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks.

Vision Generalist Model: A Survey Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.112973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.112973Z digest=sha256:4ba1a6c4929c1f82441982a183237d668d358f87c2085734b037681f068b2dfc

Observation 1254a781-eb71-4467-895a-e0596341041a · outbound

This paper cites Per-pixel classification is not all you need for semantic segmentation.

Vision Generalist Model: A Survey Per-pixel classification is not all you need for semantic segmentation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.257715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.257715Z digest=sha256:66fa2c57bee9b3e7486ba0a9faf0f9d29fe01221530d53a94467f045c601f951

Observation 6f797c5f-c547-49b1-8f41-2270661b9308 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Vision Generalist Model: A Survey Masked-attention mask transformer for universal image segmentation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.435426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.435426Z digest=sha256:635e2a6e83e17d894ca917ba5067de4f48ae1aaf9e3730c4170d94f41a9fadf2

Observation 6818b099-ca34-4e96-8c1c-65b719307360 · outbound

This paper cites Unifying vision-and-language tasks via text generation.

Vision Generalist Model: A Survey Unifying vision-and-language tasks via text generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.528315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.528315Z digest=sha256:2d6d18b4316fab1a3e8932f5e8636188c627239768c3a3484234526a4f407dd4

Observation dcc5ce5f-76eb-4247-8dea-4353601196d6 · outbound

This paper cites Conditional Positional Encodings for Vision Transformers.

Vision Generalist Model: A Survey Conditional Positional Encodings for Vision Transformers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.606136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.606136Z digest=sha256:0f0ba0c841f49b235924d321c513c9849735bdafe5a522a6cd9f53580c1713a0

Observation 394a616e-3668-4058-adba-c3c6372cd344 · outbound

This paper cites Instance-aware semantic segmentation via multi-task network cascades.

Vision Generalist Model: A Survey Instance-aware semantic segmentation via multi-task network cascades

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.694473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.694473Z digest=sha256:7465d297fdc8a37d41875f63cc35e526dc248d22edecf4dec94021eb2f595d31

Observation 5d7f965a-1865-44fc-978d-b74a856fa843 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Vision Generalist Model: A Survey BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.810289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.810289Z digest=sha256:96f586479a8ce0c7b5d31f18a842fa33d04d0bba3324e9efe9a861434c13d2ef

Observation 7c327d4b-604d-40e0-97fe-0b3d6755a824 · outbound

This paper cites Decoupling zero-shot semantic segmentation.

Vision Generalist Model: A Survey Decoupling zero-shot semantic segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.887551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.887551Z digest=sha256:e8edef79e49093758e7bb2b2434b982e0ce81d8653a8e9f688cf86fb57425a66

Observation f74f887e-624d-4bc1-b7b7-c877a16aa6c6 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Vision Generalist Model: A Survey An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:40.947027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:40.947027Z digest=sha256:0e20d8a33a7a3a1bbae85ed2cc7469664a0c2dda00fac389dabb14510f92f5cb

Observation 76496712-ffb0-4f33-a0c8-dcf07cc78315 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Vision Generalist Model: A Survey PaLM-E: An Embodied Multimodal Language Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.092048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.092048Z digest=sha256:cc2c98f87b7d8f20b1e7318027f9bbb361aef2488394eed502c875a9ad7eb491

Observation 27e29feb-4249-43a0-80c4-70078c272ae5 · outbound

This paper cites Glam: Efficient scaling of language models with mixture-of-experts.

Vision Generalist Model: A Survey Glam: Efficient scaling of language models with mixture-of-experts

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.218617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.218617Z digest=sha256:f88f24c568a0d72437a0b81b06f344587872e6f74336b582a4cd4cbe9dfe9fd7

Observation fd737cc9-e891-4ce8-8965-e2194fa52c25 · outbound

This paper cites Multi-modal alignment using representation codebook.

Vision Generalist Model: A Survey Multi-modal alignment using representation codebook

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.312439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.312439Z digest=sha256:83ef397f7c0ed362f18812edb41f773ec7c6f0a7a21e11da470a3ab87503d8af

Observation 8a70b3b4-f3b5-4b38-aa5a-dbcb44eb9a9b · outbound

This paper cites Low resource dependency parsing: Cross-lingual parameter sharing in a neural network parser.

Vision Generalist Model: A Survey Low resource dependency parsing: Cross-lingual parameter sharing in a neural network parser

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.401628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.401628Z digest=sha256:67e94a43f4a14b148ee06134a0d831d12b28d7f6bdd610bc5af4c5a90d1c0b2b

Observation 7f022ef0-546d-47ef-87af-fef343b19160 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Vision Generalist Model: A Survey Taming transformers for high-resolution image synthesis

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.537311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.537311Z digest=sha256:6a471fd118ca2d7d3e2c13cbe4de0c61fbe57cd86328523cd2efef4ea4685934

Observation 3bf1de95-4651-418a-9a3e-43aa77d84d8c · outbound

This paper cites Mmbench-video: A long-form multi-shot benchmark for holistic video understanding.

Vision Generalist Model: A Survey Mmbench-video: A long-form multi-shot benchmark for holistic video understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.642922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.642922Z digest=sha256:a67b6ba84e940c266a95f316c6eff09829bc1429e69f404c5db2735e099cf534

Observation a1c4fa58-0474-4b07-93af-27137f993a10 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Vision Generalist Model: A Survey Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.769713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.769713Z digest=sha256:97f337afbc0219280cfe29ccbde761aab47a213d9685b7953442434025280780

Observation 31a5d409-1186-47a8-b152-dc606bfc46eb · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Vision Generalist Model: A Survey MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.874281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.874281Z digest=sha256:84199412980a020fd17c4cd8cbd70b47a9ff6f1ba98390c71cf19c704a25bff5

Observation 6bce6f40-7fc0-436d-b6ef-710168cd00f2 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

Vision Generalist Model: A Survey Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:41.983897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:41.983897Z digest=sha256:aaa78ca0b56431dd6f129023166fd40dd249e4bfd0a156370206bd990257bd95

Observation 1198782c-3978-4d7f-bd57-b5177a52c381 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

Vision Generalist Model: A Survey MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.097080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.097080Z digest=sha256:f823ac6e22d62ef31612fed74dc8beed117808671c595efdcbfbe333d14f321c

Observation 9aa2c5a8-29d4-4069-bd67-2e16e1ac11fb · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

Vision Generalist Model: A Survey Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.214314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.214314Z digest=sha256:fd84f09eab1e871cab63f62231b1fe258fc6bb28cb2dbdd165d3795dc0a1c0f2

Observation 9539f517-a2bb-476d-8466-6d9f13caac59 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Vision Generalist Model: A Survey LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.303030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.303030Z digest=sha256:f396de946e47417a2e07de0c0ae3b02f9777ad789a34c3dd27b73730f60fecb2

Observation 4c99319a-658e-431f-a42e-278aeaaaed50 · outbound

This paper cites Mtl-nas: Task-agnostic neural architecture search towards general-purpose multi-task learning.

Vision Generalist Model: A Survey Mtl-nas: Task-agnostic neural architecture search towards general-purpose multi-task learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.405980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.405980Z digest=sha256:7ad533f7ed640522c6ee96e868bc670e434854017f11c68285818d72f5e11d21

Observation afa323e0-db9f-45fe-819a-ef9ed949a51a · outbound

This paper cites Instructdiffusion: A generalist modeling interface for vision tasks.

Vision Generalist Model: A Survey Instructdiffusion: A generalist modeling interface for vision tasks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.538463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.538463Z digest=sha256:dd5941e7e32036717209848254099500a03c9df66b7c10a0ae880050966d29e0

Observation a4ba17bf-996a-45b2-97b6-2119824461a4 · outbound

This paper cites Scaling open-vocabulary image segmentation with image-level labels.

Vision Generalist Model: A Survey Scaling open-vocabulary image segmentation with image-level labels

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.648552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.648552Z digest=sha256:bf3e6fe5d35807a70f9f4be5dd51cddcb94dec61743e32e8707ef67a3c900f93

Observation 59e09b5c-2843-4fe1-8a00-3791ea115380 · outbound

This paper cites Omnivore: A single model for many visual modalities.

Vision Generalist Model: A Survey Omnivore: A single model for many visual modalities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.724865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.724865Z digest=sha256:b28a08ceaeb7d77f0808fec4ca1cd45a7be0768d5c0ade97754068907d519bdd

Observation 64ad3603-6dd5-4182-8e21-9d7114ca8a36 · outbound

This paper cites Mixture of experts models.

Vision Generalist Model: A Survey Mixture of experts models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.783410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.783410Z digest=sha256:7878783d4e668eef8920b2e4e6702d94e9e5e74dc2c19f1a13abde4820a0744e

Observation 108fbebd-7adc-449c-843f-58344ee2372e · outbound

This paper cites 3d semantic segmentation with submanifold sparse convolutional networks.

Vision Generalist Model: A Survey 3d semantic segmentation with submanifold sparse convolutional networks

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.847069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.847069Z digest=sha256:52671f7b722512d6048d83f194d41e571d740db6c716a4e87400efc3aaa0235f

Observation 6c8fcf7c-9e11-41e3-988f-e7aae0c1e61c · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Vision Generalist Model: A Survey Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:42.929374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:42.929374Z digest=sha256:c7f5167b985eec4c3578dcb8e649726944446b22fb3798ece7750edf839a16fb

Observation 524f39e1-6351-4720-8dde-b722bf31fbc3 · outbound

This paper cites Deepfake video detection using recurrent neural networks.

Vision Generalist Model: A Survey Deepfake video detection using recurrent neural networks

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.022561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.022561Z digest=sha256:4150bcfd7889e2004e3456b770ca2cc0259f27ea72481940cec5e96d08620141

Observation dfc7ad55-1a4b-46bd-85ba-0aff07b70425 · outbound

This paper cites Dynamic task prioritization for multitask learning.

Vision Generalist Model: A Survey Dynamic task prioritization for multitask learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.122801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.122801Z digest=sha256:1b74bbd28d7bb1a6b4401d19da7268ac04cca4cd64bf24e2b0da90686082901d

Observation 94e049cb-e175-4c4b-84bd-d7ed481aa2a6 · outbound

This paper cites GRIT: General Robust Image Task Benchmark.

Vision Generalist Model: A Survey GRIT: General Robust Image Task Benchmark

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.270853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.270853Z digest=sha256:75a8ebec6dd1237c6445115ba2240647525e8d713d7c2802a154ae97374896e7

Observation 5909fec8-b3bb-47b2-80de-f956ce27527a · outbound

This paper cites Onellm: One framework to align all modalities with language.

Vision Generalist Model: A Survey Onellm: One framework to align all modalities with language

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.392291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.392291Z digest=sha256:5c0bebf53ea344f8c80acbf6e3c524042999a21f04625cf6bc3be62b3d1eb9f7

Observation e867694b-33ca-4192-8dcf-810361659346 · outbound

This paper cites Deep residual learning for image recognition.

Vision Generalist Model: A Survey Deep residual learning for image recognition

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.490926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.490926Z digest=sha256:3d6c01e46fdb483f463264e18b8d17353250058a3cf18a0dab7ec99c1b653d60

Observation bbeb827e-2db2-4ef1-863e-5886a6606636 · outbound

This paper cites Deep residual learning for image recognition.

Vision Generalist Model: A Survey Deep residual learning for image recognition

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.588569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.588569Z digest=sha256:1fadd51707409998ee3edb0a3d76c8591c88f6e6708e46674a2e51169105d3fc

Observation 7f6f8551-b070-484f-b446-4d752e3b9172 · outbound

This paper cites Mask r-cnn.

Vision Generalist Model: A Survey Mask r-cnn

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.728282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.728282Z digest=sha256:0bce3a9cf26ad42c9f1c749b796ea0b56ef2547818b23a7f8d4dfc2326a6cd8e

Observation 6184590f-6af9-4971-a072-c13aab04b41f · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Vision Generalist Model: A Survey Distilling the Knowledge in a Neural Network

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.821811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.821811Z digest=sha256:01c0a7891fd5314beffead274f738e52e0e41802ad7ad787b057cface3d23843

Observation 832914b1-78db-4b26-9364-e372c3f3f8a0 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Vision Generalist Model: A Survey Classifier-Free Diffusion Guidance

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:43.966930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:43.966930Z digest=sha256:3452ff18b4b1fe5c23875199e03b7e6ddb59d5f2f6aa66ec5860596e19074b03

Observation 1e3ed28a-94cc-4a54-be49-322a99ba2435 · outbound

This paper cites Denoising diffusion probabilistic models.

Vision Generalist Model: A Survey Denoising diffusion probabilistic models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:44.158944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:44.158944Z digest=sha256:d817375c8202195743a46aa5c2ec6521877b3d461ee6fd5b0ee852fd657f21f3

Observation 3c883db3-43b4-4db6-9b60-9ea54adb2b37 · outbound

This paper cites Long short-term memory.

Vision Generalist Model: A Survey Long short-term memory

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:44.383971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:44.383971Z digest=sha256:70aeeb7ba553914ef9d29e65071c7d25f11907a4e8257d1184f57965c06c8363

Observation 14bc7247-bcb2-4667-809a-57246b0a91cf · outbound

This paper cites Unit: Multimodal multitask learning with a unified transformer.

Vision Generalist Model: A Survey Unit: Multimodal multitask learning with a unified transformer

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:44.546997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:44.546997Z digest=sha256:4fac5fdae7dd5daa0ee80fc20e55484a9809e08bf1a5341fd157e8640d16821e

Observation f25d31fa-61a2-4121-a7cd-e61f9fc7b8b5 · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Vision Generalist Model: A Survey An Embodied Generalist Agent in 3D World

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:44.771145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:44.771145Z digest=sha256:d1443e03d06d12083531f4260ddbd05b2e540acf1e9e6eaf3e26c462e7dc712e

Observation 109d6306-2545-47c5-899b-c9395d515211 · outbound

This paper cites Language Is Not All You Need: Aligning Perception with Language Models.

Vision Generalist Model: A Survey Language Is Not All You Need: Aligning Perception with Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:44.973382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:44.973382Z digest=sha256:6092a8e8a8b1257b18ed1d558d63daa80bec7e7013b892d77b2ea81b09709fc1

Observation d1552308-4ba0-4b04-b8f0-3a44805f168c · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

Vision Generalist Model: A Survey Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:45.149991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:45.149991Z digest=sha256:7c4d11c6181c6348f5165504b6ca6b90bfb301d270315e1c1b8500ed44cc22cd

Observation ff85f8d2-94f1-4363-9875-2957dcc4ffa8 · outbound

This paper cites Perceiver: General perception with iterative attention.

Vision Generalist Model: A Survey Perceiver: General perception with iterative attention

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:45.347606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:45.347606Z digest=sha256:56e50073870d55bba4984df49bd71b388e6a27d14355efe54123442872cabf30

Observation 06cfafc9-e956-4a90-a212-cbacda38b73a · outbound

This paper cites Oneformer: One transformer to rule universal image segmentation.

Vision Generalist Model: A Survey Oneformer: One transformer to rule universal image segmentation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:45.503099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:45.503099Z digest=sha256:5cc81a2ea7f45d660db0a5ba4b12ec2cc0155aad954538ca933b641ba9675df5

Observation 283d2273-548b-42f1-8710-1b6284021692 · outbound

This paper cites VIMA: General Robot Manipulation with Multimodal Prompts.

Vision Generalist Model: A Survey VIMA: General Robot Manipulation with Multimodal Prompts

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:45.709626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:45.709626Z digest=sha256:882882a7202c4d73a73952cd214e5acfcf642719d5cae75c843531fd1c04ce58

Observation f4c42f47-a091-49cf-9006-eedc72a1839c · outbound

This paper cites Unified language-vision pretraining with dynamic discrete visual tokenization.

Vision Generalist Model: A Survey Unified language-vision pretraining with dynamic discrete visual tokenization

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:45.853890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:45.853890Z digest=sha256:40969a6060cbb620c961532e06ee58a07b756b83cbd0fa927d1ab920e0c06e11

Observation 19153742-5e8b-45c7-8dcf-9fa41894dff8 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions.

Vision Generalist Model: A Survey Deep visual-semantic alignments for generating image descriptions

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.072821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.072821Z digest=sha256:d4fa5a8d34781c701d92314f25c3fb80e10cd44e112474a18890b60e7092c02f

Observation f62d91b2-1762-4e9d-94ad-f0c8621573a0 · outbound

This paper cites Multi-task learning using uncertainty to weigh losses for scene geometry and semantics.

Vision Generalist Model: A Survey Multi-task learning using uncertainty to weigh losses for scene geometry and semantics

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.234329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.234329Z digest=sha256:e598e0a103fa730efd4b52625ab4ff341d2efc663e427fd5ff228e80cffe3023

Observation e66a144c-082c-4c5c-abfb-8a2ef80199ae · outbound

This paper cites Segment anything.

Vision Generalist Model: A Survey Segment anything

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.394972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.394972Z digest=sha256:2856c1a5cebba97e612acdab2ff9cfecc1228c502f74179df654e95559501763

Observation 56e22b71-6341-4b29-b253-d599cf106d49 · outbound

This paper cites Segment Anything.

Vision Generalist Model: A Survey Segment Anything

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.544041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.544041Z digest=sha256:7173b36ca7ba0c20b4ff82123f6371d46cdf0bfa14bf31c2e0e65f4ec97712cc

Observation 3e65f32f-fda9-4c12-9da9-ba44a80e60e4 · outbound

This paper cites Uvim: A unified modeling approach for vision with learned guiding codes.

Vision Generalist Model: A Survey Uvim: A unified modeling approach for vision with learned guiding codes

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.806244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.806244Z digest=sha256:7300f40677f1323020b6d50c45b1d766077fecb6bade77915f8fccd94eea12a5

Observation 05aee257-12ce-42cb-a0d3-d062bb16a7c9 · outbound

This paper cites Tabddpm: Modelling tabular data with diffusion models.

Vision Generalist Model: A Survey Tabddpm: Modelling tabular data with diffusion models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:46.982104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:46.982104Z digest=sha256:2e891e53366c9c549a110d3399d2d5c2b1090844309ad9bb839cf6b573862ef1

Observation 722487cb-9dd2-4d78-97f4-8292a40b8eb1 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Vision Generalist Model: A Survey Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.129273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.129273Z digest=sha256:8e55f8334dbe5d664f2e3b28f392911c4777e23c59dd94d4685ddf56777dc64e

Observation fcf16247-114b-4183-ac79-25779cbfa097 · outbound

This paper cites Imagenet classification with deep convolutional neural networks.

Vision Generalist Model: A Survey Imagenet classification with deep convolutional neural networks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.280227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.280227Z digest=sha256:bad005074902a00f1c4f720b03c3bba77399ada677711170946a3d88e9f205f6

Observation 60f0e233-0d61-4ae9-846c-abe12bd7ed1f · outbound

This paper cites Babytalk: Understanding and generating simple image descriptions.

Vision Generalist Model: A Survey Babytalk: Understanding and generating simple image descriptions

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.423323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.423323Z digest=sha256:f8eb9dda6f69e9466bac1c248b7bd3e29f0c928b0abcb8bfa4a07c29131428cb

Observation 12beba6b-9608-4375-aee5-8cd89e2565f0 · outbound

This paper cites Base layers: Simplifying training of large, sparse models.

Vision Generalist Model: A Survey Base layers: Simplifying training of large, sparse models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.581848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.581848Z digest=sha256:57e0c1e7f3bc535a4b8474e6fda457cb61d5dd318b611898c097b683ccdb6206

Observation d7a90b71-5da5-46df-af46-7e3b8af095e2 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Vision Generalist Model: A Survey SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.735480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.735480Z digest=sha256:cdbd6ec849a08ac670d09e5e92ac724226476c0c72a9c1a855c477176d4fa9c9

Observation 594b128c-0f69-44c4-a94c-102c4f6df394 · outbound

This paper cites Multimodal Foundation Models: From Specialists to General-Purpose Assistants.

Vision Generalist Model: A Survey Multimodal Foundation Models: From Specialists to General-Purpose Assistants

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:47.872697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:47.872697Z digest=sha256:d1746d492aad6fb57898c756a67edd3d8e009835340d4d248ae98614bc00efcd

Observation 12500a93-70ca-4ca8-80c0-8f904f0bac2c · outbound

This paper cites Mask dino: Towards a unified transformer-based framework for object detection and segmentation.

Vision Generalist Model: A Survey Mask dino: Towards a unified transformer-based framework for object detection and segmentation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:48.106927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:48.106927Z digest=sha256:a0407d1408bc67c77563ee30d0d785f451be4619149c7a5c8b2be412882d8d00

Observation 1a796a68-0850-4bff-9da1-2123fcf53024 · outbound

This paper cites Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks.

Vision Generalist Model: A Survey Uni-perceiver v2: A generalist model for large-scale vision and vision-language tasks

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:48.304969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:48.304969Z digest=sha256:417a73db3996690425322f8a24b9298b7f06c34c8c0c54b9eb42e48a2c6cb968

Observation 42360ffd-43e2-4579-a9fe-22f6808caddc · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Vision Generalist Model: A Survey Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:48.536824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:48.536824Z digest=sha256:7e4ec8608b73ac7101f78e62cd90d311b31cf12045965cdeca410dd1db9bed0b

Observation 37db749b-7cc5-4cda-a039-5ffb7e93afd4 · outbound

This paper cites Lavender: Unifying video-language understanding as masked language modeling.

Vision Generalist Model: A Survey Lavender: Unifying video-language understanding as masked language modeling

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:48.684658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:48.684658Z digest=sha256:2e5c51e592931fe86da97d9e2a2014261de7ee42ccb2318f7f4a2f2d4afa8ea8

Observation 1327ed3c-d85d-480b-ac66-2fc4e89e18cf · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Vision Generalist Model: A Survey VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:48.858023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:48.858023Z digest=sha256:0163db726cf25f17f1975e0f7104b9efb889b86c2d991710e474923d3253740f

Observation b42aa6ee-2894-45bf-99dd-8907276e4d7f · outbound

This paper cites Grounded language-image pre-training.

Vision Generalist Model: A Survey Grounded language-image pre-training

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:49.054065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:49.054065Z digest=sha256:c58900563440db9724e4c5a4e410883a8d35c9b7a1028e96af64e34b2a44c496

Observation 779fed7b-b38a-4d86-8a5e-eb7c7f3a422e · outbound

This paper cites Omg-seg: Is one model good enough for all segmentation? In CVPR, 2024.

Vision Generalist Model: A Survey Omg-seg: Is one model good enough for all segmentation? In CVPR, 2024

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:49.218178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:49.218178Z digest=sha256:b56f723b3ef38022f01b3ef308d8607fe698e72993845938d67d35c2e117164b

Pith citing papers

No inbound Pith citation observations are available.