Pith. sign in

Paper Citation Record · LEDGER

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

As of 20 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 2 inbound Pith citation observations for arXiv:2506.10395.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10395 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:34:12.154340Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T19:17:19.634242Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T18:51:15.880853Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved65
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eebf5b90-1bdb-45b1-a9e5-af521b65bd51 · outbound

This paper cites write newline.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.234925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.234925Z digest=sha256:020c72a96d3985bf9e36dfc182cc2ea5f97a9b35754e475b145c34aab0aa9c63

Observation a020a70d-a6a7-4c98-9182-37b551e521db · outbound

This paper cites The Llama 3 Herd of Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation The Llama 3 Herd of Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.321702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.321702Z digest=sha256:156d036be9119543198698677661a53109066bd8e48a9397d1a0f683c118af91

Observation abc4ea03-17c9-4738-936f-47ed169a939d · outbound

This paper cites CM3: A Causal Masked Multimodal Model of the Internet.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation CM3: A Causal Masked Multimodal Model of the Internet

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.440871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.440871Z digest=sha256:fcd179701507cd99289d09761966c3a16f6b2d50c8b62ce79a3b7d58917ef8fd

Observation 28a22b10-b9cf-4c6f-9bdd-35e81f4db077 · outbound

This paper cites Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar \' e n Simonyan.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Menick, Sebastian Borgeaud, Andy Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Kar \' e n Simonyan

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.594642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.594642Z digest=sha256:d7f6ad2513c57d10b1a8a30d45feddb582d3f0a876a7162a05e25db56dfe924e

Observation 9c6ba7fd-8ebb-4380-a453-fdc378c12828 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:59.735365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:33:59.735365Z digest=sha256:1c7625bd844f658c5fba9635474e08e7767e82d419f6ab3a48d793ccc8e97ac1

Observation 85ca09f3-7f82-4bec-868e-88e0543cf97a · outbound

This paper cites Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.049480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.049480Z digest=sha256:5ec5d51c0008640649b5bff6fbe676b7f642ffd9c88ce82dae5fca25eb3694ed

Observation 0f708b75-a8e2-4530-91a4-596384414343 · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.161811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.161811Z digest=sha256:d3a2b5a35d7c3cd3ac776480c7ffc96d4751f0b5bc6b3a0392bd12209739892e

Observation 2af2ac08-d9d3-4a7e-823b-559c41fbd359 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.326020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.326020Z digest=sha256:b37201f9f102821781593f623999c7d91ecb56fd81d9b5bad72d22e29c54c5c6

Observation 810f32d4-c9b2-4610-8376-0015a635568e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.478226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.478226Z digest=sha256:fc9864e912b0f8f78a7f74463e20869287cc0421c808c49f84edd1314ac1f2ff

Observation fdd57776-9a51-47f4-9f03-ec46aa1e85e7 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.637220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.637220Z digest=sha256:53fc3b2a5831bb752f3f8df8dbdc88d162863d2bda4a705eea8ecaef5878f8da

Observation 282b3c3f-0bf0-434e-a7b2-4d2c58bc8d31 · outbound

This paper cites Dream LLM : Synergistic multimodal comprehension and creation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Dream LLM : Synergistic multimodal comprehension and creation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.829274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.829274Z digest=sha256:ee468effa9427bdae90de862fea3a841052c30360e885e478e4a87a952047297

Observation 31c365d3-95ab-449f-a0bc-836c440bcaa2 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Taming transformers for high-resolution image synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:00.945047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:00.945047Z digest=sha256:a93209ab656a55d35165f3f2526179201261b9d6b0856b94b42215e18107f312

Observation 0a84b836-d5e9-4f8a-b917-d0214a4a70ce · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Scaling rectified flow transformers for high-resolution image synthesis

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.805912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:01.070818Z digest=sha256:6a270981ea37bf37c54fdbd8e1f1f329b9f03d0dedb12e6f2076debd99c06d0e

Observation 005b0f3d-381c-4faa-8df6-bd1e16f6c3e6 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Mme: A comprehensive evaluation benchmark for multimodal large language models, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:01.269001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:01.269001Z digest=sha256:1214682856fb95bb0d94ca6247558d4fb10f14f10d9230b43ad34e379b50b2de

Observation dd691fa1-fb58-40bf-b061-aef520ac4b56 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:01.423591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:01.423591Z digest=sha256:08606cb8676a81d3764e593b77d806651980dfa3cf7d0c97c0de63d84ef8728c

Observation 176c3921-7018-4cdd-8b03-1bf1c652af5b · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.610344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:01.595680Z digest=sha256:a61e013b429eb67d2da39f64c51855fe3e74b48d039953960daf72d1438d325d

Observation 8d0a0171-406d-4d70-be15-7df991db1f82 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.408391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:01.784116Z digest=sha256:5868559dab2afe988c46ade8464312294e314ec2181ef235af0711143cd7eb68

Observation 4cc5fe1e-282d-46d5-9df9-8d11ba897449 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Vizwiz grand challenge: Answering visual questions from blind people

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:18.190173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:01.983831Z digest=sha256:b40fb55c5f7f3cad5cbf92c9f12d2f3209164471eff51d7ad5da5d77e2364967

Observation 3ecddb5c-dceb-48f4-9232-86434787a8e4 · outbound

This paper cites Masked Autoencoders Are Scalable Vision Learners.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Masked Autoencoders Are Scalable Vision Learners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:02.155166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:02.155166Z digest=sha256:cb0c8f10ed0cf8ba20ec180112da1a76450093bac4f173c972b7a01e41568f1d

Observation 29f63196-c174-40ca-bc2c-c11ce4af7382 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.952933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:02.302887Z digest=sha256:fbc414347c7f825d1c84cdd43d1febfa37edb9a91127d016f157bbe6a59d832b

Observation c7a3e172-d7a5-4053-9f1c-a43cb76fc897 · outbound

This paper cites Unified language-vision pretraining in LLM with dynamic discrete visual tokenization.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Unified language-vision pretraining in LLM with dynamic discrete visual tokenization

Reference 22

Resolution
verified exact
doi, observed 2026-08-07T04:34:13.304814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:02.418863Z digest=sha256:ad7ff4a8d578e9f9d39ec58555f4c44af454de99a3492c798356f7dca47f8d58

Observation 61a80a70-4b79-4239-9a52-ac12d2f74da1 · outbound

This paper cites A diagram is worth a dozen images.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation A diagram is worth a dozen images

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.721711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:02.539819Z digest=sha256:f6bc90d1e1e5dac08cb7864c680f9c978e04dc5b17e47f2c3a8794dc9b7da637

Observation a52519a8-677c-4275-bdfe-737ef7d2a1d5 · outbound

This paper cites Generating images with multimodal language models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Generating images with multimodal language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.490002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:02.755733Z digest=sha256:77ddf87e0a427c7faea4711777fbf3f6944870ce108719343427ee41bab780c1

Observation e870ab70-d91a-408c-ad38-3c090f0f4a8a · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:02.916132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:02.916132Z digest=sha256:f5fe9e83349e1e28ff1c9852013f4071289f2c0cb2f7072cf32b07058faa6fc4

Observation 79fde0c5-1a91-4e10-926f-c9d7b402a850 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.070205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.070205Z digest=sha256:58e1f6bf7079e5e74811712c055172614b65f5e51dc6f84c2971c5887b846593

Observation 12899f82-3563-44a6-9a78-941a20b1afd3 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.206849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.206849Z digest=sha256:7e43a8cdae425ee8bd3611cb293cc632bc62de5577d2822e9048b3d6f3613b5a

Observation 273f10d9-7205-4d87-8284-ca2e3c2ef32a · outbound

This paper cites an unresolved cited work.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:34:17.290173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:03.339774Z digest=sha256:e2950671b89afc43196a782975c0741b0530f6317cc8f538980019e4b4a51d10

Observation 9a25378b-887d-4c8f-b888-4c69caf597a7 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.476230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.476230Z digest=sha256:32156a306e1055ca532f7a72bdbdc7196871350231d4e90264b2d7e2ac6d8ed8

Observation e07f2fb3-48c4-46de-91da-d8bf02d63ddb · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Evaluating Object Hallucination in Large Vision-Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.678107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.678107Z digest=sha256:7f1f5fed6d1e79434e691ee596432e8cb453189ce5f330c98dec2b8f2d0aa883

Observation c4ad44f6-9d62-4bf8-baf6-b5e6f06e999c · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \' a r, and C.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll \' a r, and C

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.799566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.799566Z digest=sha256:e0d06a5f72f6db34b8c9c14716b8ff8f4f095b9d50281b7625fd257f75bf82d7

Observation 98e74cdd-e023-4e57-82bb-a755b74a2fad · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:03.967311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:03.967311Z digest=sha256:40aaddad0adcf91236be21fc0158ec7fc1bf725c6cadd45e685dd1ac57e6e519

Observation f8a241d6-d815-4b89-b308-3c81e11cf694 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Improved Baselines with Visual Instruction Tuning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:04.091274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:04.091274Z digest=sha256:a9b7882d73538132f9e83029fa013a024f959a9cdaf97b19a0c20e52fbc38734

Observation fc9d2ee5-14e5-4b01-92b2-1e4f746c6f5e · outbound

This paper cites Visual Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Visual Instruction Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:04.316929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:04.316929Z digest=sha256:13f9edcd138516a71c24caf3d80a74c1b641afd8b279bc1fe3cae02d9be2d5db

Observation aa3dcf34-721e-4f12-8389-df32b9d9580c · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Llava-next: Improved reasoning, ocr, and world knowledge, January 2024 a

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:17.057930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:04.448654Z digest=sha256:6b12f4aa5a4bb52d4e685b9e9353e8e7656704581f836e13e19622ae90418f23

Observation ce5fc821-db12-452d-8655-ac7a47e1b085 · outbound

This paper cites Visual instruction tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:04.568020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:04.568020Z digest=sha256:9b6457bbc1bfff0ba0cf89087efb363fc279ff3b750e3f6a8352f6ac55220b99

Observation 396b95b1-89d1-4a06-9d74-74649dc13acf · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MMBench: Is Your Multi-modal Model an All-around Player?

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:04.747893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:04.747893Z digest=sha256:720332d558365a8e906ed0a9e89f2d16d5541384af8974b6cfd78fc397e3809e

Observation e6ea8826-9138-48fc-b20d-1235829373dc · outbound

This paper cites On the hidden mystery of ocr in large multimodal models, 2024 c.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation On the hidden mystery of ocr in large multimodal models, 2024 c

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.931157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:04.934626Z digest=sha256:4d8474328bda16edb670793eac51f73106b82455cf96a077c38abc02c4d7b688

Observation 74ab8d04-505b-4eb6-a6e8-66c7fc7cfcfe · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.044523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.044523Z digest=sha256:615522fe3cf13fc670e58edc9f39a99f8f7fa41386adf356b2ec26b5d0d5410a

Observation b37d7701-e2da-4f1e-88e1-4496d035711c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.176742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.176742Z digest=sha256:19147d32927450011ead25f4f1f3feebfd371ac8df5dcb5f9780fd381536b3f8

Observation 87ae0daa-a108-4529-bfcc-022c9b0aa4aa · outbound

This paper cites Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Macaw-LLM: Multi-Modal Language Modeling with Image, Audio, Video, and Text Integration

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.315938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.315938Z digest=sha256:27011f27bbff1537a553432f0f2771e2953e03750947de01c29f376ac5e7239d

Observation f1179f3e-2376-42ea-8792-0760c0b39fc7 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Docvqa: A dataset for vqa on document images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.474321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.474321Z digest=sha256:02d17225f46da2c0c982810f5abf73f455651f018e85afce3ba843ac3871fffe

Observation f0ddd885-f659-40fb-a918-8f9af08c3c99 · outbound

This paper cites Infographicvqa.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Infographicvqa

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.634106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.634106Z digest=sha256:1f4bfe69ba9ae418d519de37ec61fe746e6117fee1bcbf12bb46c8d43e6970b3

Observation 2838699e-86a7-441a-b0b1-7da3caea2a41 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:05.794775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:05.794775Z digest=sha256:635e3785fda2659ac65f55898c935848deac972dec50a6b52fd5a9f889b4ce5b

Observation c74be96a-30ec-4c99-aca8-4fa8592c5a13 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Learning transferable visual models from natural language supervision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.730462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:05.937934Z digest=sha256:b6c95beb3e542003e18ae0443f025b755de37d7088deb03a5f7bee28268b0c86

Observation 20385547-d54e-4e6f-ad8d-2dfec6e24874 · outbound

This paper cites an unresolved cited work.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:34:16.439186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:06.067878Z digest=sha256:011b258117296002182b5806f2133039df170c4df01cc2b24acf2f3b78b260f8

Observation 04974964-d317-4823-be2f-d05af2983131 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:06.226877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:06.226877Z digest=sha256:a12b77893c7d65e578ac064cdc783ad0b6c45c313ea15d181cec97126f2c1b18

Observation 607662ee-0c37-4bf5-a7bb-6f381a55fe69 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation High-resolution image synthesis with latent diffusion models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:06.371337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:06.371337Z digest=sha256:30be2c1b9a4d4eb6aab091bcadaabc5701f5bd66bf448a58be2ceb5ca87905a3

Observation 143f32cf-f12e-4123-b98b-f3ab38e3a9b0 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation High-resolution image synthesis with latent diffusion models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:06.504245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:06.504245Z digest=sha256:392d2dc9b194b98c94f3ef2efbbc3826bcd2762609a27740b3f0d8372bb2a912

Observation 1ee84ecc-3961-4ef5-a0ae-76d81a0a4246 · outbound

This paper cites Multimodal Instruction Tuning with Conditional Mixture of LoRA.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Multimodal Instruction Tuning with Conditional Mixture of LoRA

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:06.624192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:06.624192Z digest=sha256:e5c5b6d8b767aa0ab60cbc1085f4a880b6987543d99dd4c0ab00744ac4ad6a29

Observation ec43e49d-4005-475c-9c85-60df80513c47 · outbound

This paper cites Towards vqa models that can read.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Towards vqa models that can read

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:16.178169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:06.792015Z digest=sha256:550fdf08b0ad9df9dde70199111af7858ddac0c2097f5a69a81a3ed88cd1dc63

Observation b455ec78-c8a9-44bc-9ca5-25d79365c58f · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:06.956907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:06.956907Z digest=sha256:5fa603c702671bacb3f33d7bd86948db6fb6d87cdb947820af533d26ca485944

Observation b6fb765c-2f9b-4a37-8260-721aab517faa · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:07.113869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:07.113869Z digest=sha256:01ba57d8882be1289aed7d6c33615af533b419b33c99089fb67b10624e36b4ba

Observation b107e098-ca99-4fa3-9c50-ba1aa94d0951 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Emu: Generative Pretraining in Multimodality

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:07.236721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:07.236721Z digest=sha256:41d40220d0b47991f789de007394c78908b8dc9e611617fbb90c906c8bf563cb

Observation 9157b874-4a5e-4133-bc4e-ebce19cfbb08 · outbound

This paper cites Generative multimodal models are in-context learners.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Generative multimodal models are in-context learners

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.826498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:07.381855Z digest=sha256:a34d42794e388e00a7e3879ebb5eef47cfda6fc3fc3ac84b7c932e21e7587376

Observation b6b0a231-7596-49d1-8b53-45a7d0f0e3ad · outbound

This paper cites CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation CoDi-2: In-Context, Interleaved, and Interactive Any-to-Any Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:07.537541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:07.537541Z digest=sha256:36096b2827b9a5b4f64dcddc001842cb68774045183d7629f554621fca457b35

Observation ff032ef9-9d8c-4eac-9f1c-0536508ec074 · outbound

This paper cites Any-to-any generation via composable diffusion.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Any-to-any generation via composable diffusion

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.538723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:07.678452Z digest=sha256:4b765b3d74fe56f94d498635043a17a0b8a54da81bd4239066cad64b5259a335

Observation f3aefd29-efa1-4137-8cf6-000253965e3c · outbound

This paper cites Chameleon: Mixed-modal early-fusion foundation models, 2024.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Chameleon: Mixed-modal early-fusion foundation models, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:07.871098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:07.871098Z digest=sha256:7b453c095494e936484cd02aeac890ee762c0353a3013c122cb3736c24c3e87a

Observation ff2de168-136a-4a9c-a1e1-8385b4d6a2fe · outbound

This paper cites MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.000903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:08.000903Z digest=sha256:803ba77fee4b71441b11574ff1a163509d220692c00ebc3fbbefc78bf3b9f99a

Observation 0129f496-dcae-4c81-bc71-311c6cc4aecf · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.124461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:08.124461Z digest=sha256:6c482d5f4d515074d28b6267a8a0546a4c38aef8a95a4b09d772bd94bf8024b1

Observation 0b5b9bf6-2cc9-4da9-b2db-d3b84535404f · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:15.221743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:08.316611Z digest=sha256:2eacc12c7fb94470873e541d9ebbc1f0e8c619fe45fad05a552fcd2c05875137

Observation 969ed406-5d8c-4ea7-9be1-56d92ad587cb · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LLaMA: Open and Efficient Foundation Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.477979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:08.477979Z digest=sha256:d09ffdeea5fd276ba51540e812896a971a43887df3fef271b6146ba77a6b29cb

Observation 0ffe2784-43c0-407f-87e9-88ba41cf170c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.648626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:08.648626Z digest=sha256:a131028fc3a02cbb93a26de7f7ef3b9da5d44eac8de856d11aa8ec1be14b484b

Observation 7d658a2d-493f-4c3d-8b4e-c7ef1a5e2a14 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:08.755303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:08.755303Z digest=sha256:c637afa6776f17760c53e27d4e7f7fba56d75fa2e952941f3f764d5c32b63685

Observation 2219aa29-5f93-4240-afd0-11434893aa1e · outbound

This paper cites OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.903752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:08.885112Z digest=sha256:2be542e907dcc1da1e4cdadab72b5ed619429313b0c08ca79afaf11de3cfef4a

Observation 7c8c064b-8e4f-45de-9a24-3a57bbc9bddb · outbound

This paper cites Image as a foreign language: BEIT pretraining for vision and vision-language tasks.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Image as a foreign language: BEIT pretraining for vision and vision-language tasks

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.022182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.022182Z digest=sha256:ba4d24afec793eac36f4e39969abbfda38b04977b8e24abe6f4585b41b7330a7

Observation 1d99262d-5b9d-4c38-be12-0575e788c3eb · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.192768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.192768Z digest=sha256:452865c09f0e9d9699138100e9f1c9f1dc79e85df78e4a8c1f785dc7a4e61b8c

Observation 974ce6fa-37c5-4c72-883f-1f57d968f342 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.374451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.374451Z digest=sha256:14f909ff71def1018b71d7b75a3070ab42cf6e4678ee992bdb922dd9537c2ff7

Observation d18c754e-cc81-41f7-95a0-bf11ed74dfb9 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.517877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.517877Z digest=sha256:9715c0f72b17ac3d3a2c948fef2a3927a7eb7466b70015e354b9f60bc43018ce

Observation aa67e46c-d63e-481c-8fef-55eceba2b082 · outbound

This paper cites Grok 1.5v: The next generation of ai.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Grok 1.5v: The next generation of ai

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.727112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:09.703995Z digest=sha256:9df2029fb7c9f53e92681eea926805d490442ff410db595620ae01e4a23b8f6a

Observation ff7c1f67-5593-4f19-ae75-b98806a9f200 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.844460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.844460Z digest=sha256:6e89376868477c51eb8b94b4c202ca7d5d244322f763c25dbf25385ff03c65b3

Observation d30a2810-ca8f-4506-841d-cab57a499d09 · outbound

This paper cites LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LLaVA-UHD: an LMM Perceiving Any Aspect Ratio and High-Resolution Images

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:09.966497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:09.966497Z digest=sha256:fb126d30eeeb5d1688d79402a79144666f5b699b10128277d4f91ead071fac3e

Observation 877fdd3f-effc-4341-8f7d-aee0f9d1ff3a · outbound

This paper cites Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Multiinstruct: Improving multi-modal zero-shot learning via instruction tuning

Reference 73

Resolution
verified exact
doi, observed 2026-08-07T04:34:12.568692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:10.109785Z digest=sha256:8bc7ee4ad91053205d30c97b869db16bf54133013176a6bbbae079a4cac808bb

Observation 60240428-c3d7-42de-b785-8280a413f474 · outbound

This paper cites Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Vision-Flan: Scaling Human-Labeled Tasks in Visual Instruction Tuning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.263657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:10.263657Z digest=sha256:103765ae85c6aaedc5f6d457ab4e8686dbfb18df33fc20f0ac75579005396977

Observation 1cd2c1f0-cfcb-4753-a007-e055acbfb868 · outbound

This paper cites Modality-specialized synergizers for interleaved vision-language generalists.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Modality-specialized synergizers for interleaved vision-language generalists

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.442097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:10.391353Z digest=sha256:77cc0bc9bbfa5a4f8296af0e27bce0026b72b08360f84cdeb5d11020b0204df8

Observation 8f22ff07-7612-4a2a-b947-666c44328754 · outbound

This paper cites Retrieval-augmented multimodal language modeling.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Retrieval-augmented multimodal language modeling

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:34:14.215441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T04:34:10.582907Z digest=sha256:bf83a6e1782d55e0f6ee295a66b116a743b1c98a05ea5cbc3956adc7a1700772

Observation e7c6d5cf-9f27-42b2-ad58-643a88ed7b46 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.749553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:10.749553Z digest=sha256:735a6ffa3f953ea23e0d70b4f59cea3f854aae78fbab249dac0969ece4815e22

Observation eb2dfaed-8279-4d42-9cac-c2fa98baea49 · outbound

This paper cites LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:10.895265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:10.895265Z digest=sha256:110a4402fba64f07a31d108dbcbdd7fe22e4dddb2511913eaa7f1095a85b3ef6

Observation e9b4edd9-af5c-4184-8228-db135bf16a19 · outbound

This paper cites Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.061508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.061508Z digest=sha256:ad05e167aa8fd76664d400a1fe7415861797e9b53e7f1fff18ed385c8a1abdb7

Observation 825db274-4934-4ccb-9e22-98aa3e98a425 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.250107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.250107Z digest=sha256:2ca37a3207dbd9978e84f4ef69943049a2d747e7b3eededfac220d15f1f10886

Observation 12e13317-6f53-4ab6-ad1b-3267e3526906 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.405831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.405831Z digest=sha256:6cb96f7b8e74ff40db7abc51cf784be937c848c2f1d18b34cd3e14b052817d5b

Observation 1f35042e-3a44-45ce-86a2-9093cc06b9a2 · outbound

This paper cites Sigmoid loss for language image pre-training.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Sigmoid loss for language image pre-training

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.533516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.533516Z digest=sha256:6bb1373ebd1a61951972539aa248612f2711a48d89bdda3261ef848abf0184bd

Observation c35cc7c8-52e8-4e1f-897b-b484e3188537 · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.720293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.720293Z digest=sha256:32fdeb2136434d0fdb748e609eecc078b3630da271be00584533a33660c58b62

Observation 1137c489-997b-411f-82db-76314dc056c9 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.878440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.878440Z digest=sha256:2e12d519cda03c6bf45690ed2433e921edc54922c151d8354eeca57e666602ac

Observation df24cd45-8a6c-4865-9926-9dcdf3affe65 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.006826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:12.006826Z digest=sha256:67c084da92b5a50235ec61852d279ad4c40fae324436316775043e8568072b6e

Observation 9c091c59-76f3-4992-bda8-50607a9d87c7 · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:12.154340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:12.154340Z digest=sha256:02f4d361df9b3b0da9c8c7e114992d75976d10d7146981d36fc9413565ea5a73

Pith citing papers

Observation 00f45ee6-cf2c-4edd-b665-fcf19c93116a · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.884365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:d8f0011ffa97f04b7e5b2872c62873f7d4946ecfba35e641f47cf63c187b0b5c

Observation 98aa2f0d-0924-40d1-a253-9fb7348f8004 · inbound

Transferability Between Understanding and Generation in Unified Multimodal Models cites this paper.

Transferability Between Understanding and Generation in Unified Multimodal Models Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation

Reference 97

Resolution
unresolved
no resolver link, observed 2026-07-11T19:17:19.634242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:17:19.634242Z digest=sha256:0641c4573558d8a3dccc1a7a91086febc7555c55829f08c4533249e8fdcbc5da