Pith. sign in

Paper Citation Record · LEDGER

InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2305.05662.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.05662 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:13.117762Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:08:55.658612Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dd337dbf-d56d-4135-a28d-da0da5248840 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.598826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:b9ce7f689e66a01d9080ac067246bc22431f588dd8142b14c9920c6af96b9e2b

Observation feddb060-ac08-41fd-85b3-b92b79efb313 · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.516674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:fa0ae1046888c65f7ad83a310dda67405a5ab58a03232d7be428ddc163936af7

Observation b34065cb-728f-40bc-9c7d-ea7dd6ff454b · inbound

The Rise and Potential of Large Language Model Based Agents: A Survey cites this paper.

The Rise and Potential of Large Language Model Based Agents: A Survey InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 299

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:47:54.533970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T10:47:44.152066Z digest=sha256:360859d39fdeee793d3a0653df9513ac43be56c70eaf22ff1845015c8b232c64

Observation ee97f41b-1822-4edc-85b2-da5fcc40a46f · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.030518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:9b5a580526bf613f156726e254f01749d6dc93b4d5f5565e23cdcde2a0f0659b

Observation 91f3c253-72ec-41f9-855e-271bfc0e9c18 · inbound

Data-Centric Foundation Models in Computational Healthcare: A Survey cites this paper.

Data-Centric Foundation Models in Computational Healthcare: A Survey InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 181

Resolution
verified exact
arxiv_id, observed 2026-05-24T04:13:52.961508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T04:13:05.328492Z digest=sha256:2f1e5ac86cad0be20bd588f845f8191f619f93d43a30c419843cf3df212c053c

Observation 7336206c-68d3-48c4-b7a7-934d169d5c05 · inbound

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models cites this paper.

MoE-LLaVA: Mixture of Experts for Large Vision-Language Models InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:33:30.353423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T02:33:30.143907Z digest=sha256:deff3f5d0dc8af649e5a06ff1742bf5b1419ba47ce294ec6493a46f6eb9aac92

Observation 7d452ab8-3858-4dea-a383-87f323cf3352 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.117021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:36362b5d68fcc20c952f03695829f9f7157496ad768c177f88e2c37727d77a31

Observation 6fd60986-4138-4328-9308-0a4cbd49ab73 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.302720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:bf2a9740fd17b9687e92e6d2d7b5458ca39d130de57e7450ce9581d70010f090

Observation 9e31acdc-ec99-411c-8518-ebe57bced002 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.117762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.117762Z digest=sha256:aa3279292115820662818eb5396f594630f4a97c1d49c28d7cce0b50ac9928f3

Observation d21d2863-bf31-45e2-87fa-a91a5f5dbc6a · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:51.207678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:51.207678Z digest=sha256:a390079ac76e0e83dfa963b65d12257dc87f8b8f45b761ca7aeb00338d38d0e0

Observation 4b577851-b868-4a87-a58d-e891afe63e52 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.773169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.773169Z digest=sha256:d9bf8481e422a2e436eabc37b48f2e4800a846916e3a41975bc88576a9cd7d91

Observation 3c10a5ff-4100-4602-8d76-33c8e97745b3 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.802194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.802194Z digest=sha256:0efba1d6054f54d5cb9c2a5f6b0fb40eb36bbf95e2bd416410266d56c1040b40

Observation 2e4ab4fc-1b0e-4b94-b490-7cca333b0fa0 · inbound

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning cites this paper.

GM-PRM: A Generative Multimodal Process Reward Model for Multimodal Mathematical Reasoning InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T00:59:52.580855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:59:52.580855Z digest=sha256:32b973ccc166b1374eecf66f8551b12354241ef602da46cf52eb3bc9b3451d21

Observation 2f6371ee-f53a-4db8-9243-894cc83ee150 · inbound

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning cites this paper.

CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:33:17.105582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:28:30.864143Z digest=sha256:f245cc03b06942a0b3e7a3b8d580d522c616284ff1691d5565bc81352d5c1d30

Observation 64140c31-6391-473c-8d4b-e3a21d9a9c6d · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.122437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T09:20:54.635375Z digest=sha256:55ece87d0b4cb984af7b18114bd294dcacee569426bb636a5b5718a40e385aa4

Observation 6cac5e5d-667c-4c06-a6b4-08b10b62f271 · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:01.893221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:01.893221Z digest=sha256:e652cc1b117c2c09566799779edb971829ea2397c64e3c8ad5dda848df97fd58

Observation 1529a806-b412-4f95-b41e-e8133415e65a · inbound

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation cites this paper.

AudioX-Turbo: A Unified Framework for Efficient Anything-to-Audio Generation InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:18.704967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T08:04:48.283908Z digest=sha256:b4b2187585b3c1757726a598a10f0f0ca4239bf6b7a270a8429aa0d506af93c4

Observation 2724805f-f09a-4d6c-bea4-08dd5cc7b783 · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 102

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:08:55.660752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:01c28dcace8fd0ea17925181ca2b21eea288e3713bc298dd539865b5ec938f03