Pith. sign in

Paper Citation Record · LEDGER

ChatImage: Navigating Long-Form LLM Answers through Interactive Images

As of 10 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.05290.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.05290 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-07T19:32:20.095114Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact10
  • verified fuzzy28
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4ec586dc-a98e-407e-8b02-0a882ccabf28 · outbound

This paper cites The Claude model family.Anthropic Technical Report, 2024.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images The Claude model family.Anthropic Technical Report, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.779307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:84906606c14a8b34e880854d3fd4870d24d4187126d8007a94fcff638c329b03

Observation b0018f56-a31a-4978-bc4e-5336c9377568 · outbound

This paper cites VQA: Visual question answering.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images VQA: Visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.732264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:ddde530b7ff4aa17571eff78af9cc36e393c5f8408cad46f079958b105de2093

Observation 74c70bbd-67b6-4a25-8811-8b8f660bb9f1 · outbound

This paper cites Reranking individuals: The effect of fair classification within-groups.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Reranking individuals: The effect of fair classification within-groups

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T19:34:06.399306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:ea0738c8edf24a22eda3bf2f7c82f1e16067cdcfadd927af0d9622211f1ebc77

Observation 2844780f-372a-4a63-8a13-5ce284880f8d · outbound

This paper cites InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.748766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:f304f4d5fdba49c6440c97758c70b12c4a5873930aa4dc1a016b0296f02ca809

Observation d6d0c24e-b42a-429b-a272-b9d3487a24dd · outbound

This paper cites YOLO-World: Real-time open-vocabulary object detection.IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images YOLO-World: Real-time open-vocabulary object detection.IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.716095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:410eda861a3a221afe5560576cee0b2b28374aa608d30aef18aaabeea02c6b05

Observation 0f126077-7b68-48e2-b686-125d20680f78 · outbound

This paper cites Concatenating Binomial Codes with the Planar Code.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Concatenating Binomial Codes with the Planar Code

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T19:34:06.380852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:e50e50a81b9ea5df3205da5b1d585a9d5bd26289f1a06afa6f8e48e1e87ead70

Observation 98bbaa77-b24a-4e55-a618-50cebcf4dcec · outbound

This paper cites TransVG: Visual grounding with transformers.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images TransVG: Visual grounding with transformers

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.810593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:9c671fa2c80f85078a2071747dc4265f643aed67093a64c803e4ee4f179b6c5a

Observation 00d0de1c-f87d-4852-80bc-71594c2d534f · outbound

This paper cites LayoutGPT: Compositional visual planning and generation with large language models.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images LayoutGPT: Compositional visual planning and generation with large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.818944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:ef6d3da45a83aae9f3df05d8cd2f606c071cd71ac4108b3a88812688439b9ae3

Observation 0d620a3b-6552-4ca1-9cbb-90269bce75c7 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-07T19:34:06.371018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:ff71d54e81542c6f225664df525e10621941795eb4c97547845f4fa1c0cbd315

Observation a32ba699-9e1a-4bd8-9cbc-a51f9eeb0929 · outbound

This paper cites Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-07T19:34:06.385691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:ce8e8cecc09855ea1d1b35ae71f5ec3ac5c8b0761e1b86b1dc7b660e4f1c9ae9

Observation 914006fb-a983-4a99-9c14-e51baaa1631e · outbound

This paper cites Interactive poster visualization with PhD Online.IEEE Computer Graphics and Applications, 2004.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Interactive poster visualization with PhD Online.IEEE Computer Graphics and Applications, 2004

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.770093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:31390ce712f7fe3c620792ecefd3b362c9c8af35a090e98ce92e70e8dd7e3943

Observation a26503f0-6a60-4938-9956-752cc374522d · outbound

This paper cites Cryptography: Classical versus Post-Quantum.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Cryptography: Classical versus Post-Quantum

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T19:34:06.365987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:4107f7689f80cb1563c53e76aac845624464d8011e5d92b9750a8eeef91aa486

Observation d11de996-ca9b-49b0-85ce-2a462c0a5d13 · outbound

This paper cites SPIKE: An Adaptive Dual Controller Framework for Cost-Efficient Long-Horizon Game Agents.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images SPIKE: An Adaptive Dual Controller Framework for Cost-Efficient Long-Horizon Game Agents

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-07T19:34:06.361002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:3d1ef079a4073f932fc6031c33a6d8b98e3ab607629cc2d0919891284b78e8da

Observation 657bf2bf-32a4-4e30-b94a-da0b94e6c0ba · outbound

This paper cites MDETR: Modulated detection for end-to-end multi-modal understanding.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images MDETR: Modulated detection for end-to-end multi-modal understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.727764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:7cacfbdf41da78ca7b3d30bdc0aa9c55ad5a3c4d467dbaadb2638c9d9f30b5bd

Observation d351c8ea-411e-46a5-b25b-c23a42be83f6 · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.740141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:43cb2e6cd7c9a033c2c7ee0b924d6733171d9e383b433389b80be8cdc1c61d66

Observation 7c916d13-74a2-4c07-97cc-f775f64bb773 · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive NLP tasks.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Retrieval-augmented generation for knowledge-intensive NLP tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.753043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:4eb61cbc214788e43b493e1cc45098f1f90c99f4ed4459bc773819659fd67710

Observation 7deb228a-bc66-448d-b927-6b8f07e9a303 · outbound

This paper cites GLIGEN: Open-set grounded text-to-image generation.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images GLIGEN: Open-set grounded text-to-image generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.783486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:67c385a55ca04b5c5712ff88fcce58bfd331f425a554d497a15fa6f2c3bb5215

Observation 635065f5-2d11-4490-828c-f8b63c9d8d19 · outbound

This paper cites Visual instruction tuning.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Visual instruction tuning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.765998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:ee56bef3adf2e37d3b2eb0d1623be47b5685cf04d484a85e401d9d6ddd94ff3f

Observation db309a84-39d5-4e0c-a421-064e9c7db105 · outbound

This paper cites Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Grounding DINO: Marrying DINO with grounded pre-training for open-set object detection

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.761850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:67c5690f3cb2a14066dc89f9e177b656b9c866591df17a86420634bbc45b9ab4

Observation 3e93ea9d-1650-4983-a6cb-50cbef1d3f90 · outbound

This paper cites Counting Divisors in the Outputs of a Binary Quadratic Form.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Counting Divisors in the Outputs of a Binary Quadratic Form

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-07T19:34:06.390078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:715400e830c18494d5eefb363f5ef07aa1e571552abae0ce1ebaaf4639c3f548

Observation 89ac7dbf-e0ff-4565-bce4-e53692618d26 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.797214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:3a83458709af20523dc85f1875cc611fdc3271db24badfabba963fb95d94725b

Observation 84a7feb8-b4fb-45f7-aa7c-dfff2a0fc1f6 · outbound

This paper cites an unresolved cited work.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-07-07T19:34:06.719372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:f50f2c17dc5f5117f8d4d846e4659e0d50d4d3baf2b69407d12b76043cb1cc14

Observation 41115014-6cde-4e02-8c22-87a451d6c0cd · outbound

This paper cites CRC Press, 2014.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images CRC Press, 2014

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.801627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:34d181dcf2d5cb0d85abbff6961ca970b46d1eab403b4b1914595e38444112ae

Observation ad32b0a2-54c3-4115-9f4a-478a99c914f5 · outbound

This paper cites LocateAnything-3B: A visual grounding model.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images LocateAnything-3B: A visual grounding model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.792756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:28ab90f245e49783d886b890bccfccd4dd1aa8ffcc2fb145d64e962e742e8ba4

Observation 245c8dd1-b9cd-4d2f-a528-859a2fe81e26 · outbound

This paper cites Chart-to-Text: Generating textual descriptions of charts.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Chart-to-Text: Generating textual descriptions of charts

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.722953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:6185ba18ffea10e3d445ded09064c879f89c2e53c18c699d0d41d5b7ad44bcd7

Observation 1dbb3275-0f81-44e1-bc40-53e06ce96e68 · outbound

This paper cites GPT-4 Technical Report.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images GPT-4 Technical Report

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-07T19:34:06.407051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:bf8907a3150acea00d6d17d11e89c7ad67c002c1b728fa13c87e33ea2e6eb16c

Observation 0d4e3383-3bc1-4bf8-bb8d-24c8b2a42a91 · outbound

This paper cites SDXL: Improving latent diffusion models for high-resolution image synthesis.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images SDXL: Improving latent diffusion models for high-resolution image synthesis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.788107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:c7941eee78a687a470c1e54b8d54ba16e01d7c82e6c3263e47b9fae5f5c51bab

Observation bfcd80d0-2678-431f-be8f-8154e0697eac · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-07T19:34:06.376190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:6abd224e5af0ef880d9ace1053c8cb518aeb8827274b8194a2305cf439f76dd5

Observation b6deefb4-f147-4366-88b9-bb440bdabbd3 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images SAM 2: Segment Anything in Images and Videos

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-07T19:34:06.403344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:1f22817a15b839e0a78106cf4e6c5c31c6c47e1c398f2d5dc391df894314d77d

Observation 9ae6322d-6b31-45ca-a3e3-16160ecac009 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images High-resolution image synthesis with latent diffusion models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.744279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:8c272dfb5e845f882ced82d630db441885e8be9076f3c53e448a62a316311e84

Observation bffa7a5d-dd25-4cfb-9221-7c4287ff7ffb · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Photorealistic text-to-image diffusion models with deep language understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.756944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:983bacb3c52fa98001d420075fb55184d82eaebd9285c86bdcc122a446868228

Observation f012e20a-318f-4bd2-92f2-832f442f3e0b · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Toolformer: Language models can teach themselves to use tools

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.775142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:2fca147798db78db0d645429de3c107c5365e5f3bf42a5326c0bc4b7bc86bba1

Observation 004a7901-3849-47e3-a941-39559dafdccc · outbound

This paper cites ChartLlama: A Multimodal LLM for Chart Understanding and Generation.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images ChartLlama: A Multimodal LLM for Chart Understanding and Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-07T19:34:06.394815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:7b2639c743886a2fd64581630cc9e625363c1772bc73cd139248a4de4a113fe1

Observation 55d0d7a1-2263-4b36-a549-b0a9b68e33af · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Chain-of-thought prompting elicits reasoning in large language models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.711848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:48cef2d324bbf12c204c5bc7f49ac9debca7dcbb32b1f2a4b3d49518a097bb0f

Observation 2833e145-7e42-4507-9ec6-3a3cd5013d23 · outbound

This paper cites MiMo-VL: Xiaomi MiMo vision-language model.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images MiMo-VL: Xiaomi MiMo vision-language model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.827556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:4971e6b0bfd39c88ba25dfeffd555845445630d6d05e136e21729a71aa59515b

Observation ea4883c0-4c75-457e-9647-11178c97a973 · outbound

This paper cites an unresolved cited work.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-07-07T19:34:06.805860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:d87a264e00a042017967c52a4f76fe768f21ad0b7c86dd7fb224e261dcc0beea

Observation 2f96abde-6127-4493-9693-0841d3e0fc3e · outbound

This paper cites LLM-oriented token- adaptive knowledge distillation.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images LLM-oriented token- adaptive knowledge distillation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.814889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:845324743498b144faaa9404c5e57d3bc7c52330ccd48dbf83973334bd47873d

Observation d056bd35-82c6-4119-9796-6b17bf0a58f1 · outbound

This paper cites an unresolved cited work.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Unresolved cited work

Reference 38

Resolution
verified exact
doi, observed 2026-07-07T19:34:06.342685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:97711080cd11f71ee851fb4d3943e1108cee53ced109a8693b33dc563ef5e39b

Observation c075f7a1-ad10-4891-9855-286a7582edd3 · outbound

This paper cites UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-07T19:34:06.411651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:eb47d38d1807cd76aeba6cce8f6b97a26ed3e24ebe4db8dbdc51eef7d701814b

Observation 3e19e448-1012-4b29-81f5-5cae83ef75e9 · outbound

This paper cites UNINEXT: Universal instance perception as object-in-context prompting.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images UNINEXT: Universal instance perception as object-in-context prompting

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.735898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:5faa1f54e16e18f7487258b5fe7dab0583af89c38281565bd9ce648c3db6c7e5

Observation a5ec72a0-b7b9-4e1c-a691-dd2e57fae546 · outbound

This paper cites ReAct: Synergizing reasoning and acting in language models.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images ReAct: Synergizing reasoning and acting in language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.705003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:cec99655e4011e6dd6a442d7b603d1f127305ff67045fb1f1cdfeb5e4aec2016

Observation f55b13b3-dcd7-4d4a-9867-112d8cefcbba · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Adding conditional control to text-to-image diffusion models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.823384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:521c0096acbff19b0c32c2d8b469ef570b016462819d3d50786291cd832ce0eb

Observation b92a58e9-49d6-4897-8119-b87c2eaff329 · outbound

This paper cites Mindmap: A creative visual thinking tool for education.Journal of Educational Technology Systems, 2014.

ChatImage: Navigating Long-Form LLM Answers through Interactive Images Mindmap: A creative visual thinking tool for education.Journal of Educational Technology Systems, 2014

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T19:34:06.708182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-07T19:32:20.095114Z digest=sha256:c85222cfdc7f1e29c4a3afd80162ee504add33c5c1e3d1b57d49d17fb6bb01f9

Pith citing papers

No inbound Pith citation observations are available.