Pith. sign in

Paper Citation Record · LEDGER

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation

As of 18 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2507.08513.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08513 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:24:55.227295Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22aedcb8-1c69-4983-b2da-fb13d0a3b8c5 · outbound

This paper cites Claude v3.0, 2024.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Claude v3.0, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.970840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:26.349606Z digest=sha256:934879961b17dd394d14fd1dd69561dd35d812169a43714cd9807a16ff1ff376

Observation 23ebc90b-877b-4010-9729-4f7bafea106f · outbound

This paper cites ShapeNet: An Information-Rich 3D Model Repository.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation ShapeNet: An Information-Rich 3D Model Repository

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.443277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.443277Z digest=sha256:9ec4732a0168fd57a652a9bd6030fe96d387d1e7733f5bbfcde1c6a81af1b66f

Observation ee36dacf-f13c-487a-9253-a849d2a1a051 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.519860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.519860Z digest=sha256:5a182d2d8dbe584be985d2c1d3bda9f24dc00584bded39e614572a43a62f9438

Observation c74c8ca2-29e0-459b-8ba7-1b55b8ab2e6c · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.584334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.584334Z digest=sha256:1469c11ce8ec6c26b08a98cc74e0ae44ef1334e9c8e1483a88e115f58307e887

Observation d46bf874-ec93-4731-aaae-f46aba4d136a · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.359030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.359030Z digest=sha256:5e015051428b452040385f21af3c67e3a30aa55cd2f05ec9eeb519736654e719

Observation a214032c-b48a-448e-b668-586e250128f4 · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.450996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.450996Z digest=sha256:533bf1f0aa811c94fe266a570937992b98e4017a3a9706257a7d06ea1d62a9fb

Observation e78d2225-ca41-4d40-9ac2-964e0bcdc054 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.501932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.501932Z digest=sha256:5eb73cb3ac06ca6397b9472a191e82136f0e35b66e731d8552a2cf126fc0c3e9

Observation c0a5888e-9cac-4437-b9ef-a5f1d2f6639a · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.640240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.640240Z digest=sha256:e1c0bd2fe2a8ea618ea3459ed1b9e7b71e10826261eb649386f5edd418c5eeaa

Observation d840d744-f67b-413a-a7ef-e6193407eb5d · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Objaverse: A universe of annotated 3d objects

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.904748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:50.721353Z digest=sha256:eca5ddfd15986fd42e31a6464b6a9d5f19893d19c7f712598ac7f08638f25e8c

Observation 9ace1d4d-271b-4046-a81c-b2968a3a5186 · outbound

This paper cites Objaverse-xl: A universe of 10m+ 3d objects.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Objaverse-xl: A universe of 10m+ 3d objects

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.887292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:50.780968Z digest=sha256:0c18c07091d9712c17f4890d60408da81c9385b8fcb024636a5b8c4c80ee269f

Observation 67c8ed4e-ed69-437f-8081-08a895706f20 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.875281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.875281Z digest=sha256:82a1372f1a61cc54fba17a764ac8e5045796fdb64edc8f17c9888fa536831d99

Observation e86d4e51-015f-42ad-bb9f-5d43029e79b4 · outbound

This paper cites What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.966995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.966995Z digest=sha256:0d34e197e43048cb366f70b079fdd005d9b402fec80d7cb052ba997fdbfbacdd

Observation 336ad5c2-4776-4d06-ae92-5c7c304cd5a7 · outbound

This paper cites Training on Synthetic Data Beats Real Data in Multimodal Relation Extraction.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Training on Synthetic Data Beats Real Data in Multimodal Relation Extraction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:24:55.779515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:51.055242Z digest=sha256:9b1a84b54aaa528eb0dcbc94b25795c3c00e81f6081a255fb76c76e19317c181

Observation 4e757f10-5bb8-4a03-bdcd-822a6fc97c6c · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.096889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.096889Z digest=sha256:8535da546e1aecacf2d3204015f612d63c9127310a9ee5ca04b97d74304f8901

Observation c730f22a-0a89-4e9e-8226-a4143ac99b18 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.868412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:51.180173Z digest=sha256:039a903954fbe15facc0a73c40a8faedccb40e81d3773f666d228274023b972c

Observation 5dd674ed-5472-40dd-b3fa-5fe03143854e · outbound

This paper cites Kubric: A scalable dataset generator.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Kubric: A scalable dataset generator

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.852158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:51.252372Z digest=sha256:ec05188b148983e2f71bf1fa9dcd6c3b1adb632b59f63c10da3c176627a84085

Observation 9356ccb3-194f-4f3a-b89f-5c7733380849 · outbound

This paper cites Regiongpt: Towards region understanding vision lan- guage model.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Regiongpt: Towards region understanding vision lan- guage model

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.832620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:51.328234Z digest=sha256:09838168c0b779422601dfbb594b654c6226ee3dedcbb421772a95bfdd4463a6

Observation 953075ef-eae8-48c0-b500-13bd6b9b84be · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Lvis: A dataset for large vocabulary instance segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.814743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:51.398073Z digest=sha256:248060637a4b26f1391c282d62304589c9c23d9044a510b7b12d0c3a43ce9809

Observation 4cacf1a3-b313-423c-a996-4dda7392e2a0 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Vizwiz grand challenge: Answering visual questions from blind people

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.445774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.445774Z digest=sha256:72faa723705fc3c6fb1abfe83c26fd6d13657d9b3068f891049024b57bc0942e

Observation 9577be88-486f-4125-b157-813f5b0e6346 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.489571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.489571Z digest=sha256:179928a5733d1e3b092688ff9f88456b655a942b05e953d1a55bfffce2f7ac6b

Observation 2bf75061-9736-4928-9110-4fe24c39642c · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.541947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.541947Z digest=sha256:a1695b9eaca13740d513bc99c4d26c88bdf93892e38b32c762d76d2ec67fe3cb

Observation 75415d62-cf08-44fe-a973-1581acc6461e · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.592721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.592721Z digest=sha256:d51671fa92d116212e6837953c54bbe28abf0759570df3ef5820a8c93dab1ace

Observation 04d06e66-6699-4701-9838-9557421ee249 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Denoising dif- fusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.640713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.640713Z digest=sha256:0130d85e7579961ce2a46348e098d671f6db02d509b60f1f4b7c18748abf496e

Observation 31ab4b62-2a33-42f6-898d-1e9c7a8e186b · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.756961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:51.712219Z digest=sha256:9453fe927c61f74854045e46323ff6918c9ef859cce72239ecb6aea5aa273f70

Observation 089ba357-1387-443c-8879-f522d36d1506 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.739583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:51.765806Z digest=sha256:e6ce163b685e412e74cf2da33bfab1bfae597ef8d88b68cfb2acf08eb54d5a89

Observation 095a08c1-e4e2-4643-8893-59ab52b7ae9d · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.825630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.825630Z digest=sha256:07f798a1dc301741c942db803880053cd03abd3d5d3b4d87ef5e2b7b29d8e47a

Observation 8614b7fd-4116-423d-9fa5-881825ac47d2 · outbound

This paper cites Learning Action and Reasoning-Centric Image Editing from Videos and Simulations.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.875150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.875150Z digest=sha256:34b3fc7447771c92fa2e40bfc1bf9469049a0a8441acb6697a86cc1d373d6d69

Observation b730cbda-5bef-43f6-8798-4771a4d6016f · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.915011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.915011Z digest=sha256:2fc39439f2a7132ad244c81fb7d197dab63becae80ccc62e088becb5d961ecf4

Observation 0b07225c-fcd3-491d-a3f7-858971ef0997 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.958093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.958093Z digest=sha256:a9b37fa5be07c58447a392d5934aa35d722768b1bdbdf4576e3c27332512df18

Observation a10dd858-9346-4875-933a-7566f1fd5aac · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.066023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.066023Z digest=sha256:4b3fd3e6260047512fef92b6ed69922912bd9a7d2cdda7a4531c969ad0a28872

Observation e92f5ebe-1327-4337-bc68-0d2a152b7e54 · outbound

This paper cites StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.156204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.156204Z digest=sha256:525b505b6ffd1a7b4f612493560663be18922a0bc9ac8a63b1f3edbc06d0eb76

Observation b9b3ccb0-a351-44e7-9cea-820c995e8ae4 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Vila: On pre-training for vi- sual language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.687602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:52.337217Z digest=sha256:76d5693dca17a1c333082c508b0053849902621161d2f37d4ca22d5d04459def

Observation f35e224e-dd1e-4a27-b02c-93a3aee0d20e · outbound

This paper cites Microsoft coco: Common objects in context.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Microsoft coco: Common objects in context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.445832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.445832Z digest=sha256:620cc374707784bf53c7ce6eb86da7f89ea5bfee431f3d31d10ff6e41758f233

Observation 814c3bba-122f-4c28-9e23-cae8c89f935f · outbound

This paper cites Visual spa- tial reasoning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual spa- tial reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.659790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:52.521478Z digest=sha256:581b4a887b677cb213e0cbedf0de43aad75bfaf2077de9e527ec95db4c42b0a0

Observation 42b1530c-475b-427e-a737-908dacbf949c · outbound

This paper cites Improved baselines with visual instruction tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Improved baselines with visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.642709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:52.581229Z digest=sha256:350a66b726cb10586d5ddfae7368fa9efb3962f55fc1b5c65dc871d7caa742ef

Observation 48815cf8-cfef-4993-a696-fbfce8c671f3 · outbound

This paper cites Visual instruction tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.651502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.651502Z digest=sha256:3e333d95c655b64ebaa858fd033c3a230f5457358ca7f615f5779e44be590814

Observation 4d370933-8331-4859-a900-570bb89c1077 · outbound

This paper cites Clevr-ref+: Diagnosing visual reasoning with referring ex- pressions.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Clevr-ref+: Diagnosing visual reasoning with referring ex- pressions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.617137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:52.714850Z digest=sha256:a4d4b6492b2a88aff2dd1b2ad3bbd74041f82b4660b2f1c8e11d764b80dcbfec

Observation 0607d074-6d96-4b08-9d9c-b9ba782c61a4 · outbound

This paper cites SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.771144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.771144Z digest=sha256:a7d400cdfed221ebc59d8537606bdd655758c8690e2c4c140a6c9f759c3f86da

Observation ed303b27-b88d-462f-a43b-73855064240f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.830815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.830815Z digest=sha256:2d9e87be3a891a120d58c6244db59242599e0ad28e49b8cd359ce50d7ab909c0

Observation ab1081b8-88ad-4e2c-802a-313c49b41b59 · outbound

This paper cites Generating images with 3d annotations using diffusion models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Generating images with 3d annotations using diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.588673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:52.873041Z digest=sha256:444a7c5f151374c23bd7b0ec13f64edccb8a497cf7f51a978ce878fc7b714ba5

Observation a6cd09cf-de71-4d48-9097-e578dad25472 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.570935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:52.914267Z digest=sha256:b0fbdc013dfb86eb4870ebce3ceb7b18dcf8965de3b9a7f4a359e00868a29791

Observation 4dd85980-213a-4cb5-871d-50e3964e5b8b · outbound

This paper cites Llama 3.2 vision, 2024.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Llama 3.2 vision, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.552750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:52.948441Z digest=sha256:9325397882c10e295f33487b2e615f64a3eab19f4e16d20da9df4afced71fb9e

Observation c94a93ae-c02f-4e07-a0b7-2b79fc18a6e4 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Ocr-vqa: Visual question answering by reading text in images

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.982475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.982475Z digest=sha256:ba4f76602b7579d16b43a1ab37e9c399201f1f2c0074302060ffcd6697fb583c

Observation 1ae0ea9c-8307-4059-b057-82dd1cdba90a · outbound

This paper cites GPT-4 Technical Report.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.028696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.028696Z digest=sha256:68e92a584aeee1eb8fd4d53099ef3afe692dd1bb866f88226f55bf81cd6a53c4

Observation 1661cc17-f57a-44b8-b047-43babe783fc7 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.060999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.060999Z digest=sha256:89ce49d91ad083667baa257f19ee3f55dd4a6ca02e316910680aa9ca7d575545

Observation ca9c2755-f9d7-435a-8a1f-717b50699f4e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.119915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.119915Z digest=sha256:7cdfcdb5b5a7d2be56e39e51b7cca253d3f6d3c3580f30ce7b80e3f6c7217e0e

Observation 4052a144-3fa7-4a01-a69c-8fa3c4b4411a · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation High-resolution image synthesis with latent diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.524957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:53.198135Z digest=sha256:02c72e78f2e9133fb0152b8a254b9fa36c7beb9a65dff9594a827223458b50c2

Observation dd505345-76be-404b-9b10-275b36819337 · outbound

This paper cites Imagenet large scale visual recognition challenge.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagenet large scale visual recognition challenge

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.303467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.303467Z digest=sha256:79c72d89a92a8a6efa57d4b7408b8d945ad030a869ff6b41689a774d5754d0ec

Observation 9875cc28-2f98-468c-a8d9-b47623ed61cb · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.390031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.390031Z digest=sha256:d6deefad3fd17924da24e441f3ead9e5f1724f4a9f4378fe37eed98ea9136172

Observation d002c5df-03fe-4763-a28f-23eee2621d6b · outbound

This paper cites Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:24:55.438877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:53.424779Z digest=sha256:3625da2ac220f92cf5f5058a9dfbb18554d3aa81e3ad06f44df4ec3d2de2d337

Observation d4797100-dcb6-4087-bc71-4461a0599388 · outbound

This paper cites Towards vqa models that can read.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Towards vqa models that can read

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.489465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.489465Z digest=sha256:e494370bad1d5a05fe15ffbfb7ddfe8ca27f22b334f3162524eccf1ae799300d

Observation b8aa1ea9-f8e2-4318-9320-7e0eb3c59cde · outbound

This paper cites Denoising Diffusion Implicit Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Denoising Diffusion Implicit Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.538670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.538670Z digest=sha256:8a73d907ad46c4c19310f99d8b12d134061436224b7de01535894ee9728b1f5b

Observation c9d583e9-9e00-4bba-96a4-659e5838cd43 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.577193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.577193Z digest=sha256:7b74bde2e7658902ee0aff94c75b3a361f6ad091aae0c6ed164870f65061d2bf

Observation 4b332fb1-d509-4478-b476-b450b80f5090 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.451740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:53.637886Z digest=sha256:33a192d9e32b5cb4aca8fb82033bb758d1261cc78f0aa9da53e5c4a3fc9f1774

Observation 57934c97-273e-42da-9534-3393f50de742 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaMA: Open and Efficient Foundation Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.692677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.692677Z digest=sha256:15ec7bc8d2056c6e590b5e36b66ff2461c8da25d0420756da0d5c009ea382975

Observation 0e37eb4b-4166-45b0-b6df-e75df5fd21a9 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.783559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.783559Z digest=sha256:0c86e659d12f797a70459680fe90ce7f24c96605f920ea64e4c3849bed25709d

Observation 8423d23e-82ee-47e6-9527-49274baf82e9 · outbound

This paper cites Imagen editor and editbench: Advancing and evaluating text-guided im- age inpainting.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagen editor and editbench: Advancing and evaluating text-guided im- age inpainting

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.419298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:53.842559Z digest=sha256:3a34478f0f2a9a2e22a3782d19d608db38f5c4e9ec3e2260998aeb2c90610766

Observation fd5b5f6f-9e11-4231-831c-620c733cfe23 · outbound

This paper cites Mebow: Monocular estima- tion of body orientation in the wild.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Mebow: Monocular estima- tion of body orientation in the wild

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.390347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:53.966799Z digest=sha256:2779ee6706f25ee50861177c8fd9b6655431341c649b79e3f873534d5bc0a4f5

Observation 03e92d3b-bf6b-4034-992f-3cc3dd659b52 · outbound

This paper cites Beyond pascal: A benchmark for 3d object detection in the wild.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Beyond pascal: A benchmark for 3d object detection in the wild

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.090317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:54.042393Z digest=sha256:6434046799b18d0d2a764bb1cb34083e5de43973b8f536a696bee4e29bd14fbe

Observation 117e2623-f045-42d3-91da-bc5c6bedd6b6 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.505282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:54.129706Z digest=sha256:e0cbaf05c3f862034451ce1f2ffadcb82359e716342fb1ee93cf42969555f884

Observation 1a3ab2d6-e4ce-4703-be7f-c99d03379062 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.208155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.208155Z digest=sha256:0c32bbf8e68e31869014f09101ab0fada60d50315e9c19045a2268e09d6cc187

Observation 6a79d9d7-e38a-465d-85c8-7c341d6176ed · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.285272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.285272Z digest=sha256:de54f140cb0a03129b07814dec2b74476963d1034ac7cef346477f4dac2a8948

Observation a1ae6438-1357-4865-a2aa-6fea4b191f89 · outbound

This paper cites Modeling context in referring expres- sions.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Modeling context in referring expres- sions

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.281558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:54.353308Z digest=sha256:89cc19b4dfae3e5105f602f9d21e9f6a7464cb615e9e3b37c83c4becced2230b

Observation 720d82eb-b1da-4161-9cf6-1fba16e76b26 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.449939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.449939Z digest=sha256:566cf560b8699cd7ea4e830c04359b2677a1984340730aab65a5e69fa9bd5302

Observation 63f700ad-48dc-4077-98d8-270a195668c4 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Magicbrush: A manually annotated dataset for instruction- guided image editing

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.170158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:54.528019Z digest=sha256:f21a938faf90f124da3a59a35873de0ec550375aa5a5d5ecc1c944d8a9d368b1

Observation 8f26794e-2a13-4c39-9435-d62985f49828 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Adding conditional control to text-to-image diffusion models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.037280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:54.606014Z digest=sha256:d00a62ff539c1a1de5612906b23494db6678aa2f9ea775c4c6f72373eb50ca8e

Observation 40b253b0-067d-4f80-b54f-90b17167e97a · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.680277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.680277Z digest=sha256:eae9407ba5f407ae4390756d860bc52d80917eb556f6b07cabb2f1652cfa6a1c

Observation c59fd0ac-0659-49c7-88c2-f56a3efe1277 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.749240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.749240Z digest=sha256:d0e77ae9462cefaab1162a8c61b0c0f6aa904aa90843c84a710ea0e7e3b6e8f9

Observation f97af725-1353-4477-a6e1-299e62af70b7 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Semantic under- standing of scenes through the ade20k dataset

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.904192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:54.790724Z digest=sha256:779a6e000e957185276fb6624d32ae1d8c83da31bd4868d3f5ddae7846d404ab

Observation 7931a257-a464-4c0c-b3df-4fbbc23f9357 · outbound

This paper cites Corresponding section in main paper is Sec.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Corresponding section in main paper is Sec

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.816980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:54.830972Z digest=sha256:5c48b1359b48870f51a6b71f52956b080158ce68a2a0dfcfa6bb547b94254f25

Observation 4260af33-793b-4be7-b525-314b47663afb · outbound

This paper cites front" also means.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation front" also means

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.728795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:54.869213Z digest=sha256:9b582cba7da568ded643d1d07819f8e9df95ddfe91be54e4c7c5fccc0d42fd0a

Observation 662fd4c4-96c5-4aa1-a11b-8950e0b57382 · outbound

This paper cites SYSTEM_PROMPT_FOR_GRADING_MLLM_RESPONSE =’You are a helpful and precise assistant for checking the quality of the answer.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SYSTEM_PROMPT_FOR_GRADING_MLLM_RESPONSE =’You are a helpful and precise assistant for checking the quality of the answer

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.651563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:54.929420Z digest=sha256:93850692e5540939731b67f867fcfa0c5d55d163f51d559c00b6887265f494f1

Observation 11274d10-18db-4291-b450-0af94b271366 · outbound

This paper cites 6, we show an example of using ImageReward [60] for the dataset curation by evaluating the alignment be- tween generated image and text prompts.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 6, we show an example of using ImageReward [60] for the dataset curation by evaluating the alignment be- tween generated image and text prompts

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.549201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:54.992499Z digest=sha256:c46589eed31fa6b12bc90e95a62d94e285f27f1306b6a991bed50dbbffb09288

Observation a833ed89-9c6d-4823-a8ec-80624261cff8 · outbound

This paper cites 3, we perform quantitative comparisons on general image visual quality between different DM backbones: SD V1.5 and SDXL.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 3, we perform quantitative comparisons on general image visual quality between different DM backbones: SD V1.5 and SDXL

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.468624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:55.032116Z digest=sha256:2217cc918594d6c181d6edd27c44010f039f92175eb24e8b8836d4aaa07a8947

Observation d8e97d91-b1ca-4777-87d7-625d4fce5795 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:56.391242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:55.074253Z digest=sha256:fcf64e91659e941dec41a5ce7bb75bbb3373dca65b3c1072f6625032db857103

Observation 6fd6cb42-cbf1-427b-a8aa-f8f26501081d · outbound

This paper cites 9 shows the UI page of our user study.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 9 shows the UI page of our user study

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.320779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:55.116586Z digest=sha256:00c73ade215456137a57df2f3b53bcde4285f5b66b70f1c8989310b783d49d59

Observation f80025b9-0df0-4d73-a614-011c00d5989c · outbound

This paper cites 10 to show more qualitative comparisons between fine- tuned LLaV A model to commercial SOTAs.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 10 to show more qualitative comparisons between fine- tuned LLaV A model to commercial SOTAs

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.259557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:55.153359Z digest=sha256:0a9bc18b5d092ac1ecc11ef968c342ae7b2777b559767817f271c004695f34fe

Observation ce8877c5-f5bf-41fc-b1a7-c7c842b22480 · outbound

This paper cites 11, we shows more examples of diversity on object categories, camera-object relation, and background con- texts.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 11, we shows more examples of diversity on object categories, camera-object relation, and background con- texts

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.181755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:55.177843Z digest=sha256:6b67713c1f51b5b32705981b53a7aba4997d552a0dc40fd2bc5446108a1f539d

Observation 7b09579b-8e9e-4749-9a60-76040f022433 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:56.077268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:55.205611Z digest=sha256:ebb8d0ad2392095725594a5d29ebe31c2022d634c7cfe2ca0663920991e748ba

Observation ad2ed1f9-ff4d-4f55-9408-d2fbb7629d6e · outbound

This paper cites an unresolved cited work.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:55.986032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:24:55.227295Z digest=sha256:a983b20af3492f70fece55b408b0acf43e09fca66bec4ec3c1ddeeae8c1482e4

Pith citing papers

No inbound Pith citation observations are available.