Pith. sign in

Paper Citation Record · LEDGER

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation

As of 9 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2507.08513.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08513 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:24:55.227295Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 22aedcb8-1c69-4983-b2da-fb13d0a3b8c5 · outbound

This paper cites Claude v3.0, 2024.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Claude v3.0, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.970840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:26.349606Z digest=sha256:2e9fdf93124ab0d27516eeea11c700f3398a78d22c25740345250449c3e4ce83

Observation 23ebc90b-877b-4010-9729-4f7bafea106f · outbound

This paper cites ShapeNet: An Information-Rich 3D Model Repository.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation ShapeNet: An Information-Rich 3D Model Repository

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.443277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.443277Z digest=sha256:9d7d93c8b1690ce01804bdba05b9bea5a7ee13712b65f0f50128ad0781e38685

Observation ee36dacf-f13c-487a-9253-a849d2a1a051 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.519860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.519860Z digest=sha256:9cd650f8db819436cb1de65a1f03e690a5eddf1e5b91693da7615da85faaebe3

Observation c74c8ca2-29e0-459b-8ba7-1b55b8ab2e6c · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.584334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.584334Z digest=sha256:8e38b94cb162263a3cf47a6917c02de1ce0dd92b2ae8e6d912d22d58537ec7a3

Observation d46bf874-ec93-4731-aaae-f46aba4d136a · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.359030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.359030Z digest=sha256:e799eea251181bfed9952aa102208609066624a9bb576b2640ccee2bc0f0c68d

Observation a214032c-b48a-448e-b668-586e250128f4 · outbound

This paper cites SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SpatialRGPT: Grounded Spatial Reasoning in Vision Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.450996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.450996Z digest=sha256:6120a97b4ba51a1240466b012a87b43952a7c380ff80b855e0f090d428f5b42d

Observation e78d2225-ca41-4d40-9ac2-964e0bcdc054 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gonzalez, Ion Stoica, and Eric P

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.501932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.501932Z digest=sha256:2fa5f4e882d862b6695a79197b32137bff81e17778a82daf58d2e7200b632b4a

Observation c0a5888e-9cac-4437-b9ef-a5f1d2f6639a · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.640240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.640240Z digest=sha256:2b73618ae626a0220a55edcf42abc467b5fed8dbf1a11d47016521757a2371bb

Observation d840d744-f67b-413a-a7ef-e6193407eb5d · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Objaverse: A universe of annotated 3d objects

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.904748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:50.721353Z digest=sha256:44c7518f91e5f1ec7b18038257cfdee6a294f9fbbd093e420312da8a03b9607f

Observation 9ace1d4d-271b-4046-a81c-b2968a3a5186 · outbound

This paper cites Objaverse-xl: A universe of 10m+ 3d objects.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Objaverse-xl: A universe of 10m+ 3d objects

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.887292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:50.780968Z digest=sha256:b45f24bc52be26e2cdc219dbad99529a747ca6e7d9ce312e2b7b266e4e87b651

Observation 67c8ed4e-ed69-437f-8081-08a895706f20 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.875281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.875281Z digest=sha256:db668d55d4fad09d545001189d3f7aec1d02857ac7e043ef9528d5e304a7de07

Observation e86d4e51-015f-42ad-bb9f-5d43029e79b4 · outbound

This paper cites What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:50.966995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:50.966995Z digest=sha256:3162145df89c3e1e4f2808a133a9f8efe36d4c99efa90f8b8ef22b601bcab008

Observation 336ad5c2-4776-4d06-ae92-5c7c304cd5a7 · outbound

This paper cites Training on Synthetic Data Beats Real Data in Multimodal Relation Extraction.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Training on Synthetic Data Beats Real Data in Multimodal Relation Extraction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:24:55.779515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:51.055242Z digest=sha256:440f505c04cea4816bb065900e10d36862ba4877239a5c2abb8be0caa4171806

Observation 4e757f10-5bb8-4a03-bdcd-822a6fc97c6c · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.096889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.096889Z digest=sha256:515358ad25b949fb2653f73b1235074689091eaddba546f42591de872a9d88be

Observation c730f22a-0a89-4e9e-8226-a4143ac99b18 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.868412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:51.180173Z digest=sha256:d08de011356f5f94086b15e79f83488d357cb90e62e881cf81140a5e212540aa

Observation 5dd674ed-5472-40dd-b3fa-5fe03143854e · outbound

This paper cites Kubric: A scalable dataset generator.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Kubric: A scalable dataset generator

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.852158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:51.252372Z digest=sha256:4003eab7f4f4e2f2a35d6beab793a4af162193e57363b0be0ace3b5eadadfbac

Observation 9356ccb3-194f-4f3a-b89f-5c7733380849 · outbound

This paper cites Regiongpt: Towards region understanding vision lan- guage model.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Regiongpt: Towards region understanding vision lan- guage model

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.832620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:51.328234Z digest=sha256:22f1b28275109994400dc43210f1307c4c9a3200ccce3265ff4c7d07845809df

Observation 953075ef-eae8-48c0-b500-13bd6b9b84be · outbound

This paper cites Lvis: A dataset for large vocabulary instance segmentation.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Lvis: A dataset for large vocabulary instance segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.814743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:51.398073Z digest=sha256:ef49b4f89b1306368df510e8bed56a788568003356b346d9c9c3be9cfbdd843a

Observation 4cacf1a3-b313-423c-a996-4dda7392e2a0 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Vizwiz grand challenge: Answering visual questions from blind people

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.445774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.445774Z digest=sha256:f37bd9b8b35bb3d9c5482584e80804492a92de22a4258fd9088bf26972cde8e0

Observation 9577be88-486f-4125-b157-813f5b0e6346 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.489571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.489571Z digest=sha256:b40f2358153d06ca1db4f8c53b165955f7fa0a9329c1c9b321e13d2e72ef13bb

Observation 2bf75061-9736-4928-9110-4fe24c39642c · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.541947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.541947Z digest=sha256:ebe0d6f9ddecff93a0854d967e3040733d6c17d67ee72ffd26b50675c955f1fd

Observation 75415d62-cf08-44fe-a973-1581acc6461e · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.592721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.592721Z digest=sha256:3839e96e119e1e088d09aae22aaa98d4e8a576613c53d3899faf4b7196628faa

Observation 04d06e66-6699-4701-9838-9557421ee249 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Denoising dif- fusion probabilistic models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.640713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.640713Z digest=sha256:8e034218cf25a771cb6120d8df9f60abe5207a58b7e6e863f5ad60ed9d022f74

Observation 31ab4b62-2a33-42f6-898d-1e9c7a8e186b · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.756961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:51.712219Z digest=sha256:54cc04a4b2f3760998439fe8e7cfa5491df11cabc93c3b4ef551db4865275d80

Observation 089ba357-1387-443c-8879-f522d36d1506 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.739583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:51.765806Z digest=sha256:a5598e13dea4872d8c15bef6dcc9e2705a3b01d9b141e4b8efaef9263cd78b6b

Observation 095a08c1-e4e2-4643-8893-59ab52b7ae9d · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.825630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.825630Z digest=sha256:2b16fa88910964f1ebcd635ca67da94efe01e9c22a796499da85e5ef654826ab

Observation 8614b7fd-4116-423d-9fa5-881825ac47d2 · outbound

This paper cites Learning Action and Reasoning-Centric Image Editing from Videos and Simulations.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Learning Action and Reasoning-Centric Image Editing from Videos and Simulations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.875150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.875150Z digest=sha256:b5b3859e3cb9e6285844ca0a8fca5c6b554d785647782fd6cb3a01c7c6c2d69a

Observation b730cbda-5bef-43f6-8798-4771a4d6016f · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.915011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.915011Z digest=sha256:a9ee3f68713a4ce7c575ea2389d5782f7fad38e5b326a7116a0c2ea40458864f

Observation 0b07225c-fcd3-491d-a3f7-858971ef0997 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.958093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:51.958093Z digest=sha256:8fc96009a20ad090e8b642b3c8eb9cc70972c2e01941ae0525f647bc8749455a

Observation a10dd858-9346-4875-933a-7566f1fd5aac · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.066023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.066023Z digest=sha256:ac5a31618a8cb6f62fdd269e246547acc01ef69c9b42f6668b776770784ebc25

Observation e92f5ebe-1327-4337-bc68-0d2a152b7e54 · outbound

This paper cites StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation StableLLaVA: Enhanced Visual Instruction Tuning with Synthesized Image-Dialogue Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.156204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.156204Z digest=sha256:526bb2c89012b47053ae1be523f0b81eff773ab92c56232b2cb2c43df4822acc

Observation b9b3ccb0-a351-44e7-9cea-820c995e8ae4 · outbound

This paper cites Vila: On pre-training for vi- sual language models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Vila: On pre-training for vi- sual language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.687602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:52.337217Z digest=sha256:4cf0a70e74b0b34f6d440993426c9dcb2db4b083cb15ec52e3b08a4071ba6dc2

Observation f35e224e-dd1e-4a27-b02c-93a3aee0d20e · outbound

This paper cites Microsoft coco: Common objects in context.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Microsoft coco: Common objects in context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.445832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.445832Z digest=sha256:c30eccd5c6e33983fda1eae0e3e0b65cd5c6c39132d84db0850f4cab13a943c7

Observation 814c3bba-122f-4c28-9e23-cae8c89f935f · outbound

This paper cites Visual spa- tial reasoning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual spa- tial reasoning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.659790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:52.521478Z digest=sha256:10d59ba636019f4e97820db370581bb172b04e6d5870825e0962eb9a390f751d

Observation 42b1530c-475b-427e-a737-908dacbf949c · outbound

This paper cites Improved baselines with visual instruction tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Improved baselines with visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.642709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:52.581229Z digest=sha256:dbeb7d61d504a419cb4deb76ab67677f73adcb433f3bd2fca16d6fd94744c40d

Observation 48815cf8-cfef-4993-a696-fbfce8c671f3 · outbound

This paper cites Visual instruction tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.651502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.651502Z digest=sha256:70119de9509edef1939e0fe639bb1b68c56e5c627ebf62f9b618489732ce43b8

Observation 4d370933-8331-4859-a900-570bb89c1077 · outbound

This paper cites Clevr-ref+: Diagnosing visual reasoning with referring ex- pressions.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Clevr-ref+: Diagnosing visual reasoning with referring ex- pressions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.617137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:52.714850Z digest=sha256:611d138c314aca3777811bcbc9f0eab79e14719adb9e0b1fdebd620bb8893093

Observation 0607d074-6d96-4b08-9d9c-b9ba782c61a4 · outbound

This paper cites SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SynthVLM: Towards High-Quality and Efficient Synthesis of Image-Caption Datasets for Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.771144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.771144Z digest=sha256:fde80486bcfe698c73838b120057b10c26fc3f46452b39abe1488d6c52e66d49

Observation ed303b27-b88d-462f-a43b-73855064240f · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.830815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.830815Z digest=sha256:3476a15a1ae70f1326091bf983f1f3d3d0b416eb067146cc5c9c075e5b68c59f

Observation ab1081b8-88ad-4e2c-802a-313c49b41b59 · outbound

This paper cites Generating images with 3d annotations using diffusion models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Generating images with 3d annotations using diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.588673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:52.873041Z digest=sha256:4388b0f5dd7be01391bc298b2e622fd19743d4376992ccc0586b92b838ce211e

Observation a6cd09cf-de71-4d48-9097-e578dad25472 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.570935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:52.914267Z digest=sha256:bf098418bfa96ac55829e54df3b0b3d728a2bb19a28ba8039ccd94999b7668e3

Observation 4dd85980-213a-4cb5-871d-50e3964e5b8b · outbound

This paper cites Llama 3.2 vision, 2024.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Llama 3.2 vision, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.552750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:52.948441Z digest=sha256:ca88f85492ffd2ecac4a391dcaa49a4c5cf54627d9d83493b0714989aaa6fe7b

Observation c94a93ae-c02f-4e07-a0b7-2b79fc18a6e4 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Ocr-vqa: Visual question answering by reading text in images

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:52.982475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:52.982475Z digest=sha256:ae6db0d6fd2d8103fe596cb66aa5285067b0c911c47c351f0595130c718955c8

Observation 1ae0ea9c-8307-4059-b057-82dd1cdba90a · outbound

This paper cites GPT-4 Technical Report.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation GPT-4 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.028696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.028696Z digest=sha256:7dd9d17fb4dace402950fce9b21cccdfacfa2dc47e308631db11e7d5562c1ae5

Observation 1661cc17-f57a-44b8-b047-43babe783fc7 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.060999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.060999Z digest=sha256:3e4f018e22cadb27d47fbf518ad590dc36e6c08ee9b0ef5f64ee5955efd5d98e

Observation ca9c2755-f9d7-435a-8a1f-717b50699f4e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.119915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.119915Z digest=sha256:2cdcf72b1d252a901b025faace867016117209a6caa863aa31e51f2991306e8d

Observation 4052a144-3fa7-4a01-a69c-8fa3c4b4411a · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation High-resolution image synthesis with latent diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.524957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:53.198135Z digest=sha256:557bfcfe53d98f9b3d6ba2f3501226e8ebb074925074c307f549044ce80a1615

Observation dd505345-76be-404b-9b10-275b36819337 · outbound

This paper cites Imagenet large scale visual recognition challenge.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagenet large scale visual recognition challenge

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.303467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.303467Z digest=sha256:ae2a4e16536eb82bbf05f2bf6e2e80fd37678f171aa3b178cf28a41e9f2ab9ae

Observation 9875cc28-2f98-468c-a8d9-b47623ed61cb · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.390031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.390031Z digest=sha256:b45c95598ac8a36178f8af6c5e0f490ecc72204689ee9479ff193136ab7c36d9

Observation d002c5df-03fe-4763-a28f-23eee2621d6b · outbound

This paper cites Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Synth$^2$: Boosting Visual-Language Models with Synthetic Captions and Image Embeddings

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:24:55.438877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:53.424779Z digest=sha256:6cd6c82b7ed4e83d758baaf995f366dfe291957a33ea590efc92a8663d23180b

Observation d4797100-dcb6-4087-bc71-4461a0599388 · outbound

This paper cites Towards vqa models that can read.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Towards vqa models that can read

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.489465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.489465Z digest=sha256:69296cc974f18428d6ac46e983e606cec1eaca6878e97567c94fbe8ea9a6dfea

Observation b8aa1ea9-f8e2-4318-9320-7e0eb3c59cde · outbound

This paper cites Denoising Diffusion Implicit Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Denoising Diffusion Implicit Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.538670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.538670Z digest=sha256:52745391011ee3c734087b2391daf7b890583765300e8b485b759c4a0d05aed6

Observation c9d583e9-9e00-4bba-96a4-659e5838cd43 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.577193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.577193Z digest=sha256:5753a1975af973bac9869961ed3c63f3c23c874ef5bf6b3133b3b3926b9e8193

Observation 4b332fb1-d509-4478-b476-b450b80f5090 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.451740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:53.637886Z digest=sha256:d0fdf0e3a787098d010b619f5bffec6e6847c98446245c71a747d9d4ff16432c

Observation 57934c97-273e-42da-9534-3393f50de742 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaMA: Open and Efficient Foundation Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.692677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.692677Z digest=sha256:7d0603ea4de6537eeef8dc933096dc26903dc2da6b4a370b172e24240e21e24e

Observation 0e37eb4b-4166-45b0-b6df-e75df5fd21a9 · outbound

This paper cites To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation To See is to Believe: Prompting GPT-4V for Better Visual Instruction Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:53.783559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:53.783559Z digest=sha256:eb102a78e4f3bd3f1147d3eefb41ba83d1eff561aad0ac4f2ef09006e377ab15

Observation 8423d23e-82ee-47e6-9527-49274baf82e9 · outbound

This paper cites Imagen editor and editbench: Advancing and evaluating text-guided im- age inpainting.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagen editor and editbench: Advancing and evaluating text-guided im- age inpainting

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.419298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:53.842559Z digest=sha256:23ff62c21d9b84c623e2ce12fa0f1252ff35222960c6d505f8ba6b9babad7229

Observation fd5b5f6f-9e11-4231-831c-620c733cfe23 · outbound

This paper cites Mebow: Monocular estima- tion of body orientation in the wild.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Mebow: Monocular estima- tion of body orientation in the wild

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.390347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:53.966799Z digest=sha256:ace42b18d9ecb3b821bfba976266033c321a08d4ea464454ff142b2d0036b500

Observation 03e92d3b-bf6b-4034-992f-3cc3dd659b52 · outbound

This paper cites Beyond pascal: A benchmark for 3d object detection in the wild.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Beyond pascal: A benchmark for 3d object detection in the wild

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:25:19.090317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:54.042393Z digest=sha256:1a87e8e0772a687e291aa58f42aeda2c978a4217cb9be4c88b4810cd02a1ac84

Observation 117e2623-f045-42d3-91da-bc5c6bedd6b6 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.505282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:54.129706Z digest=sha256:0ac5a4b1853ac625ee571808eea14eec426055a48bf2cd80a35346d7f98c1792

Observation 1a3ab2d6-e4ce-4703-be7f-c99d03379062 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.208155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.208155Z digest=sha256:8fd8dd89ae835159be72c6bf0bc4cb7be403ba4672449c9fea7f102b7fc3b837

Observation 6a79d9d7-e38a-465d-85c8-7c341d6176ed · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.285272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.285272Z digest=sha256:90a50e2bb07147801047294fa1e6f2125af5c1acd54495df3e15138cb7d00009

Observation a1ae6438-1357-4865-a2aa-6fea4b191f89 · outbound

This paper cites Modeling context in referring expres- sions.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Modeling context in referring expres- sions

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.281558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:54.353308Z digest=sha256:6d69b220543302f920e678048a0e8aa4414f133b290e98cf6031e0d34aeb672b

Observation 720d82eb-b1da-4161-9cf6-1fba16e76b26 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.449939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.449939Z digest=sha256:4c37e2936d631d6c363501d87e8a0f0d451fb9e29824a970fe73fee9a0e2fa2d

Observation 63f700ad-48dc-4077-98d8-270a195668c4 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Magicbrush: A manually annotated dataset for instruction- guided image editing

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.170158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:54.528019Z digest=sha256:56ebc23539648216b39ab3243221c0ff0a1e3772fea59b759a72802b67b71ea0

Observation 8f26794e-2a13-4c39-9435-d62985f49828 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Adding conditional control to text-to-image diffusion models

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:57.037280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:54.606014Z digest=sha256:9136d76fb16a42bed23463ed6ae066d2066f8761ac258f83693ae08ec8457f81

Observation 40b253b0-067d-4f80-b54f-90b17167e97a · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.680277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.680277Z digest=sha256:c3f5052676b06ff8a2afa3ae683f5eb9d200cc743a8d09df9886cb5b9321ed3f

Observation c59fd0ac-0659-49c7-88c2-f56a3efe1277 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:54.749240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:54.749240Z digest=sha256:abb119b4cbfb9cf29d09c8d4c75cfc8cae3fe889e7ca7e7ce27cb0d6e869caf3

Observation f97af725-1353-4477-a6e1-299e62af70b7 · outbound

This paper cites Semantic under- standing of scenes through the ade20k dataset.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Semantic under- standing of scenes through the ade20k dataset

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.904192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:54.790724Z digest=sha256:e9a043d8cedc98eeb099cac4f5d4902b854804e1251afb9735230ca6645b72dd

Observation 7931a257-a464-4c0c-b3df-4fbbc23f9357 · outbound

This paper cites Corresponding section in main paper is Sec.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Corresponding section in main paper is Sec

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.816980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:54.830972Z digest=sha256:00e42b254532fc9c03c0365504ba83a5cd79d26d53fd85c9cc64f7e52449de93

Observation 4260af33-793b-4be7-b525-314b47663afb · outbound

This paper cites front" also means.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation front" also means

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.728795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:54.869213Z digest=sha256:48ac30a29c4d76347fee466e728c4bb24df8cbe2d44e45f3c96674af4a928297

Observation 662fd4c4-96c5-4aa1-a11b-8950e0b57382 · outbound

This paper cites SYSTEM_PROMPT_FOR_GRADING_MLLM_RESPONSE =’You are a helpful and precise assistant for checking the quality of the answer.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation SYSTEM_PROMPT_FOR_GRADING_MLLM_RESPONSE =’You are a helpful and precise assistant for checking the quality of the answer

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.651563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:54.929420Z digest=sha256:1108381e6781d43f8157f26c2bb83195be1f7da107b508f648bfe5be9c905df5

Observation 11274d10-18db-4291-b450-0af94b271366 · outbound

This paper cites 6, we show an example of using ImageReward [60] for the dataset curation by evaluating the alignment be- tween generated image and text prompts.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 6, we show an example of using ImageReward [60] for the dataset curation by evaluating the alignment be- tween generated image and text prompts

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.549201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:54.992499Z digest=sha256:d82c39d4dcbb4959674cdc706deb119f4484516a3b77c725251df2c05e612bb3

Observation a833ed89-9c6d-4823-a8ec-80624261cff8 · outbound

This paper cites 3, we perform quantitative comparisons on general image visual quality between different DM backbones: SD V1.5 and SDXL.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 3, we perform quantitative comparisons on general image visual quality between different DM backbones: SD V1.5 and SDXL

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.468624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:55.032116Z digest=sha256:c526468ca42fe6f6dd3fd9b4133d1f6ed4092186424a81f47a6d2a9d9de6ae1d

Observation d8e97d91-b1ca-4777-87d7-625d4fce5795 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:56.391242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:55.074253Z digest=sha256:efc0f6531529ea5bec33424fa296101613cd8252db18cfa63a989495013f6b76

Observation 6fd6cb42-cbf1-427b-a8aa-f8f26501081d · outbound

This paper cites 9 shows the UI page of our user study.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 9 shows the UI page of our user study

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.320779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:55.116586Z digest=sha256:2e49699887ddc4caa376ed60951fff21c732cc35101c4cbd66b055b013f09f14

Observation f80025b9-0df0-4d73-a614-011c00d5989c · outbound

This paper cites 10 to show more qualitative comparisons between fine- tuned LLaV A model to commercial SOTAs.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 10 to show more qualitative comparisons between fine- tuned LLaV A model to commercial SOTAs

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.259557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:55.153359Z digest=sha256:fead60fd7502b36637b04561f417243befbac7ced62abd22d2a7df06e9ae8b93

Observation ce8877c5-f5bf-41fc-b1a7-c7c842b22480 · outbound

This paper cites 11, we shows more examples of diversity on object categories, camera-object relation, and background con- texts.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation 11, we shows more examples of diversity on object categories, camera-object relation, and background con- texts

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:24:56.181755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:55.177843Z digest=sha256:d9ba739409128a49976da7703b22bd22e083db6da2ed98ed4edf1910d49c0ffb

Observation 7b09579b-8e9e-4749-9a60-76040f022433 · outbound

This paper cites an unresolved cited work.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:56.077268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:55.205611Z digest=sha256:2797c867d59517b472fe9d1d98175fd6237e5183fbf555b842756507629d9c0a

Observation ad2ed1f9-ff4d-4f55-9408-d2fbb7629d6e · outbound

This paper cites an unresolved cited work.

Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:24:55.986032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T18:24:55.227295Z digest=sha256:e52613a839deb030b46addda6f6dd96b98634ca3d944f1e88d371085d6873c6c

Pith citing papers

No inbound Pith citation observations are available.