Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2309.09958.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.09958 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:51:51.535228Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T11:46:55.421431Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fa1ac60c-8104-4d18-90e9-82b5f38ebdaa · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.517467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:210928f6cffa61346a09c04fcbf0f352cbfd932c820f39d5f06d3ff03e6e799a

Observation a38909ca-2420-467b-8523-072020ed638c · inbound

Aligning Large Multimodal Models with Factually Augmented RLHF cites this paper.

Aligning Large Multimodal Models with Factually Augmented RLHF An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:58:17.884879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T17:58:17.699042Z digest=sha256:645151cf5ddd71761e86749efd16a7485137b2eb20b1a822c81d718284e1c739

Observation 787620c9-1f97-4da3-9880-ff724cba2328 · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:33.934444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:5cfd788055a907c989af0ca0b52dfa51bcafa2114f778dd99c4ddc135c1db6e2

Observation c9f63bcf-432e-49e3-9dc3-e9c5d50b076a · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.054073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:aa96c41d2c1dae9fe1d731a0d8fefd65b21e9dc43ffbce07a4ba51a7c5855069

Observation 036e0362-408a-4ef8-b126-f2fd49856bf5 · inbound

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts cites this paper.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.535228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.535228Z digest=sha256:bbe2f44adfc3f7f6b5f1d50427a9375e8da21144216bb2ed2afdc2a8497ca589

Observation 36740492-60e7-415b-8874-720003050ed2 · inbound

UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages cites this paper.

UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:20:31.418746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:20:31.418746Z digest=sha256:799d21cddad12d4259af9a1291c056348698e3e46b9089a832340ab362073eef

Observation d97a5506-0c4f-4577-87f1-af88335367f4 · inbound

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion cites this paper.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.380002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.380002Z digest=sha256:781a1aa761971fb5c753507db87376556eb2937507e37cce6b23545ac56a0c2f

Observation bb619d1e-1ece-43ea-a7d2-407c7a995493 · inbound

Liquid: Language Models are Scalable and Unified Multi-modal Generators cites this paper.

Liquid: Language Models are Scalable and Unified Multi-modal Generators An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:40.375788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:35:40.375788Z digest=sha256:b2b5a013f19707473d4aa75a99f05ae63a2506852d3b30ed50e8a131201dad16

Observation fbfcbc23-8912-4db2-95e2-113b5a6bfab1 · inbound

FineVQ: Fine-Grained User Generated Content Video Quality Assessment cites this paper.

FineVQ: Fine-Grained User Generated Content Video Quality Assessment An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:53:00.769440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:53:00.769440Z digest=sha256:d31a4f7b58416f248fba484ae7b2fca6f3dfa41243f6ff1eac1e1c28e75113e9

Observation 73eea413-02bd-4805-9e29-ebe6079d4b4b · inbound

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos cites this paper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.744204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.744204Z digest=sha256:be3e16d6baa3e01ac45e98ab86aa8d0dcdc709d601b02ed6bb209ef44630d2ae

Observation 4e2d60e5-80c3-4e40-92a4-26d27bc026b4 · inbound

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval cites this paper.

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T12:47:05.523805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:47:05.523805Z digest=sha256:5070574ff313ef51ea28bd46ca5048d2f820da47f120a7e7eabccf39ad4e898e

Observation 93d63d5a-131c-4ba5-bb99-15459bb4c382 · inbound

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models cites this paper.

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T19:11:42.842973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:11:42.842973Z digest=sha256:79a01a0163e8de56dc5b6b1bd731dbf35b6e92de98ee47fc8bc409915c350d2f

Observation 15ea76cd-a1b2-4006-9f8d-4528d65d5e48 · inbound

Balancing Image Compression and Generation with Bootstrapped Tokenization cites this paper.

Balancing Image Compression and Generation with Bootstrapped Tokenization An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.423044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T03:07:33.054518Z digest=sha256:9619bbf84398b2bdd8ebda7ba15197c9b122d5a3ea7f02568934afa28c1bf2b4