Pith. sign in

Paper Citation Record · LEDGER

An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2309.09958.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.09958 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:51:51.535228Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T11:46:55.421431Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fa1ac60c-8104-4d18-90e9-82b5f38ebdaa · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.517467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:d583b075a59664f7be2ea2200285e321e33784b454817fdb075541088d09041f

Observation a38909ca-2420-467b-8523-072020ed638c · inbound

Aligning Large Multimodal Models with Factually Augmented RLHF cites this paper.

Aligning Large Multimodal Models with Factually Augmented RLHF An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:58:17.884879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T17:58:17.699042Z digest=sha256:87ccaca01de02d52d79c4f075f6eed741c823af86b046b1a74599d7cc560ce39

Observation 787620c9-1f97-4da3-9880-ff724cba2328 · inbound

Improved Baselines with Visual Instruction Tuning cites this paper.

Improved Baselines with Visual Instruction Tuning An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:11:33.934444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T19:11:33.783746Z digest=sha256:a834e0b0027956031dff36c1f15f8492c4ce6f94cd2e106604208010e45dcf56

Observation c9f63bcf-432e-49e3-9dc3-e9c5d50b076a · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:46:10.054073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:379fe8f3bac6c836e90884805b948f7c9d6abee3aac0306c9589934fd726f0ee

Observation 036e0362-408a-4ef8-b126-f2fd49856bf5 · inbound

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts cites this paper.

Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:51:51.535228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:51:51.535228Z digest=sha256:fbd143c2bc52031548cf3a6447d68261ea9153d5c8cdc0c03360816d8b396602

Observation 36740492-60e7-415b-8874-720003050ed2 · inbound

UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages cites this paper.

UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:20:31.418746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:20:31.418746Z digest=sha256:e56e56a3091236a9b1de93723c7d2ade234e03230c0dac305f703bae2a968890

Observation d97a5506-0c4f-4577-87f1-af88335367f4 · inbound

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion cites this paper.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.380002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.380002Z digest=sha256:c4f46ef39d14dd82c8ef3c1d00cf4b4d13d251d3e9455d13eede432acb78df85

Observation bb619d1e-1ece-43ea-a7d2-407c7a995493 · inbound

Liquid: Language Models are Scalable and Unified Multi-modal Generators cites this paper.

Liquid: Language Models are Scalable and Unified Multi-modal Generators An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:40.375788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:35:40.375788Z digest=sha256:cb6e6b8f789c0e6cd821f14bff241e17d9eff8e5d4b166726332b47e789abb4b

Observation fbfcbc23-8912-4db2-95e2-113b5a6bfab1 · inbound

FineVQ: Fine-Grained User Generated Content Video Quality Assessment cites this paper.

FineVQ: Fine-Grained User Generated Content Video Quality Assessment An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T00:53:00.769440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:53:00.769440Z digest=sha256:cccf72e7db5427e4eacfcb25cf964c3d59bfd99671f31ca7f4c42b39e474ca01

Observation 73eea413-02bd-4805-9e29-ebe6079d4b4b · inbound

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos cites this paper.

Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:33:41.744204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:33:41.744204Z digest=sha256:7c5f2a7610ee9448bfe689c58ebbbb940329b1ca580c882754b2a8089b5453e4

Observation 4e2d60e5-80c3-4e40-92a4-26d27bc026b4 · inbound

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval cites this paper.

Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T12:47:05.523805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:47:05.523805Z digest=sha256:86d84521fc36477a7a1d170d42da9794591cbb03432a6c7bebcc199a7a62b687

Observation 93d63d5a-131c-4ba5-bb99-15459bb4c382 · inbound

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models cites this paper.

Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T19:11:42.842973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:11:42.842973Z digest=sha256:87fd90ea1b0547cfaab115bc40079c6d090374ef34d7fc837c75a5c49c3f589c

Observation 15ea76cd-a1b2-4006-9f8d-4528d65d5e48 · inbound

Balancing Image Compression and Generation with Bootstrapped Tokenization cites this paper.

Balancing Image Compression and Generation with Bootstrapped Tokenization An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:46:55.423044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T03:07:33.054518Z digest=sha256:4201159bd6bec4124c6177f49666cb933ed0433ae125bd7ba0ca0ab367c36160