Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2309.09958.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:51:51.535228Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T11:46:55.421431Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation fa1ac60c-8104-4d18-90e9-82b5f38ebdaa · inbound
A Survey on Multimodal Large Language Models An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a38909ca-2420-467b-8523-072020ed638c · inbound
Aligning Large Multimodal Models with Factually Augmented RLHF An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 787620c9-1f97-4da3-9880-ff724cba2328 · inbound
Improved Baselines with Visual Instruction Tuning An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c9f63bcf-432e-49e3-9dc3-e9c5d50b076a · inbound
InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 036e0362-408a-4ef8-b126-f2fd49856bf5 · inbound
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36740492-60e7-415b-8874-720003050ed2 · inbound
UnifiedCrawl: Aggregated Common Crawl for Affordable Adaptation of LLMs on Low-Resource Languages An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d97a5506-0c4f-4577-87f1-af88335367f4 · inbound
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb619d1e-1ece-43ea-a7d2-407c7a995493 · inbound
Liquid: Language Models are Scalable and Unified Multi-modal Generators An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbfcbc23-8912-4db2-95e2-113b5a6bfab1 · inbound
FineVQ: Fine-Grained User Generated Content Video Quality Assessment An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73eea413-02bd-4805-9e29-ebe6079d4b4b · inbound
Being-H0: Vision-Language-Action Pretraining from Large-Scale Human Videos An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e2d60e5-80c3-4e40-92a4-26d27bc026b4 · inbound
Describe, Adapt and Combine: Empowering CLIP Encoders for Open-set 3D Object Retrieval An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93d63d5a-131c-4ba5-bb99-15459bb4c382 · inbound
Transport Discrepancy as a Reliability Signal for Vision-Language-Action Models An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15ea76cd-a1b2-4006-9f8d-4528d65d5e48 · inbound
Balancing Image Compression and Generation with Bootstrapped Tokenization An Empirical Study of Scaling Instruct-Tuned Large Multimodal Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.