Pith. sign in

Paper Citation Record · LEDGER

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures

As of 21 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2507.18009.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18009 v1

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:43:00.032297Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 02b908cb-b9f0-41d1-8dfb-5d4185b77753 · outbound

This paper cites GPT-4 Technical Report.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.986783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.986783Z digest=sha256:f3e4125ea7683bc807fcef1ea3feeedd13c5b785c18942b8940ffa3b90774c8f

Observation 4959d51a-6a5c-41ad-af12-928e4361ad27 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.993290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.993290Z digest=sha256:9597435547c41fbf89e5f87dbd60c472a8c74a25f51244591e9501d815a9149e

Observation e4af88cd-e715-43e2-8ae2-81172465c4b6 · outbound

This paper cites GLU Variants Improve Transformer.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures GLU Variants Improve Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.000440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.000440Z digest=sha256:808d660dc412ba4004bac756f273a0602022834b4d15192e51964be59cf81d5d

Observation 366b64a3-709b-4993-abe5-1beac0f64fe9 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Gemma 2: Improving Open Language Models at a Practical Size

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.005418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.005418Z digest=sha256:ede59a11df60d8f5af1458b2c9a20ffa165cef387499a632e34f0d5ed76d2782

Observation 2bf0d6a8-50f7-4d32-bdd5-0a79c85994fc · outbound

This paper cites Nemotron-4 340B Technical Report.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Nemotron-4 340B Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.011505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.011505Z digest=sha256:0eea1b41e9cd101943089b8be6894e7899ae379597acdb5bcbe8d67a0bf38892

Observation 105069bb-a8f4-4e98-8efc-ce049bef17fa · outbound

This paper cites Qwen2 Technical Report.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Qwen2 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.015406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.015406Z digest=sha256:03227d35cc7e91b3cc68da1726dfac7ce29a9964257929b0ff1de1becc231d0a

Observation e6dfe577-ba5c-448c-aeeb-3d3faf8fc4f0 · outbound

This paper cites an unresolved cited work.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:43:00.211792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T14:43:00.018854Z digest=sha256:48aa194e422cea7a45e72fd20df2bee58d4bc53b69afe87d772bfff711e1cda7

Observation dc74ace1-043e-4c9d-bc44-68ec55b530a9 · outbound

This paper cites PubMedClip: How much does CLIP benefit visual question answering in the medical domain? In Findings of the Association for Computational Linguistics: EACL 2023, pages 1181–1193,.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures PubMedClip: How much does CLIP benefit visual question answering in the medical domain? In Findings of the Association for Computational Linguistics: EACL 2023, pages 1181–1193,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:43:00.200477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T14:43:00.022201Z digest=sha256:281c1cc31147eab371eb13954e79e2f70c3b2c47a052da408f6328b962e2ef21

Observation dc60dfb6-d565-4060-9dee-5f4b50b356c5 · outbound

This paper cites Decoupled Weight Decay Regularization.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Decoupled Weight Decay Regularization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.028975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.028975Z digest=sha256:4dc9f36f2cc92e3a4ad5e8636904c2f9d70dc7c52bb5a7be013f3a9e38ca6be7

Observation 780af361-7dfb-4aec-aae2-d6d3920c797b · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.032297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.032297Z digest=sha256:fec14853277fe9c83665a6ecb6e5ac4b94c51d40f6a8991bad63f10725330776

Observation 38bce55e-d6ab-4720-8219-927013cdbeb7 · outbound

This paper cites The Llama 3 Herd of Models.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures The Llama 3 Herd of Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.978684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.978684Z digest=sha256:c52a6eb720fe83d56ba20bfee92aa92e2a2b6a26166ec9a898d630a6d8ff207e

Observation b2d2b7a8-e42a-41f3-9700-33d02ff64959 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures Imagenet: A large-scale hierarchical image database

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:00.025506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:00.025506Z digest=sha256:6cb8f8d982e2a56c62a7c5de85ae6819d7f7f2079dd6e6db8509dbd5c4cc021b

Observation 7439bea8-daca-4d88-8499-61210ebfd298 · outbound

This paper cites VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.996885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.996885Z digest=sha256:cac81b3d0a508a7ff38f4712679ad622926efbe0fe6822ee7892ca181f607336

Observation b405ada9-b939-422c-b40a-30ae82eda205 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.989946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.989946Z digest=sha256:eeb69ab8582b7b8fd6bd68e28a02694649d010842414c9f1dfed6b571e983b03

Observation 385e81e0-0776-48e7-a04a-dae1ff7b274f · outbound

This paper cites DeepSeek-V3 Technical Report.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures DeepSeek-V3 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.982607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.982607Z digest=sha256:6c4fd0890059062295be4f3fdb62a944116f38c76979da74076f42f3647ff745

Pith citing papers

No inbound Pith citation observations are available.