Pith. sign in

Paper Citation Record · LEDGER

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing

As of 19 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 4 inbound Pith citation observations for arXiv:2506.19848.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19848 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:21.640667Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T23:55:50.973218Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T00:10:52.889170Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved52
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1194e040-1202-4a3b-a3ab-e4493c099e2e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.346925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.346925Z digest=sha256:95f43fbe52aedbf100fec7de484398c473619bac88ac2340432d1f42807e3e05

Observation 0fd1d742-2250-4161-bcd0-a63d84662a02 · outbound

This paper cites Qwen2.5-VL Technical Report.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.352597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.352597Z digest=sha256:f67dfad2cbe694217181a0cf8572cc8071ad828215aa6c1f8ebe811c2352fd07

Observation 899d65ea-605f-4324-8dbc-9a462037e244 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Hallucination of Multimodal Large Language Models: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.357011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.357011Z digest=sha256:484a7763203c9e6317b57a4fa226bed32e75859c798b0b5c45e796b953bff474

Observation cc31f824-cfde-40f6-8a57-3e95bef40b48 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.361234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.361234Z digest=sha256:fb7bbf6130c03af30d4e6f93627905e72cb2baa2fcc5f785dae5ad56467dc09b

Observation 86b0af9f-37cd-4421-80de-f080c3d507a1 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.365268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.365268Z digest=sha256:15406fec1e514f715746d8b44ab3acde1b4ed91fb444ce965f1ab58be75a34e4

Observation 8ddf7c82-2a1f-4a43-b308-42963a0c94ce · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Sharegpt4v: Improving large multi-modal models with better captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.369879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.369879Z digest=sha256:58694398b72e468ef45bdba3188239077be37da47b033dc3f579ef4a2165ed65

Observation 232bb052-a488-4139-b90a-db696ee72945 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495, 2024.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.374112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.374112Z digest=sha256:222907a23e8b49ebc45853aedbed0ca03ee336cb4f313535317fab4f97c6991f

Observation 956d6ff3-2e1c-4782-87d2-19cb45576514 · outbound

This paper cites Open-llava-next: An open-source implementation of llava-next series for facilitat- ing the large multi-modal model community.https://github.com/xiaoachen98/Open-LLaVA-NeXT , 2024.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Open-llava-next: An open-source implementation of llava-next series for facilitat- ing the large multi-modal model community.https://github.com/xiaoachen98/Open-LLaVA-NeXT , 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.464216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.378126Z digest=sha256:a750a91f5daabdb4355bbe3f4bb3c28d62ab5f146164278dc2ba3f325f5e4878

Observation f74ad12e-996c-4c0c-932d-f67a3b5dfee8 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.382358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.382358Z digest=sha256:9d4e64923a187cf8d0ea556f594662eb373b2e4a877e4550da344ce0a353a859

Observation 58999624-24c1-4df0-8a88-88d2e369d97c · outbound

This paper cites Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.386756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.386756Z digest=sha256:51a8cd8e59543ca6d3500379418c2ca411ef64d21724d59f7b74883b0fb0e410

Observation 14802a9e-513a-4870-bb1c-641c6bef301a · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.391223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.391223Z digest=sha256:f3073b919f13cf2071443ffba3f9168030f978e7102a870aa75df7251e0ce922

Observation 7eb07c46-21a3-4682-9278-b0c890ced8d4 · outbound

This paper cites MM-IFEngine: Towards Multimodal Instruction Following.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing MM-IFEngine: Towards Multimodal Instruction Following

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.395096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.395096Z digest=sha256:9b10310c97fb63fe8c2c0700a1188b10caa4f24e20167966c107a07d868f72f5

Observation ba3f5f63-7133-45d7-83aa-f1f0dc1c0b6c · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Benchmarking and Improving Detail Image Caption

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.398824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.398824Z digest=sha256:5361fd3e4577db0a3856a308c25fbab6f6d700557f683f6caf33c010c6d64217

Observation dfdcdcb9-dde5-4a40-ad37-0f57f8a02c56 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.402911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.402911Z digest=sha256:a727cde3a6dfa989a35839a2288dbc3ad655915e217d94a46f4598ffe6f0f0f9

Observation 249a3e78-117b-4c22-9298-b07413a8067c · outbound

This paper cites Improving clip training with language rewrites.Advances in Neural Information Processing Systems, 36:35544–35575, 2023.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Improving clip training with language rewrites.Advances in Neural Information Processing Systems, 36:35544–35575, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.444486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.406919Z digest=sha256:016442409b4ba55417ef44ea42efa4c11ceba245a0aae1ca432b96c9998e996a

Observation f5509eb8-46c1-446e-bb9a-5d5541675a7b · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Eva: Exploring the limits of masked visual representation learning at scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.410549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.410549Z digest=sha256:76a2ac96f65ff3de374a8a9763f50ab3011ea144d3d5fcbf16c285bcab7412d6

Observation 49a73119-b28f-49ef-af48-e4ad5c1f9b3a · outbound

This paper cites Ic9600: A benchmark dataset for automatic image complexity assessment.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(01):1–17, 2023.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Ic9600: A benchmark dataset for automatic image complexity assessment.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(01):1–17, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.423241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.414333Z digest=sha256:94848233545c6ee220da06dee5cdbdd3a69a38aec39518060a69b8fb93fa7b5f

Observation c9d7b64a-425c-4940-bbed-bff2ab2b7a46 · outbound

This paper cites Imageinwords: Unlocking hyper-detailed image descriptions, 2024.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Imageinwords: Unlocking hyper-detailed image descriptions, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.410900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.417966Z digest=sha256:726222dd9367b291a575e7fe78cb6e3fab6bb5181cbf4291b23aa80baab41786

Observation dc754695-bbd4-46a3-8f63-b0955ad04015 · outbound

This paper cites The Llama 3 Herd of Models.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.421614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.421614Z digest=sha256:a93e842ba875c8bd5ff2f018bcbe8ce0e8e1ac330b1da085b5c34397b23c8e9d

Observation 95e41511-0bb0-44f7-8dc9-d0b7ee1dc27a · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.425500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.425500Z digest=sha256:c0dd844ddf498c44309094cb5086764dd89cad70878c0c93bb323f920b177b6f

Observation cb56781f-ce3b-41b9-8530-dd91f0f952f8 · outbound

This paper cites Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.429233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.429233Z digest=sha256:0adc80c71aeb34e3636c04c93fad401bef099c212aa153dac6e5da5778a28210

Observation 538615cf-305a-4137-9b6d-9e3a569fbaa3 · outbound

This paper cites Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.433181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.433181Z digest=sha256:14abcdbad25846ec475f84cc2d46c4ab2567ee7b6c5b4f71c2a1ccdc2bb136ce

Observation 6a6508ee-8599-46df-a5e0-f797a2eba695 · outbound

This paper cites Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.437515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.437515Z digest=sha256:03bec73fd2317947e801f7ef1600317a013d8818e0d58f5587f89dceab53e7d7

Observation a078e7be-9a77-4f55-a236-7c4e43a8ea2f · outbound

This paper cites CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.441372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.441372Z digest=sha256:b0e56fbbae61891ed1133027c9b86587702a471cebf0a5a6e7cc4ce23a79b4b1

Observation f031ab7f-1aa4-4258-9539-659958223b2c · outbound

This paper cites Adu: Adaptive detection of unknown categories in black-box domain adaptation.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Adu: Adaptive detection of unknown categories in black-box domain adaptation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.382225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.445439Z digest=sha256:79c7f64bcb48fda245696f0a58413c0143c31a385df46de549a236cee6253b88

Observation bd2a8f2c-a29f-435e-8031-592e311144f1 · outbound

This paper cites Veclip: Improving clip training via visual-enriched captions.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Veclip: Improving clip training via visual-enriched captions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.449624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.449624Z digest=sha256:e915ff7855c3752aee0305398a5c8517ff373b77b82761d78c90fe432c6bb5b4

Observation 4262ea31-b8cc-42bb-bd54-1ed105f262c9 · outbound

This paper cites LAION-5B: A New Era of Open Large-scale Multimodal Datasets, 2022.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing LAION-5B: A New Era of Open Large-scale Multimodal Datasets, 2022

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.363061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.452923Z digest=sha256:52099c12c4ca62c9b4f54fbe84f5863a705e89717b4950331ccf435f99cb175d

Observation 40d48064-ee00-4ebc-984a-1f07bbe261d6 · outbound

This paper cites Mitigating object hallucinations in large vision-language models through visual contrastive decoding.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Mitigating object hallucinations in large vision-language models through visual contrastive decoding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.456427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.456427Z digest=sha256:fef2a0eb60e25e1bb23d41ba2806409b3870dd9891cc93d74e33b29cb215a6d2

Observation 28b79ed9-5551-4b82-b2e8-fd27e63b5c65 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.343134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.459848Z digest=sha256:f298df584d3c6d8a624136d08c7c5a8af63cdb4b233007821b7699429075cfcd

Observation c4d50d11-d1ff-4973-92ac-d88b9885b378 · outbound

This paper cites DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.464143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.464143Z digest=sha256:ee9b0a5a15b9486c32d848364776e734fb6b05f71fd3bc3d3570376255844e41

Observation f1ca1f7c-f38e-4397-a5d0-a0c1492abd04 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Evaluating Object Hallucination in Large Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.467770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.467770Z digest=sha256:47dde5067dcada14bfb359b60319aaca3f1e22e642bed27486a77e5a12a21f21

Observation a2a961e1-a5cd-4bdf-86c7-4a305cb9aaf9 · outbound

This paper cites an unresolved cited work.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:30:22.330559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.471445Z digest=sha256:057447365bc6213f573675ee13b52851098c50639dab8fb5a321ed10e43f161c

Observation 010ba671-e19f-4418-8be8-493dc8c5cd0e · outbound

This paper cites Grounding 3D Scene Affordance From Egocentric Interactions.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Grounding 3D Scene Affordance From Egocentric Interactions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.475174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.475174Z digest=sha256:cf1c0b9e551888076db8986ee7b6205070d43d071461f417a99196cecf4379d5

Observation 51e84157-12f6-4cc8-883d-8f55af0794ae · outbound

This paper cites Improved baselines with visual instruction tuning.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Improved baselines with visual instruction tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.479353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.479353Z digest=sha256:3d699c28cb22e63430281a961d276f606203d8f58c28993a3d58975f45af7ab4

Observation 3e1eedf8-ceab-4d1a-ab28-8e8c2422557c · outbound

This paper cites Visual Instruction Tuning.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.483286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.483286Z digest=sha256:55f29cfa0819fbb88b0a87a01348f63fd267726348e6d8bc35b8e242a43b598a

Observation b1edc3c0-1566-4916-8a7b-7a8d2d214b49 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing A Survey on Hallucination in Large Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.487659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.487659Z digest=sha256:283ec1217afe90d6051510c838ee283b7383472256d08aec06c4f29ae74e1e45

Observation ff9c46d3-5958-4ca8-bd8c-4c935b12e1bb · outbound

This paper cites Paying more attention to image: A training-free method for alleviating hallucination in lvlms.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Paying more attention to image: A training-free method for alleviating hallucination in lvlms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.492214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.492214Z digest=sha256:5b8fcc6cf65e114ef7640d7dd5734868f85723884d7f36b814209c1279cd9380

Observation 8983ad3d-ffe6-423a-8e4c-056a0479e28c · outbound

This paper cites MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.496411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.496411Z digest=sha256:9864a1c62b3c451beed79cd3abff23dbb9fb7abb63582213ff36414dadda77b7

Observation ac6a5000-de90-4487-8fa0-eb6106fd6337 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.501310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.501310Z digest=sha256:4b7ad7aed03234f12b3bc48cc66a4c765a369b5b5fbf92cc584be4cf7ef3fc00

Observation 92905a7b-6ba2-4c25-93fe-49f1133f1126 · outbound

This paper cites MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.507165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.507165Z digest=sha256:668f9b1ece50e4e466d0cf8c5c414bfa64e425ccd466b337d9e138a65c0dc3f6

Observation 35b2dd15-f3a9-46ee-b058-8902a4d84537 · outbound

This paper cites LLaMA 4 Models, 2025.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing LLaMA 4 Models, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.304938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.512298Z digest=sha256:a7eddc200867a1fde30093b8934890ae52b3eeaf30e4dcc91022b0db9eb04062

Observation b910c38b-2a8c-4195-ad35-4ff7e7546599 · outbound

This paper cites Docci: Descriptions of connected and contrasting images.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Docci: Descriptions of connected and contrasting images

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.516625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.516625Z digest=sha256:5758d676a0f6ecae2ce6aae52e800f673cb2fc72929a629e8a24b0f8a3618e73

Observation 1cfa5e0a-e2a3-4f4d-a2c0-53b09093fb38 · outbound

This paper cites Chatgpt.https://chat.openai.com/, 2023.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Chatgpt.https://chat.openai.com/, 2023

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.284873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.520838Z digest=sha256:ffec43e467a0501d666ac39f7616deb385573e6027d0316f8f6ace93049af784

Observation e8687744-f4d8-4ddd-aecb-8e6acf57ac4d · outbound

This paper cites Gpt-4v(ision) system card.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Gpt-4v(ision) system card

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.525501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.525501Z digest=sha256:03787fd39fffee501061b6099f973610365109263c6d9d42f641b55db85e5983

Observation 688261e9-cdbd-4cb3-a5d7-cfecd01d483d · outbound

This paper cites Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.532404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.532404Z digest=sha256:e3ed95552d1ffdcc49039c755790fd40de858033b75fff3b8fb85403e7faad60

Observation 336d5bca-329a-45bd-a2d9-d90580afc681 · outbound

This paper cites Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.537076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.537076Z digest=sha256:0f6ecd492b006a6d8880aff55b33636ec2573566f5b465fffc74d9416eec606f

Observation 1131b7a0-48d8-4fba-82e4-9a6c75aa0aa3 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Glamm: Pixel grounding large multimodal model

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.542132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.542132Z digest=sha256:25b75e351ced7f3e20addd69e6ad44d9e5ca5f96c209fabb4fb2b2d0b4409a58

Observation 102e882d-e025-46db-a6ad-55f00a30f7f8 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.257432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.546515Z digest=sha256:51c1c065bc50bbebceffce5ce366024c04bdb1e37d8f2fa9fb0f4651717d334a

Observation e69665ae-af7c-4524-82ba-19e7859f0dd6 · outbound

This paper cites Descriptive caption enhancement with visual specialists for multimodal perception.arXiv preprint arXiv:2412.14233, 2024.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Descriptive caption enhancement with visual specialists for multimodal perception.arXiv preprint arXiv:2412.14233, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.550601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.550601Z digest=sha256:951adb438ad3496b0621a4e897977b0ba82eab6688cfa69928828f7c05acd740

Observation 03e39ec1-c5dc-4dcb-a655-5aad110f1693 · outbound

This paper cites X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.554689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.554689Z digest=sha256:be2b606576b64864021d234c6f0218d6c49ea6716579327b5408399f6849cb7a

Observation b9eda35a-64a6-4952-a0ab-2e5a5bfd0135 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.558769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.558769Z digest=sha256:448df63c917e46953e17fe18ae65da4e9db9c85b105577b0407b206fa64957ad

Observation aa5e912d-7b66-46b2-bc96-acdecc3f72d6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.562862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.562862Z digest=sha256:1adfacf296d211b2b065a022ed8c049a16409c0ab74504c930de78feaba7d757

Observation bc5f46d0-dc93-4738-a8ef-aaf49c27b2e0 · outbound

This paper cites A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.243677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.566972Z digest=sha256:266dbefe4cc888a5bfd212ea3ad1f13e0b460057baf0a3273b74e12c4dfef969

Observation eea95389-3355-455a-a68d-066ff04460be · outbound

This paper cites Common subspace for model and similarity: Phrase learning for caption generation from images.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Common subspace for model and similarity: Phrase learning for caption generation from images

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.228582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.570723Z digest=sha256:f0c61c226bbf59a4d54842e89f675ba0c4637f52fa8a96c76b46e3b25b333f43

Observation a97b1593-3422-49d3-82da-ee561ba61b14 · outbound

This paper cites Contrastive region guidance: Improving grounding in vision-language models without training.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Contrastive region guidance: Improving grounding in vision-language models without training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.209162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.574689Z digest=sha256:612f351d2c7560981f82d000f40113faa4cf31c8eadbdfc1144cc4502938d07f

Observation 59c74586-ed16-4de0-9da6-2c34d197fbed · outbound

This paper cites Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.579345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.579345Z digest=sha256:47a3c64f87d9ce685af1cb3bb1ca95c1df9ebc369bf52852b05da06968c13d63

Observation f167734f-7c30-4b59-8533-132b5d573bba · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.583543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.583543Z digest=sha256:5e477c41cb811fde0a2579c408c9772e730f629668226a263e07462b98386607

Observation 382215c7-67c8-4b50-abca-e0ceaa465502 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.587807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.587807Z digest=sha256:cf4591aec77c381aba88d397f74dac6960ca0058d31cf5fe037481b9a8d25c7a

Observation 64806f49-161e-4d56-aa8e-cb26f5dd936c · outbound

This paper cites Altogether: Image Captioning via Re-aligning Alt-text.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Altogether: Image Captioning via Re-aligning Alt-text

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.592235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.592235Z digest=sha256:42089a324baef59e9a2fe62b673893dc7b335a415b9bedf9f5210c201f68ed51

Observation cc18bfe2-8182-4c96-99cf-dd688cd20071 · outbound

This paper cites Woodpecker: Hallucination correction for multimodal large language models.Science China Information Sciences, 67(12):220105, 2024.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Woodpecker: Hallucination correction for multimodal large language models.Science China Information Sciences, 67(12):220105, 2024

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.596932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.596932Z digest=sha256:23e593549e28a7d09734aa2a31987eff434226a49bfe3a6eb2db5d3aa11b4d26

Observation 1895fc64-31f8-46d5-8074-6069960cf683 · outbound

This paper cites Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.600960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.600960Z digest=sha256:0f02ea49cf14b2e9fbb3b1efd95157fb1925e763d0981884a3b714b1806b5b0f

Observation dbc28bfb-e29b-4a74-93fd-ba4ddd4207ca · outbound

This paper cites Capsfusion: Rethinking image-text data at scale.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Capsfusion: Rethinking image-text data at scale

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.181030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.604860Z digest=sha256:d92139f850d8e4567aad78c46899ffb48582e633a7e71c3be1b3f2d7237ef249

Observation cb9eac3c-d3a8-4456-85a5-82a3f4b5f632 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.609708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.609708Z digest=sha256:098a5f0f1fb738de8607464e57ce29a64b3d4f0f62ce7c15e3b3f75c763c67fd

Observation 7f2a9247-4cf8-42d9-9435-fe23fb62ed5a · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.613881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.613881Z digest=sha256:8ac3b12d227d7b573e5d493d20333564111bc84c71125b9f2996dcb2d0ff3e10

Observation 6a49984a-583c-4989-8c47-235ffc566225 · outbound

This paper cites Recognize anything: A strong image tagging model.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Recognize anything: A strong image tagging model

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.617932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.617932Z digest=sha256:8a539f6e0f898bd482d916fefb6ffcaba2f2f71f15c6c128ca64a85bfad21f90

Observation 320140f0-88f5-461d-ae07-408a96261f17 · outbound

This paper cites OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-15T18:30:21.622247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:30:21.622247Z digest=sha256:cbb59e36a07b1f15bd4b399386ddc55bafb02780ac329c25cde3e761f0ef7ffe

Observation 0ca300c7-a156-4f3f-affb-1acd58425647 · outbound

This paper cites Describe more details about the [Object].

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Describe more details about the [Object]

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.155723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.626340Z digest=sha256:5b9d20885ff1dddf05a90fecf0a2fa33bb0360c40943619aac1a45ef552f8b1c

Observation b246e20a-7d31-4762-898a-77100b275c04 · outbound

This paper cites Instructions: Describe more details about the oven.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Instructions: Describe more details about the oven

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.143730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.630653Z digest=sha256:f46b3aab7a01d57e1e7301df64bdabcdd2a60e90d6a55f64dbfef1b4f29bb44d

Observation 2dea20cf-015d-4cf2-9a37-2921312b64e1 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:30:22.116183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.640667Z digest=sha256:11a6f0247c49de9258bb9b05298539491ed62a485530412c506229e28e719ade

Observation b5cb4164-54b6-40ad-939c-81202bf559b6 · outbound

This paper cites Limitations.

ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Limitations

Reference 256

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T18:30:22.131012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:30:21.635084Z digest=sha256:38d41ee6fbc38e812760c4334cca53211148962546aef11fbcc4d63409b5adbc

Pith citing papers

Observation 61b793eb-4b1f-449e-a480-548d2d78de90 · inbound

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience cites this paper.

SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T23:55:50.973218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:55:50.973218Z digest=sha256:afebfb55b09f3cba79978b06a52241cb8a3d6e017ef783c1d7a76d87eb04a1ff

Observation 909b480f-cc88-4ceb-92a9-a48cf0da6589 · inbound

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning cites this paper.

CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:50.902512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:50.902512Z digest=sha256:5101beb398d0fe012c268d19faa9ec4e5f3c80c8c967217af95d3987b63fad70

Observation b1857817-b6bf-460c-b2cf-5b9ddea88f20 · inbound

Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning cites this paper.

Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T12:31:19.632698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:31:19.632698Z digest=sha256:74b1fc99f220f20ded18edf015d0e1a04f552049c4ae9e8f5a70d5c42e1e9b00

Observation 06ac6dee-7941-4b8d-b165-2aa7ce6d0158 · inbound

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions cites this paper.

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:10:52.894245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:39:29.592129Z digest=sha256:afe10cf2a6168278937638e3fde4be6ce8109d8fec650be7ce5c59ae074da2ee