Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:21.640667Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 4 inbound Pith citation observations for arXiv:2506.19848.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:30:21.640667Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T23:55:50.973218Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T00:10:52.889170Z
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1194e040-1202-4a3b-a3ab-e4493c099e2e · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fd1d742-2250-4161-bcd0-a63d84662a02 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 899d65ea-605f-4324-8dbc-9a462037e244 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Hallucination of Multimodal Large Language Models: A Survey
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc31f824-cfde-40f6-8a57-3e95bef40b48 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86b0af9f-37cd-4421-80de-f080c3d507a1 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ddf7c82-2a1f-4a43-b308-42963a0c94ce · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Sharegpt4v: Improving large multi-modal models with better captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 232bb052-a488-4139-b90a-db696ee72945 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Sharegpt4video: Improving video understanding and generation with better captions.Advances in Neural Information Processing Systems, 37:19472–19495, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 956d6ff3-2e1c-4782-87d2-19cb45576514 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Open-llava-next: An open-source implementation of llava-next series for facilitat- ing the large multi-modal model community.https://github.com/xiaoachen98/Open-LLaVA-NeXT , 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f74ad12e-996c-4c0c-932d-f67a3b5dfee8 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58999624-24c1-4df0-8a88-88d2e369d97c · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Holistic Analysis of Hallucination in GPT-4V(ision): Bias and Interference Challenges
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14802a9e-513a-4870-bb1c-641c6bef301a · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7eb07c46-21a3-4682-9278-b0c890ced8d4 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing MM-IFEngine: Towards Multimodal Instruction Following
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba3f5f63-7133-45d7-83aa-f1f0dc1c0b6c · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Benchmarking and Improving Detail Image Caption
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfdcdcb9-dde5-4a40-ad37-0f57f8a02c56 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 249a3e78-117b-4c22-9298-b07413a8067c · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Improving clip training with language rewrites.Advances in Neural Information Processing Systems, 36:35544–35575, 2023
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f5509eb8-46c1-446e-bb9a-5d5541675a7b · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Eva: Exploring the limits of masked visual representation learning at scale
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49a73119-b28f-49ef-af48-e4ad5c1f9b3a · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Ic9600: A benchmark dataset for automatic image complexity assessment.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(01):1–17, 2023
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c9d7b64a-425c-4940-bbed-bff2ab2b7a46 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Imageinwords: Unlocking hyper-detailed image descriptions, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dc754695-bbd4-46a3-8f63-b0955ad04015 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing The Llama 3 Herd of Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95e41511-0bb0-44f7-8dc9-d0b7ee1dc27a · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb56781f-ce3b-41b9-8530-dd91f0f952f8 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538615cf-305a-4137-9b6d-9e3a569fbaa3 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Deciphering Cross-Modal Alignment in Large Vision-Language Models with Modality Integration Rate
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a6508ee-8599-46df-a5e0-f797a2eba695 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a078e7be-9a77-4f55-a236-7c4e43a8ea2f · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing CATCH: Complementary Adaptive Token-level Contrastive Decoding to Mitigate Hallucinations in LVLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f031ab7f-1aa4-4258-9539-659958223b2c · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Adu: Adaptive detection of unknown categories in black-box domain adaptation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation bd2a8f2c-a29f-435e-8031-592e311144f1 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Veclip: Improving clip training via visual-enriched captions
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4262ea31-b8cc-42bb-bd54-1ed105f262c9 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing LAION-5B: A New Era of Open Large-scale Multimodal Datasets, 2022
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 40d48064-ee00-4ebc-984a-1f07bbe261d6 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Mitigating object hallucinations in large vision-language models through visual contrastive decoding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28b79ed9-5551-4b82-b2e8-fd27e63b5c65 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c4d50d11-d1ff-4973-92ac-d88b9885b378 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1ca1f7c-f38e-4397-a5d0-a0c1492abd04 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Evaluating Object Hallucination in Large Vision-Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2a961e1-a5cd-4bdf-86c7-4a305cb9aaf9 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 010ba671-e19f-4418-8be8-493dc8c5cd0e · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Grounding 3D Scene Affordance From Egocentric Interactions
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51e84157-12f6-4cc8-883d-8f55af0794ae · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Improved baselines with visual instruction tuning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e1eedf8-ceab-4d1a-ab28-8e8c2422557c · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Visual Instruction Tuning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1edc3c0-1566-4916-8a7b-7a8d2d214b49 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing A Survey on Hallucination in Large Vision-Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff9c46d3-5958-4ca8-bd8c-4c935b12e1bb · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Paying more attention to image: A training-free method for alleviating hallucination in lvlms
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8983ad3d-ffe6-423a-8e4c-056a0479e28c · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac6a5000-de90-4487-8fa0-eb6106fd6337 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Visual-RFT: Visual Reinforcement Fine-Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92905a7b-6ba2-4c25-93fe-49f1133f1126 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing MIA-DPO: Multi-Image Augmented Direct Preference Optimization For Large Vision-Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35b2dd15-f3a9-46ee-b058-8902a4d84537 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing LLaMA 4 Models, 2025
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b910c38b-2a8c-4195-ad35-4ff7e7546599 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Docci: Descriptions of connected and contrasting images
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cfa5e0a-e2a3-4f4d-a2c0-53b09093fb38 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Chatgpt.https://chat.openai.com/, 2023
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e8687744-f4d8-4ddd-aecb-8e6acf57ac4d · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Gpt-4v(ision) system card
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 688261e9-cdbd-4cb3-a5d7-cfecd01d483d · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Tailor3D: Customized 3D Assets Editing and Generation with Dual-Side Images
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 336d5bca-329a-45bd-a2d9-d90580afc681 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1131b7a0-48d8-4fba-82e4-9a6c75aa0aa3 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Glamm: Pixel grounding large multimodal model
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 102e882d-e025-46db-a6ad-55f00a30f7f8 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e69665ae-af7c-4524-82ba-19e7859f0dd6 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Descriptive caption enhancement with visual specialists for multimodal perception.arXiv preprint arXiv:2412.14233, 2024
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03e39ec1-c5dc-4dcb-a655-5aad110f1693 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9eda35a-64a6-4952-a0ab-2e5a5bfd0135 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa5e912d-7b66-46b2-bc96-acdecc3f72d6 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc5f46d0-dc93-4738-a8ef-aaf49c27b2e0 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eea95389-3355-455a-a68d-066ff04460be · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Common subspace for model and similarity: Phrase learning for caption generation from images
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a97b1593-3422-49d3-82da-ee561ba61b14 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Contrastive region guidance: Improving grounding in vision-language models without training
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 59c74586-ed16-4de0-9da6-2c34d197fbed · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f167734f-7c30-4b59-8533-132b5d573bba · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing VideoRoPE: What Makes for Good Video Rotary Position Embedding?
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 382215c7-67c8-4b50-abca-e0ceaa465502 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64806f49-161e-4d56-aa8e-cb26f5dd936c · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Altogether: Image Captioning via Re-aligning Alt-text
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc18bfe2-8182-4c96-99cf-dd688cd20071 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Woodpecker: Hallucination correction for multimodal large language models.Science China Information Sciences, 67(12):220105, 2024
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1895fc64-31f8-46d5-8074-6069960cf683 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc28bfb-e29b-4a74-93fd-ba4ddd4207ca · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Capsfusion: Rethinking image-text data at scale
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cb9eac3c-d3a8-4456-85a5-82a3f4b5f632 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f2a9247-4cf8-42d9-9435-fe23fb62ed5a · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a49984a-583c-4989-8c47-235ffc566225 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Recognize anything: A strong image tagging model
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 320140f0-88f5-461d-ae07-408a96261f17 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ca300c7-a156-4f3f-affb-1acd58425647 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Describe more details about the [Object]
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b246e20a-7d31-4762-898a-77100b275c04 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Instructions: Describe more details about the oven
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2dea20cf-015d-4cf2-9a37-2921312b64e1 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b5cb4164-54b6-40ad-939c-81202bf559b6 · outbound
ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing Limitations
Reference 256
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 61b793eb-4b1f-449e-a480-548d2d78de90 · inbound
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 909b480f-cc88-4ceb-92a9-a48cf0da6589 · inbound
CODA: Coordinating the Cerebrum and Cerebellum for a Dual-Brain Computer Use Agent with Decoupled Reinforcement Learning ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1857817-b6bf-460c-b2cf-5b9ddea88f20 · inbound
Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06ac6dee-7941-4b8d-b165-2aa7ce6d0158 · inbound
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions ScaleCap: Inference-Time Scalable Image Captioning via Dual-Modality Debiasing
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.