Pith. sign in

Paper Citation Record · LEDGER

Detailed Object Description with Controllable Dimensions

As of 15 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2411.19106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19106 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:35:48.056317Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:15.313911Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:26:16.116871Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact5
  • verified fuzzy35
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01625707-a1fb-4b23-80fe-c7e910ae8ee4 · outbound

This paper cites Describing objects by their attributes,.

Detailed Object Description with Controllable Dimensions Describing objects by their attributes,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.686409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.604403Z digest=sha256:5fe96d38aa4d7350dea147ca6033d039745c838ef69578d64bdd84eda99ea2a3

Observation 98a7afe5-4f63-4af1-9b41-b71994d8ea98 · outbound

This paper cites A new image captioning approach for visually impaired people,.

Detailed Object Description with Controllable Dimensions A new image captioning approach for visually impaired people,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.668889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.610996Z digest=sha256:655c3cbfe8c8c1f70b42468bcb6c74a18a6f4618b8404999d3d76e6617d3886a

Observation f8072616-6ca4-43f8-aca7-d66831197f51 · outbound

This paper cites Visual instruction tuning,.

Detailed Object Description with Controllable Dimensions Visual instruction tuning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.617693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.617693Z digest=sha256:2996fc51c2d15e8dbe6fe8a053f7ca508c24d926a29fed1952fa3ea6ac39dbec

Observation db1c39d6-f2ed-4bbf-9a2f-2d597d61f64d · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Detailed Object Description with Controllable Dimensions InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.624064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.624064Z digest=sha256:3fa4bdfe0bf63c9850a9155d62ee17f4488b39385aeec1479e233edc1b8d3080

Observation 3378b141-db95-40ab-b2a0-12a6ec8f5036 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Detailed Object Description with Controllable Dimensions Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.633004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.633004Z digest=sha256:71a7b1affa8ab2b80417311e565eb2c65762b8d7fa5f51ce3ed685cf4276683b

Observation 3d5890bf-6d27-48b3-ab1c-2e70abc6e4a9 · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity,.

Detailed Object Description with Controllable Dimensions Ferret: Refer and ground anything anywhere at any granularity,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.639549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.638597Z digest=sha256:0ec86689283b60be6f06a1e95003190c560f3261cdc886e7319140f82e450e83

Observation 56c4f11e-d411-4f34-93f5-f0d79587aaa6 · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning,.

Detailed Object Description with Controllable Dimensions Osprey: Pixel understanding with visual instruction tuning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.623325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.645570Z digest=sha256:5a4a000d5ab49df570e95c8298a3084d84cb33365f4653f5939b3ae1fe25bfab

Observation de5b4f10-c62b-4e9a-9b6b-ea1ea97dc266 · outbound

This paper cites Gpt4roi: Instruction tuning large language model on region-of- interest,.

Detailed Object Description with Controllable Dimensions Gpt4roi: Instruction tuning large language model on region-of- interest,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.605571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.652677Z digest=sha256:ad3aaa406f571aa46dab7d5805c7970386d995598f3802d1ad3726f1d1ec0ab7

Observation a909808c-fc55-4a0d-90d8-cb50dcd1cae7 · outbound

This paper cites an unresolved cited work.

Detailed Object Description with Controllable Dimensions Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:35:49.587938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.659736Z digest=sha256:c2b530a1e475704a868b682a3d8a17cde8a820914ab31afe82dfd97174175de7

Observation ba6fc9cf-6e13-4474-8ffa-a4af28239ae7 · outbound

This paper cites GPT-4 Technical Report.

Detailed Object Description with Controllable Dimensions GPT-4 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.665164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.665164Z digest=sha256:5eb27fcd546eeecdd26380b539511b6e327a884bcfc205205c107d32d5b8d46f

Observation 40b2ef8c-f778-4c2e-9326-47cf79e61c84 · outbound

This paper cites Say as you wish: Fine- grained control of image caption generation with abstract scene graphs,.

Detailed Object Description with Controllable Dimensions Say as you wish: Fine- grained control of image caption generation with abstract scene graphs,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.572137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.673001Z digest=sha256:d032dc128ba490af42847345afc7bc97da776c0edddbc8693c0da1b14b9461f6

Observation cb689d64-bb52-4b56-a096-6e6f5a6bb125 · outbound

This paper cites Human-like controllable image captioning with verb-specific semantic roles,.

Detailed Object Description with Controllable Dimensions Human-like controllable image captioning with verb-specific semantic roles,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.552554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.677938Z digest=sha256:328480a0824a08d203cb760273a9b8d96e410c16228d4ade9ad2011b5e452822

Observation 064b0a90-495b-4473-ba0e-3aeea78d620a · outbound

This paper cites Length-Controllable Image Captioning.

Detailed Object Description with Controllable Dimensions Length-Controllable Image Captioning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:35:48.872869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.683228Z digest=sha256:b18154817449fe3052ae18e65243682f96095301f116ff809e986d5aed3f2003

Observation 843e0503-0ea8-4a99-9f54-0194369f9ddc · outbound

This paper cites Image captioning with controllable and adaptive length levels,.

Detailed Object Description with Controllable Dimensions Image captioning with controllable and adaptive length levels,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.534740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.689741Z digest=sha256:a8812de5695efbe5d875faaf4b50f3c6e8bb9b2a1aeeb1a2a8945c0c2c63beb5

Observation 78a9ab64-72f7-4c6a-9eba-bd8a51811d3e · outbound

This paper cites SentiCap: Generating Image Descriptions with Sentiments.

Detailed Object Description with Controllable Dimensions SentiCap: Generating Image Descriptions with Sentiments

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:35:48.843335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.694808Z digest=sha256:acad217fc644121171ead965af06f443fd8f3293ef2976bf6a3727b385d637f0

Observation 353414c4-068d-43bb-90c6-dff0efe6be2d · outbound

This paper cites Alpha-clip: A clip model focusing on wherever you want,.

Detailed Object Description with Controllable Dimensions Alpha-clip: A clip model focusing on wherever you want,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.516778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.700614Z digest=sha256:77f58a4bdff4155ea440a61bfb599d1e97de42946ec4b7f4821364922c70a46e

Observation 17c5457e-12b3-41a9-999d-fe902ba830b8 · outbound

This paper cites Kosmos-2: Grounding multimodal large language models to the world,.

Detailed Object Description with Controllable Dimensions Kosmos-2: Grounding multimodal large language models to the world,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.498047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.705960Z digest=sha256:57989cd7e6d41257c2498417760496520d3c248069573b70467f3c2dd0312d3b

Observation 82254e72-4d11-4a24-9f3b-a45684d2e4ae · outbound

This paper cites Open-vocabulary attribute detection,.

Detailed Object Description with Controllable Dimensions Open-vocabulary attribute detection,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.480672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.710740Z digest=sha256:fa370fa15069b11402961a72b315aec7c964c53c1820486adaf57a56833b3318

Observation 28c062d1-fdf8-4f6f-ba7c-ca472fed1984 · outbound

This paper cites OvarNet: Towards Open-vocabulary Object Attribute Recognition.

Detailed Object Description with Controllable Dimensions OvarNet: Towards Open-vocabulary Object Attribute Recognition

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:35:48.810071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.716653Z digest=sha256:970e2e6ababad366b59d67237f5554f2fc8d014981d0652cb108ea62792286b1

Observation 5e38b85b-607e-4f10-8b20-9054436ded70 · outbound

This paper cites Hierarchical visual attribute learning in the wild,.

Detailed Object Description with Controllable Dimensions Hierarchical visual attribute learning in the wild,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.722495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.722495Z digest=sha256:e52e956c72f7cde94b25fbae09c83c2f5e530c7c65e166b6e5395b17775f548d

Observation fc1fc82d-f697-458a-9b53-36ee71e43d57 · outbound

This paper cites Coco attributes: Attributes for peo- ple, animals, and objects,.

Detailed Object Description with Controllable Dimensions Coco attributes: Attributes for peo- ple, animals, and objects,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.463474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.727724Z digest=sha256:0be4500587c920cc384cd6279ef49eda812123e71b20cd510ba8e9e7a8039f83

Observation 7563b39c-d52d-4f5d-843d-c298fc1ee45e · outbound

This paper cites Reflective decod- ing network for image captioning,.

Detailed Object Description with Controllable Dimensions Reflective decod- ing network for image captioning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.445416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.733902Z digest=sha256:2ddc24d160c438d313cf1898637e27f81a7b2d096a1570fba561f32e436706c1

Observation af887a95-cd45-4343-a65b-eec3ae886a29 · outbound

This paper cites Show, control and tell: A framework for generating controllable and grounded captions,.

Detailed Object Description with Controllable Dimensions Show, control and tell: A framework for generating controllable and grounded captions,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.425986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.739604Z digest=sha256:eb0762cb8e9405ea0d44276025a0ab88c3cb8d5bb354d4e709afdee6b24f83c6

Observation d0997027-8721-4b6a-9a6f-8fd0da4cc149 · outbound

This paper cites Open-set image tagging with multi-grained text supervision,.

Detailed Object Description with Controllable Dimensions Open-set image tagging with multi-grained text supervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.745838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.745838Z digest=sha256:776e4680e6b7d18c0d9b5a573684f91a8b94d7124c7ac2d8d849e9d674320431

Observation 7d5c422c-c1a1-43d1-917e-6618f65fb836 · outbound

This paper cites Recognize Anything: A Strong Image Tagging Model.

Detailed Object Description with Controllable Dimensions Recognize Anything: A Strong Image Tagging Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.752131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.752131Z digest=sha256:ba44c4b2a3055bcd6c350185c7df979e95c7a4af4c21c79ad8f8da00c8fe88fa

Observation 2cea303c-2bea-4d47-b397-696cac09c554 · outbound

This paper cites Tag2Text: Guiding Vision-Language Model via Image Tagging.

Detailed Object Description with Controllable Dimensions Tag2Text: Guiding Vision-Language Model via Image Tagging

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.759661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.759661Z digest=sha256:bb4e5fd86c95e5f135184ae3a163615ccf2819907912ffb6317e92473bd6e5e6

Observation 9675df14-9443-47ca-9e2e-1783d99e4690 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,.

Detailed Object Description with Controllable Dimensions Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.398305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.767206Z digest=sha256:4811752b0afca4b01a776a5674c82ae293d6f8105d5229f7816b63e50095e657

Observation 594f5d7a-b24f-4360-8630-34a301c74082 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models,.

Detailed Object Description with Controllable Dimensions Laion- 5b: An open large-scale dataset for training next generation image-text models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.381669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.773535Z digest=sha256:0bd1d9e0d9b4a2ac7ce976ee2343e80c4d79f7d2b99fbfda834f642574656529

Observation ad253e3f-9fce-4674-a0d4-ff678e1d337e · outbound

This paper cites Docci: Descriptions of connected and contrasting images,.

Detailed Object Description with Controllable Dimensions Docci: Descriptions of connected and contrasting images,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.364637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.778783Z digest=sha256:c6cb93235292a98cd6fb5e1c7217cca593d9db540dd3fcbbe1f0740f2100e7cf

Observation 19b32b7a-0943-4593-8318-435522822cea · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models,.

Detailed Object Description with Controllable Dimensions Dense and aligned captions (dac) promote compositional reasoning in vl models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.346553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.784510Z digest=sha256:f78e0a91a24bf1499c25bf4597f3fec9524e256f62644af8441466ee97671d1d

Observation a94cc381-02df-4ff5-b333-c36e5e55d52f · outbound

This paper cites Pixlore: A dataset-driven approach to rich image caption- ing,.

Detailed Object Description with Controllable Dimensions Pixlore: A dataset-driven approach to rich image caption- ing,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.327577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.789767Z digest=sha256:9311314371b1c2c53190ea8bc7608290f1a6f8cd7830f2b7bf80eaec74a32f26

Observation 13d4cd19-ccc3-41bf-88b0-4a282f552ba7 · outbound

This paper cites SPICE: Semantic Propositional Image Caption Evaluation.

Detailed Object Description with Controllable Dimensions SPICE: Semantic Propositional Image Caption Evaluation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.795510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.795510Z digest=sha256:c68b3e049730d027b0971d962cfbd2e35aa49eeaf38b900b052c550d17e2dcd6

Observation e3d82ea8-504c-413c-ae82-8aeb7f70dc96 · outbound

This paper cites Cdkm: Common and distinct knowledge mining network with content interaction for dense captioning,.

Detailed Object Description with Controllable Dimensions Cdkm: Common and distinct knowledge mining network with content interaction for dense captioning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.310238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.801592Z digest=sha256:7a35827fdf7f1e327e27fb9f0c840a663e4e44dbf63ac22125abf2ab0eae1ba9

Observation 8fdda0a4-519f-4c14-b4dc-2064fc8e02c3 · outbound

This paper cites Icocap: Improving video captioning by compounding images,.

Detailed Object Description with Controllable Dimensions Icocap: Improving video captioning by compounding images,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.291115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.807248Z digest=sha256:a09dd939474b1cbb8f6e33644beb8432681cf2574ed802ecf1dceae743afa7dc

Observation 5f221a8c-83e2-4ce4-8d64-a6894a515d8b · outbound

This paper cites Fine-grained image captioning with global-local discriminative objective,.

Detailed Object Description with Controllable Dimensions Fine-grained image captioning with global-local discriminative objective,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.274636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.813906Z digest=sha256:0210d878cab5ad22ee8ddacd005d6d091442aa7a3fe95c232e068db4af8b285d

Observation 3a4e2ebd-d14d-477b-9362-7c56e870d7ad · outbound

This paper cites Mul- titask learning for cross-domain image captioning,.

Detailed Object Description with Controllable Dimensions Mul- titask learning for cross-domain image captioning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.256275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.820027Z digest=sha256:1ddcba02bdc405d2150ec9b325a9478aeef873fb41eda75ee499ddc29f7687f7

Observation 4db71749-ac45-4575-9548-0e84ad3439c6 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Detailed Object Description with Controllable Dimensions Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.825939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.825939Z digest=sha256:0115fde854ac676910e003b058a70bce855e795f2e063f28ab15e716c41378d1

Observation 40bc00fa-2b03-46d7-a8f4-5543bde6e78d · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions.

Detailed Object Description with Controllable Dimensions ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.832245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.832245Z digest=sha256:1a3475aa10dc6d7fe778f20595f3d42c75edc4a26b76f4dd6fd23d80504a921b

Observation 952c4115-cedb-43a8-a491-a31df8724bf0 · outbound

This paper cites Boosting entity-aware image captioning with multi-modal knowledge graph,.

Detailed Object Description with Controllable Dimensions Boosting entity-aware image captioning with multi-modal knowledge graph,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.227154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.837917Z digest=sha256:4ab14c0a311217b76ab4a215e474fb6f6648fdb182014409900505656e5cb091

Observation 47b3914d-8339-4831-905d-14e8ededf831 · outbound

This paper cites Describing like humans: On diversity in image captioning,.

Detailed Object Description with Controllable Dimensions Describing like humans: On diversity in image captioning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.209834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.843340Z digest=sha256:88fcd0cd5e3101e8765660444d846ae0e12942aec158295642ba7568e6ffd764

Observation 9462c86b-b03c-4ed8-834c-07fdad930a34 · outbound

This paper cites Benchmarking complex instruction-following with multiple constraints composition,.

Detailed Object Description with Controllable Dimensions Benchmarking complex instruction-following with multiple constraints composition,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.849028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.849028Z digest=sha256:d2742ca07d6ce5820e74898b21b61bd68de667780033b668f163e149f4cf4352

Observation d6c24d2b-b081-401b-961d-5806342bf03d · outbound

This paper cites Densecap: Fully convolutional localization networks for dense captioning,.

Detailed Object Description with Controllable Dimensions Densecap: Fully convolutional localization networks for dense captioning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.180368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.862137Z digest=sha256:9eef0464af60bc36d2930780692609e19f7fa362ae8c386e37a484cafab95859

Observation b26ba455-6353-4ab9-b9bb-8492f294d36b · outbound

This paper cites Controlcap: Controllable region-level captioning,.

Detailed Object Description with Controllable Dimensions Controlcap: Controllable region-level captioning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.163671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.867248Z digest=sha256:99a3e5907ef35eb21e1deb5c4f190e4c17d5ba5c6d15a7cb9ec0c85b667638a9

Observation 7dc12e9f-7bb0-41b4-b2cb-279501b2d840 · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

Detailed Object Description with Controllable Dimensions OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.874399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.874399Z digest=sha256:3ea98b9e8d760ac0fddc46620e9582f7d9014ae55f8155446cd15f3c172b2063

Observation 9d266fd7-6f17-43ec-aa48-8eafa2903e2a · outbound

This paper cites Prompt Highlighter: Interactive Control for Multi-Modal LLMs.

Detailed Object Description with Controllable Dimensions Prompt Highlighter: Interactive Control for Multi-Modal LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.884042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.884042Z digest=sha256:4c829482508eb9d7a21583621a4f95569df7876a189cd2db0246f4fe99ba3cb9

Observation 241b12de-afc9-4a69-8aa0-7c188425219a · outbound

This paper cites FlexCap: Describe Anything in Images in Controllable Detail.

Detailed Object Description with Controllable Dimensions FlexCap: Describe Anything in Images in Controllable Detail

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:35:48.419082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.890020Z digest=sha256:b7e1a31784f4b52a2b125fc332da36fcd8e1d705881b33069c7c15533a485348

Observation a17ba614-9c42-406a-be2a-490781254861 · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

Detailed Object Description with Controllable Dimensions Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.895609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.895609Z digest=sha256:56a389d4d8c20e200517db9ae98a43e4f98ab50834d22afb6eecc1e3b851531b

Observation 6b999e30-97e3-4414-ab62-820aaafc0919 · outbound

This paper cites Debiasing pretrained generative models by uniformly sampling semantic attributes,.

Detailed Object Description with Controllable Dimensions Debiasing pretrained generative models by uniformly sampling semantic attributes,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.144490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.902410Z digest=sha256:888598553b9b5fd5ff63563ecb4b98f5f11faf9bf485dcc3c4d1cea41d2ebe41

Observation 971ced30-02e6-434f-98ec-559cd1f75ba9 · outbound

This paper cites Unifying visual attribute learning with object recognition in a multiplicative framework,.

Detailed Object Description with Controllable Dimensions Unifying visual attribute learning with object recognition in a multiplicative framework,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.125290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.908845Z digest=sha256:8b52c387bfaddcfe59ad7eddbe53df3f25dc86bdb92fea12f864382d75e3056b

Observation f3c2e607-1ce7-49fd-825d-db348c506a6e · outbound

This paper cites Learning to parameterize visual attributes for open-set fine-grained retrieval,.

Detailed Object Description with Controllable Dimensions Learning to parameterize visual attributes for open-set fine-grained retrieval,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.104239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.914817Z digest=sha256:ba05780fe696dd96afa4194f2670355086e2f87da5cd4c93116a2119dae66abc

Observation 2f5b3f3b-7c3f-4c2e-a7ed-d661bd9b06ba · outbound

This paper cites Benchmarking Segmentation Models with Mask-Preserved Attribute Editing.

Detailed Object Description with Controllable Dimensions Benchmarking Segmentation Models with Mask-Preserved Attribute Editing

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:35:48.354732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.920039Z digest=sha256:3ef1772ab767e126b9f3aee7f9153fe79407f8f15cda38c0ce06b77c0d91048c

Observation d0bf7c53-756a-426f-81a1-f75c003871b1 · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering,.

Detailed Object Description with Controllable Dimensions Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.083310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.927030Z digest=sha256:0e78baeed324ab30f9d2bf4cb16eb4b8ea3e98f890ec55dc5eec22d19a3c78a7

Observation 5855857a-2d3e-48db-b118-5ffda950003e · outbound

This paper cites MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs.

Detailed Object Description with Controllable Dimensions MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.933082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.933082Z digest=sha256:f90e2e965af06ca1f34b7406cf41171ecb566e8e3ddce206eb6d680efc9d354f

Observation 5b524447-7507-44f3-be44-13baf064773e · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models,.

Detailed Object Description with Controllable Dimensions Mme: A comprehensive evaluation benchmark for multimodal large language models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.062367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.942139Z digest=sha256:24e23d14aebe269937b22fb77c2a8aabe52cf9cbfe5b57f1cab296fefaea3385

Observation 7ef05d19-4bdf-4ba8-a320-65cb317a3dd5 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Detailed Object Description with Controllable Dimensions SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.949483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.949483Z digest=sha256:068037e79e9a6f138b771863b5974db179e66c1171100759a6241f3c101b40c5

Observation a54a24cf-6522-4bd2-83e7-d1116c974f2b · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

Detailed Object Description with Controllable Dimensions SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.960523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.960523Z digest=sha256:b9d478437e4c49b3dd172ed9afa8b952d0dfed871068767f474b8e0cdcfce189

Observation 18e991f0-6e74-440d-ac84-0659174b7709 · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

Detailed Object Description with Controllable Dimensions SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.966757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.966757Z digest=sha256:461e4b14e8e65478a91c2af556f07262aec8d668977b320a4a1467f43c2450e3

Observation 0f595b07-6d6b-471b-a02a-5b098af6f983 · outbound

This paper cites Woodpecker: Hallucination Correction for Multimodal Large Language Models.

Detailed Object Description with Controllable Dimensions Woodpecker: Hallucination Correction for Multimodal Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.975805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.975805Z digest=sha256:f609363df00cdae69f457cd5ae74dbf53a2c0719f4e8c78afaed03736c68eec1

Observation c992c0a2-7dc5-421f-ac95-62ad52703d06 · outbound

This paper cites FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models.

Detailed Object Description with Controllable Dimensions FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.985941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.985941Z digest=sha256:88d4b3e237074490a8f1baf11d2d1df4995db11d0ec3c07d40478aadec2bd171

Observation 514c7eca-6f46-40c8-bfdc-8f0243fb3bfa · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Detailed Object Description with Controllable Dimensions Evaluating Object Hallucination in Large Vision-Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.992435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.992435Z digest=sha256:65725cc90a40ead51d856ade5216192f67d3b27796eab3684d84b3de81204a89

Observation 73ef9f90-9d6e-4db9-99a7-1e95f222cbb7 · outbound

This paper cites Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation,.

Detailed Object Description with Controllable Dimensions Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.042580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:47.999331Z digest=sha256:c8c9e4d717261e3783871860d7c9beaa1032b0db0638bc6629139ada73050359

Observation 3cb156dc-3df5-43f7-993b-0464db26413a · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Detailed Object Description with Controllable Dimensions Bleu: a method for automatic evaluation of machine translation,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.024308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:48.005430Z digest=sha256:6c1c54c64360a52556a537cf6a1a0fc9664f4489315659d6d8cb4c6fccef611c

Observation 3e2b5125-5d48-40b7-8675-7ba40dd47cde · outbound

This paper cites METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,.

Detailed Object Description with Controllable Dimensions METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.005724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T10:35:48.011868Z digest=sha256:d8ecd7f13ebf753b0697d3178e049f0172f523cce219e5f5613ac22b560de85b

Observation 20e61f48-e320-4dad-b4c3-44551232a430 · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation.

Detailed Object Description with Controllable Dimensions CIDEr: Consensus-based Image Description Evaluation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.017689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.017689Z digest=sha256:0954da4c0852eec6afdc1f290a75d724bc9b0d1274a24bfd1dbac40938824b1d

Observation f3332ba3-5785-42de-b973-7e501621c279 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Detailed Object Description with Controllable Dimensions CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.024015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.024015Z digest=sha256:9e3aaf61f9c95f3a7ca7084c623a6ca8f80e220c6c8b5f032caeb10f41fecf75

Observation a0577e74-aa53-4d04-8388-75a6a5aa47dc · outbound

This paper cites Improved baselines with visual instruction tuning,.

Detailed Object Description with Controllable Dimensions Improved baselines with visual instruction tuning,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.030002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.030002Z digest=sha256:49c137ad13a4179e9a89e4e7d7a9d779f0f995b416f1ee874a08fc171943449a

Observation c1152873-3296-4039-8d8f-dd6aee64ba02 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt- 4 with 90%* chatgpt quality,.

Detailed Object Description with Controllable Dimensions Vicuna: An open-source chatbot impressing gpt- 4 with 90%* chatgpt quality,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.039061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.039061Z digest=sha256:673d07d8fd0221b50c414fc56046d792137da2d338ca8363833964b9bd572beb

Observation a779451d-4b33-46cf-8b22-2473c7e98cfa · outbound

This paper cites Llama 3 model card,.

Detailed Object Description with Controllable Dimensions Llama 3 model card,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.046825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.046825Z digest=sha256:edc8bb2bfe5ee991ade4bae684a8944272f3d9547fb06ed2f6b50b799845a398

Observation bf3e4c25-36dc-4702-a277-c20a1ec407f8 · outbound

This paper cites Microsoft coco: Common objects in context,.

Detailed Object Description with Controllable Dimensions Microsoft coco: Common objects in context,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.056317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.056317Z digest=sha256:6626df225ce01da86c6bc568eeae7f851a18bd83a1e6c26851e8c9d70769ed67

Observation 64d73cef-163e-4ba0-ab26-dc194996979f · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

Detailed Object Description with Controllable Dimensions Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.855603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.855603Z digest=sha256:2a9d9583f1580ab2100e91b8c609a6ce9fa61ee4179fd0087eafb327641153b9

Pith citing papers

Observation 144dfe14-17ee-4373-af5d-381632a1b5ca · inbound

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation cites this paper.

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation Detailed Object Description with Controllable Dimensions

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:26:16.179513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T15:26:15.313911Z digest=sha256:5bc75b1d8158a3af6316b01af513e506cc7a06edd1c0200fd25606c6b78335c3