Pith. sign in

Paper Citation Record · LEDGER

Detailed Object Description with Controllable Dimensions

As of 20 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2411.19106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19106 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:35:48.056317Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:15.313911Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:26:16.116871Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact5
  • verified fuzzy35
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 01625707-a1fb-4b23-80fe-c7e910ae8ee4 · outbound

This paper cites Describing objects by their attributes,.

Detailed Object Description with Controllable Dimensions Describing objects by their attributes,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.686409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.604403Z digest=sha256:7e51dd43350b1781b47a857fad68c5a6dea666b2b6d0ff9249413e9d0f87c206

Observation 98a7afe5-4f63-4af1-9b41-b71994d8ea98 · outbound

This paper cites A new image captioning approach for visually impaired people,.

Detailed Object Description with Controllable Dimensions A new image captioning approach for visually impaired people,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.668889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.610996Z digest=sha256:0cbff92c2e478c808bf0a4a00988d6cccd91e8f027d86f0190c995f4e360a99e

Observation f8072616-6ca4-43f8-aca7-d66831197f51 · outbound

This paper cites Visual instruction tuning,.

Detailed Object Description with Controllable Dimensions Visual instruction tuning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.617693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.617693Z digest=sha256:5b5228ac5c59c345654ff21eebcc1663ad7888ef0c6ffc754a2fd37f448411cb

Observation db1c39d6-f2ed-4bbf-9a2f-2d597d61f64d · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Detailed Object Description with Controllable Dimensions InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.624064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.624064Z digest=sha256:113784d2f8cb6d7c0f1261c0f667c91fd54e24da2d7361c4c7573b0b614fed8a

Observation 3378b141-db95-40ab-b2a0-12a6ec8f5036 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Detailed Object Description with Controllable Dimensions Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.633004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.633004Z digest=sha256:3859f24547ed12e6b8a2ef784e8b8bb709d3647487328ab92d98f6eb88391930

Observation 3d5890bf-6d27-48b3-ab1c-2e70abc6e4a9 · outbound

This paper cites Ferret: Refer and ground anything anywhere at any granularity,.

Detailed Object Description with Controllable Dimensions Ferret: Refer and ground anything anywhere at any granularity,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.639549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.638597Z digest=sha256:0dfe0ca0827d2da812e20367e31443ef1f0a921c46085d79f23ca3afb91d6a54

Observation 56c4f11e-d411-4f34-93f5-f0d79587aaa6 · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning,.

Detailed Object Description with Controllable Dimensions Osprey: Pixel understanding with visual instruction tuning,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.623325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.645570Z digest=sha256:60c1a15a688582498a37023fba1c0e064a9b260ad22f69d13139db739d90b34f

Observation de5b4f10-c62b-4e9a-9b6b-ea1ea97dc266 · outbound

This paper cites Gpt4roi: Instruction tuning large language model on region-of- interest,.

Detailed Object Description with Controllable Dimensions Gpt4roi: Instruction tuning large language model on region-of- interest,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.605571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.652677Z digest=sha256:fef6d67ef35569f90556a8ade7bea87a83c79a3a97d1ef43555d71f6a917d9f6

Observation a909808c-fc55-4a0d-90d8-cb50dcd1cae7 · outbound

This paper cites an unresolved cited work.

Detailed Object Description with Controllable Dimensions Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-12T10:35:49.587938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.659736Z digest=sha256:90e0199c3e0bc1b20df5c56e9c67f912b077f614343e063ea3c7224f61ab48c8

Observation ba6fc9cf-6e13-4474-8ffa-a4af28239ae7 · outbound

This paper cites GPT-4 Technical Report.

Detailed Object Description with Controllable Dimensions GPT-4 Technical Report

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.665164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.665164Z digest=sha256:e65a80740999041571ef90660483268ae123f8a00f1c22cb0cf48834058879d8

Observation 40b2ef8c-f778-4c2e-9326-47cf79e61c84 · outbound

This paper cites Say as you wish: Fine- grained control of image caption generation with abstract scene graphs,.

Detailed Object Description with Controllable Dimensions Say as you wish: Fine- grained control of image caption generation with abstract scene graphs,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.572137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.673001Z digest=sha256:16aff21057f7a021af0ff96c6696f02b53a27d5f36911e9f9633070e90193c5c

Observation cb689d64-bb52-4b56-a096-6e6f5a6bb125 · outbound

This paper cites Human-like controllable image captioning with verb-specific semantic roles,.

Detailed Object Description with Controllable Dimensions Human-like controllable image captioning with verb-specific semantic roles,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.552554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.677938Z digest=sha256:94afa11a3b47dae721df09730695528a73311d83adb204e86ee8360b2937ab96

Observation 064b0a90-495b-4473-ba0e-3aeea78d620a · outbound

This paper cites Length-Controllable Image Captioning.

Detailed Object Description with Controllable Dimensions Length-Controllable Image Captioning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:35:48.872869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.683228Z digest=sha256:c64aa0786ef8352137916ac01eff7f6b1e0a0d31bc769e2f9bf59842a1f71afb

Observation 843e0503-0ea8-4a99-9f54-0194369f9ddc · outbound

This paper cites Image captioning with controllable and adaptive length levels,.

Detailed Object Description with Controllable Dimensions Image captioning with controllable and adaptive length levels,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.534740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.689741Z digest=sha256:4acf74032719804e370c010347ed8e388cf4a75eb05d6cadee070bba7391a94a

Observation 78a9ab64-72f7-4c6a-9eba-bd8a51811d3e · outbound

This paper cites SentiCap: Generating Image Descriptions with Sentiments.

Detailed Object Description with Controllable Dimensions SentiCap: Generating Image Descriptions with Sentiments

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:35:48.843335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.694808Z digest=sha256:080be88c708d4ef00f7acabb00696d18666404e71eeffede9ed3503455ba081b

Observation 353414c4-068d-43bb-90c6-dff0efe6be2d · outbound

This paper cites Alpha-clip: A clip model focusing on wherever you want,.

Detailed Object Description with Controllable Dimensions Alpha-clip: A clip model focusing on wherever you want,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.516778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.700614Z digest=sha256:85cd3b4090d9dcaeb2358b5498e058a0d82a1b43c8256e1ff8d548b448236661

Observation 17c5457e-12b3-41a9-999d-fe902ba830b8 · outbound

This paper cites Kosmos-2: Grounding multimodal large language models to the world,.

Detailed Object Description with Controllable Dimensions Kosmos-2: Grounding multimodal large language models to the world,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.498047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.705960Z digest=sha256:f5227b3fd1792fcb35d2eb98c2cc3f0ea92e14f43de70636c6a48bde8ca83db7

Observation 82254e72-4d11-4a24-9f3b-a45684d2e4ae · outbound

This paper cites Open-vocabulary attribute detection,.

Detailed Object Description with Controllable Dimensions Open-vocabulary attribute detection,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.480672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.710740Z digest=sha256:1867737e067d5d35daafb766ea1486353ed32f60471052550a707d04b76e2d9c

Observation 28c062d1-fdf8-4f6f-ba7c-ca472fed1984 · outbound

This paper cites OvarNet: Towards Open-vocabulary Object Attribute Recognition.

Detailed Object Description with Controllable Dimensions OvarNet: Towards Open-vocabulary Object Attribute Recognition

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:35:48.810071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.716653Z digest=sha256:c3ae5e6db94472937e5495e7cdca6aaf7be2315a5c03db4086325450ff04f62e

Observation 5e38b85b-607e-4f10-8b20-9054436ded70 · outbound

This paper cites Hierarchical visual attribute learning in the wild,.

Detailed Object Description with Controllable Dimensions Hierarchical visual attribute learning in the wild,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.722495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.722495Z digest=sha256:6cad648b45cc5dfcfddcb73dd7711a63cdb930c4a94e2c373654336464e13344

Observation fc1fc82d-f697-458a-9b53-36ee71e43d57 · outbound

This paper cites Coco attributes: Attributes for peo- ple, animals, and objects,.

Detailed Object Description with Controllable Dimensions Coco attributes: Attributes for peo- ple, animals, and objects,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.463474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.727724Z digest=sha256:6791f8a712d8d1322d4e329e186649b6f606ea60dd2f87c01561cc0f3aeb66fc

Observation 7563b39c-d52d-4f5d-843d-c298fc1ee45e · outbound

This paper cites Reflective decod- ing network for image captioning,.

Detailed Object Description with Controllable Dimensions Reflective decod- ing network for image captioning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.445416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.733902Z digest=sha256:81e6cebae2d8f33142e32fdb10fd890ea224a832b0f2769a10c654411e3f60a0

Observation af887a95-cd45-4343-a65b-eec3ae886a29 · outbound

This paper cites Show, control and tell: A framework for generating controllable and grounded captions,.

Detailed Object Description with Controllable Dimensions Show, control and tell: A framework for generating controllable and grounded captions,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.425986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.739604Z digest=sha256:df4f6928f97842d85ed585251405c08dc4e746de09f888e909bfb1481d50a693

Observation d0997027-8721-4b6a-9a6f-8fd0da4cc149 · outbound

This paper cites Open-set image tagging with multi-grained text supervision,.

Detailed Object Description with Controllable Dimensions Open-set image tagging with multi-grained text supervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.745838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.745838Z digest=sha256:4ab46096b1ff7e32df58e3a25aa66c100487f127dfb20fc281fc1eee03f99f28

Observation 7d5c422c-c1a1-43d1-917e-6618f65fb836 · outbound

This paper cites Recognize Anything: A Strong Image Tagging Model.

Detailed Object Description with Controllable Dimensions Recognize Anything: A Strong Image Tagging Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.752131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.752131Z digest=sha256:ad014b7791922d7109aa0c997f83554f73f7ff07d16a6974360c81504d986bcc

Observation 2cea303c-2bea-4d47-b397-696cac09c554 · outbound

This paper cites Tag2Text: Guiding Vision-Language Model via Image Tagging.

Detailed Object Description with Controllable Dimensions Tag2Text: Guiding Vision-Language Model via Image Tagging

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.759661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.759661Z digest=sha256:952c273db9321f6f4e613679707ccabf530e486acf28afca58c32c7b7ecd73be

Observation 9675df14-9443-47ca-9e2e-1783d99e4690 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,.

Detailed Object Description with Controllable Dimensions Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.398305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.767206Z digest=sha256:085c482de48ba523da0685392a951b0d503cffc98697758bc3ae4d1570742dc1

Observation 594f5d7a-b24f-4360-8630-34a301c74082 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models,.

Detailed Object Description with Controllable Dimensions Laion- 5b: An open large-scale dataset for training next generation image-text models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.381669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.773535Z digest=sha256:a8f6cf09d6bc9260e28bdfdd99801378c0fd650ad792f6fb375e3a786c861bbf

Observation ad253e3f-9fce-4674-a0d4-ff678e1d337e · outbound

This paper cites Docci: Descriptions of connected and contrasting images,.

Detailed Object Description with Controllable Dimensions Docci: Descriptions of connected and contrasting images,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.364637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.778783Z digest=sha256:acf8975275c43cdab2a5813465f4464524f98d61dfb4ba0d5190f59d2d5d87e7

Observation 19b32b7a-0943-4593-8318-435522822cea · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models,.

Detailed Object Description with Controllable Dimensions Dense and aligned captions (dac) promote compositional reasoning in vl models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.346553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.784510Z digest=sha256:e35a3f9571d6a0a0f81594fd491aaf6ee14fdbc6e6a40e2069a843f84abf7494

Observation a94cc381-02df-4ff5-b333-c36e5e55d52f · outbound

This paper cites Pixlore: A dataset-driven approach to rich image caption- ing,.

Detailed Object Description with Controllable Dimensions Pixlore: A dataset-driven approach to rich image caption- ing,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.327577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.789767Z digest=sha256:f47fcc732b6c022bbd59167b8ef4ccddd436887dc3de17ff90ee6849f438bf35

Observation 13d4cd19-ccc3-41bf-88b0-4a282f552ba7 · outbound

This paper cites SPICE: Semantic Propositional Image Caption Evaluation.

Detailed Object Description with Controllable Dimensions SPICE: Semantic Propositional Image Caption Evaluation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.795510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.795510Z digest=sha256:a5b0cebbda524110c63e936108c3ca8e5915e9ddd7984b4e24a2c694fe75ce25

Observation e3d82ea8-504c-413c-ae82-8aeb7f70dc96 · outbound

This paper cites Cdkm: Common and distinct knowledge mining network with content interaction for dense captioning,.

Detailed Object Description with Controllable Dimensions Cdkm: Common and distinct knowledge mining network with content interaction for dense captioning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.310238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.801592Z digest=sha256:6e0a3042245292c425996758a909862a3325b07bc2914733519eef0a4e7e0af2

Observation 8fdda0a4-519f-4c14-b4dc-2064fc8e02c3 · outbound

This paper cites Icocap: Improving video captioning by compounding images,.

Detailed Object Description with Controllable Dimensions Icocap: Improving video captioning by compounding images,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.291115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.807248Z digest=sha256:f0418c327aad9f4d87eb4f3c8ceae5d0c7d69faf0f090a50c8526b2ae98ca714

Observation 5f221a8c-83e2-4ce4-8d64-a6894a515d8b · outbound

This paper cites Fine-grained image captioning with global-local discriminative objective,.

Detailed Object Description with Controllable Dimensions Fine-grained image captioning with global-local discriminative objective,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.274636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.813906Z digest=sha256:5077ea5f03626abfc518ce401e5b6201b02e184e091d9b7bb410012c51be473b

Observation 3a4e2ebd-d14d-477b-9362-7c56e870d7ad · outbound

This paper cites Mul- titask learning for cross-domain image captioning,.

Detailed Object Description with Controllable Dimensions Mul- titask learning for cross-domain image captioning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.256275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.820027Z digest=sha256:aadd9951e8a7194d67d974314e2766c01b288bda5e71f32aa7fbcfd2629388b8

Observation 4db71749-ac45-4575-9548-0e84ad3439c6 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Detailed Object Description with Controllable Dimensions Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.825939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.825939Z digest=sha256:c95801d7de1f634debd3b23a9d5610d24ad799b97c7e6645499a81fd877ed1e1

Observation 40bc00fa-2b03-46d7-a8f4-5543bde6e78d · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions.

Detailed Object Description with Controllable Dimensions ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.832245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.832245Z digest=sha256:2543da59ea91098dd44f1b407538a75c513527dcdfc401ddd3aee69ce914e244

Observation 952c4115-cedb-43a8-a491-a31df8724bf0 · outbound

This paper cites Boosting entity-aware image captioning with multi-modal knowledge graph,.

Detailed Object Description with Controllable Dimensions Boosting entity-aware image captioning with multi-modal knowledge graph,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.227154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.837917Z digest=sha256:bd5969fca3017ad17ac6c566d6c139711d3f3ebe85c65721521448d3f8426287

Observation 47b3914d-8339-4831-905d-14e8ededf831 · outbound

This paper cites Describing like humans: On diversity in image captioning,.

Detailed Object Description with Controllable Dimensions Describing like humans: On diversity in image captioning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.209834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.843340Z digest=sha256:c5fcfc054a699a75aea768f517c157e53418d5d011e1ec894515cbbfae538d8d

Observation 9462c86b-b03c-4ed8-834c-07fdad930a34 · outbound

This paper cites Benchmarking complex instruction-following with multiple constraints composition,.

Detailed Object Description with Controllable Dimensions Benchmarking complex instruction-following with multiple constraints composition,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.849028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.849028Z digest=sha256:a9805b0d1c3f038610fe999fb5928d98dd792d5e0fcda9e747890ac4b2a81b09

Observation d6c24d2b-b081-401b-961d-5806342bf03d · outbound

This paper cites Densecap: Fully convolutional localization networks for dense captioning,.

Detailed Object Description with Controllable Dimensions Densecap: Fully convolutional localization networks for dense captioning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.180368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.862137Z digest=sha256:bafd13c55f80e8e151a33e9c6ebe9899522d9037bc97913ca77afa8495fbc9dd

Observation b26ba455-6353-4ab9-b9bb-8492f294d36b · outbound

This paper cites Controlcap: Controllable region-level captioning,.

Detailed Object Description with Controllable Dimensions Controlcap: Controllable region-level captioning,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.163671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.867248Z digest=sha256:9e71dd2d9cbfc1793d0e44ec22fb05709750dd257f06ad41b30bb4ada6dbc423

Observation 7dc12e9f-7bb0-41b4-b2cb-279501b2d840 · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

Detailed Object Description with Controllable Dimensions OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.874399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.874399Z digest=sha256:6b1e2cd64cc5a1bc54dd1ad1aaa06e51ced92d559eaf949a98aba79571b807dc

Observation 9d266fd7-6f17-43ec-aa48-8eafa2903e2a · outbound

This paper cites Prompt Highlighter: Interactive Control for Multi-Modal LLMs.

Detailed Object Description with Controllable Dimensions Prompt Highlighter: Interactive Control for Multi-Modal LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.884042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.884042Z digest=sha256:dbc7f1210e241126bf456395ff5c5c257ffa64778c07967a10de393174c5a622

Observation 241b12de-afc9-4a69-8aa0-7c188425219a · outbound

This paper cites FlexCap: Describe Anything in Images in Controllable Detail.

Detailed Object Description with Controllable Dimensions FlexCap: Describe Anything in Images in Controllable Detail

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:35:48.419082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.890020Z digest=sha256:44bf5e1616b3296bfd2cfce89b91de212540b228a0620ec73d1c9105c093062a

Observation a17ba614-9c42-406a-be2a-490781254861 · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

Detailed Object Description with Controllable Dimensions Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.895609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.895609Z digest=sha256:5bab3b2a2ba274ee5677ad9ce53ad132a95a9f9de7917a3843bbc36419224a6b

Observation 6b999e30-97e3-4414-ab62-820aaafc0919 · outbound

This paper cites Debiasing pretrained generative models by uniformly sampling semantic attributes,.

Detailed Object Description with Controllable Dimensions Debiasing pretrained generative models by uniformly sampling semantic attributes,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.144490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.902410Z digest=sha256:2992ee1354c52142c36ac9192ca92a17e17b4e08711116de78eeb62bfba40abd

Observation 971ced30-02e6-434f-98ec-559cd1f75ba9 · outbound

This paper cites Unifying visual attribute learning with object recognition in a multiplicative framework,.

Detailed Object Description with Controllable Dimensions Unifying visual attribute learning with object recognition in a multiplicative framework,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.125290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.908845Z digest=sha256:39a3c65edcfbd5cadda353cdcc25ed1794afbab0d07d731bee60f110a1516dfb

Observation f3c2e607-1ce7-49fd-825d-db348c506a6e · outbound

This paper cites Learning to parameterize visual attributes for open-set fine-grained retrieval,.

Detailed Object Description with Controllable Dimensions Learning to parameterize visual attributes for open-set fine-grained retrieval,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.104239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.914817Z digest=sha256:9ab12d18106290dc4258c8440a85c9a0f170db36e1da8993b314bf9b55908036

Observation 2f5b3f3b-7c3f-4c2e-a7ed-d661bd9b06ba · outbound

This paper cites Benchmarking Segmentation Models with Mask-Preserved Attribute Editing.

Detailed Object Description with Controllable Dimensions Benchmarking Segmentation Models with Mask-Preserved Attribute Editing

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-12T10:35:48.354732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.920039Z digest=sha256:d197b736068588ca5ac447f7f9e41fa6d933830e622c7658d4cc2ab11e13405d

Observation d0bf7c53-756a-426f-81a1-f75c003871b1 · outbound

This paper cites Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering,.

Detailed Object Description with Controllable Dimensions Tifa: Accurate and interpretable text-to-image faithfulness evaluation with question answering,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.083310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.927030Z digest=sha256:8df16b99bad20c165051c1f19788a8eb688746f4c7dc4448c99ab218367da0e5

Observation 5855857a-2d3e-48db-b118-5ffda950003e · outbound

This paper cites MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs.

Detailed Object Description with Controllable Dimensions MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.933082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.933082Z digest=sha256:3842cd918880e45ba1964ab645d4ed9cf7d341f5881ca2d9d7b87537bc15359c

Observation 5b524447-7507-44f3-be44-13baf064773e · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models,.

Detailed Object Description with Controllable Dimensions Mme: A comprehensive evaluation benchmark for multimodal large language models,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.062367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.942139Z digest=sha256:2768762030f9f5d29f3366020bf6ecae6dfd3b38639a63b2184d1ae855924331

Observation 7ef05d19-4bdf-4ba8-a320-65cb317a3dd5 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Detailed Object Description with Controllable Dimensions SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.949483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.949483Z digest=sha256:5ec0a48ba36511053a07f6f19608e4678b006631db9060860b189f5afac97d5e

Observation a54a24cf-6522-4bd2-83e7-d1116c974f2b · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

Detailed Object Description with Controllable Dimensions SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.960523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.960523Z digest=sha256:e62f1d7ec2d4678710f4b8680ad191fb60006d1f40422ebd730ae0da65e78d29

Observation 18e991f0-6e74-440d-ac84-0659174b7709 · outbound

This paper cites SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension.

Detailed Object Description with Controllable Dimensions SEED-Bench-2-Plus: Benchmarking Multimodal Large Language Models with Text-Rich Visual Comprehension

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.966757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.966757Z digest=sha256:49cc504ea4f88febd31059f0d3c385e705d43ae543beddfb95f1a8b08d594943

Observation 0f595b07-6d6b-471b-a02a-5b098af6f983 · outbound

This paper cites Woodpecker: Hallucination Correction for Multimodal Large Language Models.

Detailed Object Description with Controllable Dimensions Woodpecker: Hallucination Correction for Multimodal Large Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.975805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.975805Z digest=sha256:2e96486af513bce22cad2dd65e07bcb80218fc7ad922dfacdd9f1c150cd49645

Observation c992c0a2-7dc5-421f-ac95-62ad52703d06 · outbound

This paper cites FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models.

Detailed Object Description with Controllable Dimensions FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.985941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.985941Z digest=sha256:10a981fe5c36d4e7ba2e650545c99e05890d802bdab6f750317dd583d93348e0

Observation 514c7eca-6f46-40c8-bfdc-8f0243fb3bfa · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Detailed Object Description with Controllable Dimensions Evaluating Object Hallucination in Large Vision-Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.992435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.992435Z digest=sha256:f991468e48f64afc9768ae5ae2b11f2e643cefa4de5a12f150cf9564fa71da32

Observation 73ef9f90-9d6e-4db9-99a7-1e95f222cbb7 · outbound

This paper cites Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation,.

Detailed Object Description with Controllable Dimensions Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.042580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:47.999331Z digest=sha256:dc0728d1d1e12638ef3810b45efe1f389e7908ee234f116243f8df140283cb31

Observation 3cb156dc-3df5-43f7-993b-0464db26413a · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

Detailed Object Description with Controllable Dimensions Bleu: a method for automatic evaluation of machine translation,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.024308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:48.005430Z digest=sha256:1c7c8a061908d8f53a5bc359d19f6d0811cf737df36307349ae3794ebf29dd20

Observation 3e2b5125-5d48-40b7-8675-7ba40dd47cde · outbound

This paper cites METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,.

Detailed Object Description with Controllable Dimensions METEOR: An automatic metric for MT evaluation with improved correlation with human judgments,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:35:49.005724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T10:35:48.011868Z digest=sha256:a64d79709b9d1d45f994ea0f2294a42286613fcf752c0f400cb71515598ce575

Observation 20e61f48-e320-4dad-b4c3-44551232a430 · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation.

Detailed Object Description with Controllable Dimensions CIDEr: Consensus-based Image Description Evaluation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.017689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.017689Z digest=sha256:5dd5687ef6987516b5e30e5dd02c43bbbfa09ae752679d08d98dad2b2c677cce

Observation f3332ba3-5785-42de-b973-7e501621c279 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Detailed Object Description with Controllable Dimensions CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.024015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.024015Z digest=sha256:7e21dde50724ef288c5a9d86589d0cce373b9b5f431387d90728169f9a24c8be

Observation a0577e74-aa53-4d04-8388-75a6a5aa47dc · outbound

This paper cites Improved baselines with visual instruction tuning,.

Detailed Object Description with Controllable Dimensions Improved baselines with visual instruction tuning,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.030002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.030002Z digest=sha256:f93e0281a33baefbc6aed3340b3ae073418fb96613a3cb5a66f7dfa8026b56e0

Observation c1152873-3296-4039-8d8f-dd6aee64ba02 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt- 4 with 90%* chatgpt quality,.

Detailed Object Description with Controllable Dimensions Vicuna: An open-source chatbot impressing gpt- 4 with 90%* chatgpt quality,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.039061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.039061Z digest=sha256:a262101f6fb83eb746d6777ae1be19335e5613ca6b9a9e8055423da582c84977

Observation a779451d-4b33-46cf-8b22-2473c7e98cfa · outbound

This paper cites Llama 3 model card,.

Detailed Object Description with Controllable Dimensions Llama 3 model card,

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.046825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.046825Z digest=sha256:bcc617047f638e3f3c776a6d72b7d942c7db377cdc26f64acae4198ed8202234

Observation bf3e4c25-36dc-4702-a277-c20a1ec407f8 · outbound

This paper cites Microsoft coco: Common objects in context,.

Detailed Object Description with Controllable Dimensions Microsoft coco: Common objects in context,

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:48.056317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:48.056317Z digest=sha256:053d46f5ffbea39aa54f617507ffd6409298947bbbe0bfba0a3b7ca7c2c3e615

Observation 64d73cef-163e-4ba0-ab26-dc194996979f · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

Detailed Object Description with Controllable Dimensions Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T10:35:47.855603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:35:47.855603Z digest=sha256:7370364264639e2a1c24533f508a448a3f5597ef45d9cb909f397bab7e1f4d4f

Pith citing papers

Observation 144dfe14-17ee-4373-af5d-381632a1b5ca · inbound

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation cites this paper.

Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation Detailed Object Description with Controllable Dimensions

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:26:16.179513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T15:26:15.313911Z digest=sha256:478149015df5edbd298c0def1b5337e529f1f078dbe76c016a27df0515327e40