Pith. sign in

Paper Citation Record · LEDGER

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework

As of 10 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2502.00711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00711 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:03:53.348354Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:33:44.595993Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T14:33:46.092736Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact6
  • verified fuzzy30
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 649cd177-51ed-4f04-9e08-d8538d083313 · outbound

This paper cites PaliGemma 2: A Family of Versatile VLMs for Transfer.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PaliGemma 2: A Family of Versatile VLMs for Transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.108911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.108911Z digest=sha256:f3812b9548191792e02c7ae94a2c41b44ef4cf781150d8567f9e4788bc884d02

Observation d1e22241-9f19-4b57-a981-2e8df316fd40 · outbound

This paper cites Visual instruction tuning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual instruction tuning,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.114192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.114192Z digest=sha256:be7dd177b2cbdeb7c6a66fe54f90aa23abb020649e5ecab34900e4ade330b7e2

Observation ddb8f5b2-24eb-467b-9b82-dc9abbc70039 · outbound

This paper cites Mdetr-modulated detection for end-to-end multi-modal understanding,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Mdetr-modulated detection for end-to-end multi-modal understanding,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.273883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.118804Z digest=sha256:8f80806f99efeeac85e2c559207e6dd361b774597f69272e9b8934a6edbc3dd5

Observation cad07e75-4fab-4b70-b696-a2df3173f6ad · outbound

This paper cites Omni-smola: Boosting generalist multimodal models with soft mixture of low-rank experts,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Omni-smola: Boosting generalist multimodal models with soft mixture of low-rank experts,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.259335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.123583Z digest=sha256:386623d3acb495cde1c9979200aae5f76c3ef397c1779af73fd9c14a2fd7f930

Observation 78d15b40-0858-4e91-b723-e66af2381bdf · outbound

This paper cites Cola: A benchmark for compositional text-to-image retrieval,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Cola: A benchmark for compositional text-to-image retrieval,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.244820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.128033Z digest=sha256:d6508799799fe7a66165d9a536d6a521fe63760d275a1ce4daa0f98843cb20a5

Observation fd748ad8-7833-406c-b731-c2e9de31d20c · outbound

This paper cites Visual programming: Compositional visual reasoning without training,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual programming: Compositional visual reasoning without training,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.230148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.132829Z digest=sha256:fab19d78ebdbb152a1ca9dc9790fefecb8be03753f3d8d929270f7e8e441e663

Observation 13ca5d95-21ba-4f0c-83a7-a6ef98b275c2 · outbound

This paper cites Interpretable visual reasoning: A survey,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Interpretable visual reasoning: A survey,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.216095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.138224Z digest=sha256:7645e38d06d604fbddd2077c5169ba32254a4b00735620fdc3603598eaaaa058

Observation 9d897cae-9bd3-4bce-925a-d4d70a53e570 · outbound

This paper cites Rapper: Reinforced rationale-prompted paradigm for natural language explanation in visual question answering,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Rapper: Reinforced rationale-prompted paradigm for natural language explanation in visual question answering,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.200839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.143782Z digest=sha256:3cfd250456ec3aabba97d8662e2422b53671781bbd38c475e20a1cf8817aa2ef

Observation 87d224cb-6bc4-47bf-90d1-cea4ba5904f5 · outbound

This paper cites Rephrase, augment, reason: Visual grounding of questions for vision-language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Rephrase, augment, reason: Visual grounding of questions for vision-language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.183457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.148002Z digest=sha256:97aea64a543cf63b7bd8935a4e783c145f2c1b721d4da170d366a930f1ca641b

Observation 9c97c9b4-04b5-43ad-876c-31afd58a9186 · outbound

This paper cites Toward multi-granularity decision- making: Explicit visual reasoning with hierarchical knowledge,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Toward multi-granularity decision- making: Explicit visual reasoning with hierarchical knowledge,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.152331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.157639Z digest=sha256:d97e4bec81ee21738f1760de291cd199fd8aabd501e0671eb66db0dd8f43de4e

Observation e6cc5687-cb05-4ac4-89df-ad9c12f108cb · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.162330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.162330Z digest=sha256:0226151ed2a9eead425a51d7a7d9ac1a5859ade757cfe9bc307d9605071faf31

Observation 4ecd5353-0806-429f-bd7d-adc11a0573bb · outbound

This paper cites Vision–language model for visual question answering in medical imagery,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Vision–language model for visual question answering in medical imagery,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.137221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.167035Z digest=sha256:5557e3d17bde3eccb2cec800284060466ad1a058ebd1425a2d042bd8da8ef8ed

Observation 09bf68a5-9800-46fb-a189-f9a5aa0935ad · outbound

This paper cites Lingoqa: Visual question answering for autonomous driving,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Lingoqa: Visual question answering for autonomous driving,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.120736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.172078Z digest=sha256:de5692f44a04c56da0ff7b5dfb31b26a520381dfc68e25483821057c3f7aa169

Observation 493c5b92-5199-4a6a-9220-f8dc3ab48917 · outbound

This paper cites VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework VQA and Visual Reasoning: An Overview of Recent Datasets, Methods and Challenges

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.176903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.176903Z digest=sha256:f3137d70a26b758c6022044e946b531b733e7ce8d4afbd4c887cac8b88381125

Observation a8ac84f1-acd6-40fd-a085-f22403b49a64 · outbound

This paper cites Dealing with Semantic Underspecification in Multimodal NLP.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Dealing with Semantic Underspecification in Multimodal NLP

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.181943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.181943Z digest=sha256:2344f5979de0de8e4cef495a8cdf8ab84c7fb17c70d99d589cd3439fce24757c

Observation 201d67f5-58d4-40b1-87a2-6cb46d878e3b · outbound

This paper cites Open visual knowledge extraction via relation-oriented multimodality model prompting,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Open visual knowledge extraction via relation-oriented multimodality model prompting,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.103734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.187050Z digest=sha256:5d94b5a4092136b43a8b23d2a6ef95f0d5a244f60f41baf04d9d7bcbcd3d2de8

Observation 45b14339-5b4d-41df-aaf5-dcf88d0e2447 · outbound

This paper cites PV2TEA: Patching Visual Modality to Textual-Established Information Extraction.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PV2TEA: Patching Visual Modality to Textual-Established Information Extraction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.675681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.196382Z digest=sha256:71d6cf189b0b35649fff12d6a2de11691dda2e4d56ec7fd931aec99ca93c5c33

Observation c41710f0-af35-4fab-8584-779222491bc2 · outbound

This paper cites Recurrent fusion network for image captioning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Recurrent fusion network for image captioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.073290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.201585Z digest=sha256:00ee9c76308263aa9faa252d76fa2c8f08aaf0d77e0909cae95037df49bcf7e2

Observation 9cc4f3b7-a6f7-44f6-a865-61ff0b598e07 · outbound

This paper cites Boosting image captioning with attributes,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Boosting image captioning with attributes,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.057716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.206094Z digest=sha256:bf19453591cf8ada76501055710792d0fa7ab384bbb19843e3cc525b5b43a285

Observation ec9a6f63-0b01-4463-8b84-a56fea172bf6 · outbound

This paper cites Promptcap: Prompt-guided image captioning for vqa with gpt-3,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Promptcap: Prompt-guided image captioning for vqa with gpt-3,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.042647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.210754Z digest=sha256:a7a9e24532e2abea34f53c5fdf5c93a74f0e5a49bea03f45945c512c77966e63

Observation 2b3de275-8ce0-4c20-9023-989fc64a8130 · outbound

This paper cites Injecting semantic concepts into end-to-end image captioning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Injecting semantic concepts into end-to-end image captioning,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.026025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.215491Z digest=sha256:c7befff2832613a49f5f089fcac336e7a7c7ec792d4a90b479f2c8979bce718e

Observation 44a583f8-05d8-490e-9b87-d344108e3e9b · outbound

This paper cites Visual Commonsense based Heterogeneous Graph Contrastive Learning.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Visual Commonsense based Heterogeneous Graph Contrastive Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.652545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.220310Z digest=sha256:91aa9afce82d7529f857e517fe6f64518f463d92418a430bfa1a17122238fec7

Observation 576ee7d7-bb2a-411f-bda0-bc7818ca20da · outbound

This paper cites Covlm: Composing visual entities and relationships in large language models via communicative decoding,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Covlm: Composing visual entities and relationships in large language models via communicative decoding,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.011138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.225287Z digest=sha256:c3235ae5949f48e7a647da17fde7f17b2eea8313e15e6a0c192e698857cf2bd5

Observation 859becd4-8893-49f5-8051-0a6a55c14891 · outbound

This paper cites Bridging knowledge graphs to generate scene graphs,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Bridging knowledge graphs to generate scene graphs,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.995215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.229765Z digest=sha256:bd50a0c364ee3fabc9bb294b0311f93b51323730b25cbeca73c46454021a6c10

Observation 17da6dd9-5f85-4145-b893-db9cff8abcfb · outbound

This paper cites Co-training improves prompt-based learning for large language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Co-training improves prompt-based learning for large language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.979908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.234699Z digest=sha256:121f0555bcfd3b0e4429f05750d4fb4bae0265698f1b2ef3dd8888600c9c97dc

Observation 352cb6f8-0e86-4031-9d49-c0d03e1ee2c6 · outbound

This paper cites Large language model as attributed training data generator: A tale of diversity and bias,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Large language model as attributed training data generator: A tale of diversity and bias,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.964531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.239804Z digest=sha256:1fe26c624dbaba337fd06fed491358873326c35878150116cef912ae5fcb7a9c

Observation f23dd74d-2d9e-4d04-965a-e696057e3844 · outbound

This paper cites Large language models are zero-shot reasoners,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Large language models are zero-shot reasoners,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.244544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.244544Z digest=sha256:a67e32da39d060ef62be9a4d766d71754c92ac0d29f0a5660df6b1baa8ca28fb

Observation 72faaa99-7ef1-4082-8c89-e4dd2dface47 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.249363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.249363Z digest=sha256:b7c170d35f46178667f54b24786c7e67448f130b6fe9c50cad50dfbae37aed09

Observation 3bd099ee-9f1a-4352-b2f0-45ae8faa3e49 · outbound

This paper cites e-vil: A dataset and benchmark for natural language explanations in vision-language tasks,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework e-vil: A dataset and benchmark for natural language explanations in vision-language tasks,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.937640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.254268Z digest=sha256:a611850e590098dd977719426bef1a6ce1128448e85fa36c04cf59b5265d4f46

Observation 37485df7-ccff-4907-a53b-d104118fabc5 · outbound

This paper cites NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework NLX-GPT: A Model for Natural Language Explanations in Vision and Vision-Language Tasks

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.411763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.258591Z digest=sha256:4c9c0663a20da8ebcdf3237b65689eaec7b489829738f557c6c55561e2fb8884

Observation 7e32c95e-1594-4df8-a12c-5606734b071c · outbound

This paper cites Improving vision-and-language reasoning via spatial relations modeling,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Improving vision-and-language reasoning via spatial relations modeling,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.088306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.262983Z digest=sha256:5e72128e67f9589cb803813c01dc47396d391f0865c7c4ad7b6edd2111fd9df8

Observation 581bbdad-363a-4dde-a6dd-c3153a860408 · outbound

This paper cites GPT-4 Technical Report.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework GPT-4 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.267186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.267186Z digest=sha256:088805076dc353931865da6cf53524cf91c35586848b26ad57c46fa7275e0e5f

Observation 49a49c73-976b-48ed-b03b-edf5ca7db438 · outbound

This paper cites Chain of thought prompting elicits reasoning in large language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Chain of thought prompting elicits reasoning in large language models,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.920798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.271528Z digest=sha256:4c01649e0fc9a2206f08c466367b3244fa5c67d3b76f247320a5a7aaf7f8ce83

Observation 6301d550-8086-41db-9af7-2e5291357021 · outbound

This paper cites Re- flexion: Language agents with verbal reinforcement learning,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Re- flexion: Language agents with verbal reinforcement learning,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.276014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.276014Z digest=sha256:26e952da073f070930439001b62f6400914f25a155bbfd96ee2e507576e72bf4

Observation 48cb50bd-3dce-4d0a-b28a-5375e5202022 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.280509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.280509Z digest=sha256:558743664892ac0e02ad6814a0968c47cbb86068b94c083816ba86f06d9120e5

Observation ffc1b197-7c6f-4a7c-ab72-dbc64b4aa10b · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowledge,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework A-okvqa: A benchmark for visual question answering using world knowledge,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.883686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.284831Z digest=sha256:bc8f5de44ef58239b4335d50537c12f9c4382073d5da5e3bba75d71c521f7dd1

Observation 2c7d8b91-5454-4649-8acd-ffe132606bac · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Vizwiz grand challenge: Answering visual questions from blind people,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.866801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.288845Z digest=sha256:49f623119509a38b9fa100dc1b83ed16dfd759d010cad2072c95dad916d5afca

Observation b4a91c25-d32e-4a6f-a5d3-5e61ae026181 · outbound

This paper cites @ crepe: Can vision-language foundation models reason compositionally?.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework @ crepe: Can vision-language foundation models reason compositionally?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.293183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.293183Z digest=sha256:9f6014a4c44eb4705f0f0b94d2c3634d729388825123e87d62862f67152472e7

Observation 4ebb616e-16a6-4190-a4bd-660078d895be · outbound

This paper cites Qwen2.5-vl,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Qwen2.5-vl,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.297452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.297452Z digest=sha256:e839227bfece30bfb50fed1bb909544076cbad1678ec12fc58824f2cadf2bed5

Observation 1f87daaa-787a-4922-9028-3e10613dc958 · outbound

This paper cites 4o mini: Advancing cost-efficient intelligence, 2024,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework 4o mini: Advancing cost-efficient intelligence, 2024,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.838477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.301534Z digest=sha256:5659ab51e2f439caaeb8bffd233fd4cf4ea57bb4630bd6d4424e11b079e143e7

Observation 39ae7dd0-f211-4e54-874e-c68fc8233f56 · outbound

This paper cites Hello gpt-4o,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Hello gpt-4o,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.823122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.305912Z digest=sha256:8ac5e03987441dc785a8dc8fa9c434882a7473621482d5b8cc1a089d278e2bf5

Observation b03738ee-646c-4502-8f5f-188b317569f9 · outbound

This paper cites Introducing gpt-5,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Introducing gpt-5,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.807507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.310555Z digest=sha256:44d80fb0b2345df11a3b5c8a9236a33eee893b36e34cf7f04fba5233e6b46c45

Observation 51f45e83-d199-4222-bc23-0a51e053ed2b · outbound

This paper cites Learning to localize objects improves spatial reasoning in visual-llms,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Learning to localize objects improves spatial reasoning in visual-llms,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.791861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.314572Z digest=sha256:721be16eb44ceabc05fd0965eef779a99768ff168eacea39f0fe61b768059c2a

Observation 8a1837d5-e9a4-484d-a90b-0c32bec08abb · outbound

This paper cites HiMix: Reducing Computational Complexity in Large Vision-Language Models.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework HiMix: Reducing Computational Complexity in Large Vision-Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.497001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.318565Z digest=sha256:b005d077c79d28a099df2b06878692f033454f4b72e06e933d005c7660e97fe0

Observation d453f790-a4cd-49b5-b5db-c2bb756e6185 · outbound

This paper cites Eliminating the language bias for visual question answering with fine- grained causal intervention,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Eliminating the language bias for visual question answering with fine- grained causal intervention,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:53.774850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.323075Z digest=sha256:1cde5530fe94b02fd815bd027f6173c3c7855b434f2fcf8419cf11d4cc6bcd1e

Observation bbc7420b-076e-4705-9814-f9e1961ef556 · outbound

This paper cites Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Multi-Agents Based on Large Language Models for Knowledge-based Visual Question Answering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.327797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.327797Z digest=sha256:70de3fd705f8dcc00bbacc9e62c88131986e75c4f91baea1322d8b2db41ce640

Observation e538e280-5cfb-46b7-a517-36b48a81d7a5 · outbound

This paper cites Diversify, Rationalize, and Combine: Ensembling Multiple QA Strategies for Zero-shot Knowledge-based VQA.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Diversify, Rationalize, and Combine: Ensembling Multiple QA Strategies for Zero-shot Knowledge-based VQA

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.459183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.333286Z digest=sha256:e6cdbb5550685d76c1a5f1b055aa0d8a2712c382cb9a115436ab9223a8b034dd

Observation cfe634d1-97ed-4d42-a49c-e3f291b3d187 · outbound

This paper cites Harnessing the Power of Multi-Task Pretraining for Ground-Truth Level Natural Language Explanations.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Harnessing the Power of Multi-Task Pretraining for Ground-Truth Level Natural Language Explanations

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.338237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.338237Z digest=sha256:356a6eff5b687e6e10866e0ada6442ec5fc80ff7f86b04be7cdbafc43d23c846

Observation 614997c3-fa05-4221-81c6-761128f6e041 · outbound

This paper cites Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Coarse-to-Fine Contrastive Learning in Image-Text-Graph Space for Improved Vision-Language Compositionality

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-09T18:03:53.436188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.343259Z digest=sha256:e8cd448963a350e7f3df2851f23a6081ea0637bf204890375f5c29138d625768

Observation 886cfc9d-39ab-4256-aaa3-4f24b4803d38 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T18:03:53.348354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:03:53.348354Z digest=sha256:1e60eef7b99ed945f318f283d811823a7568337de719a8dee75bc33191009f16

Observation 0426853d-c62e-4887-8569-5ba81f7ec38d · outbound

This paper cites Available: https://openreview.net/forum?id=L4nOxziGf9.

VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework Available: https://openreview.net/forum?id=L4nOxziGf9

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T18:03:54.168000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-09T18:03:53.152560Z digest=sha256:0c2869bb20cdd59696841f1c5014e696c4a238faef511a6680e0c33cdcc2c905

Pith citing papers

Observation 798e4ab8-facb-4bdd-b162-b30819d9d17f · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review VIKSER: Visual Knowledge-Driven Self-Reinforcing Reasoning Framework

Reference 129

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:33:46.154777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T14:33:44.595993Z digest=sha256:848d2a917b1a1a4548e40262e8acf08325dd910c414bcab3a45b222c4760a6c1