Pith. sign in

Paper Citation Record · LEDGER

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation

As of 18 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2412.09817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09817 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:45:44.870849Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab2c5cd8-8afd-46f6-aff5-a1120b759a2c · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.659646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.659646Z digest=sha256:07cf97817416f4d17c10a843e6613c446230534f71681f50a198fc40b3c4b796

Observation 26459d2f-3047-470f-9171-71096ab547f7 · outbound

This paper cites write newline.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.664145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.664145Z digest=sha256:4fac0194307764d9a1174b3ed4cc3a157c2bfe05ebd90157b817a417f02318e0

Observation aaf457e3-f2df-4deb-8169-82c9c15c62a4 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.497001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.668737Z digest=sha256:a31a506ad26481ecaf068f97dbff158191e755c462f4ffb61ce238927079d0ac

Observation de3f5a70-9bd8-4bca-bd00-eb1f083de74b · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.672692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.672692Z digest=sha256:825255bf9191536b5b54335d2c0fa3acd49faffe3d01f130f9e745c943d73162

Observation c065c839-cc18-4729-a2f2-bd9c412471da · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.677283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.677283Z digest=sha256:2af982d0775fa4b1101ed177ed8dba32b4f7ecba0eb6dc9e610d37fcabd703b1

Observation 8deb9786-40bc-4052-b46b-5b4eeeda68e0 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.486002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.681789Z digest=sha256:2b35c5ab67c762143fa1223e5665feb37ca4c3529eb1ed3b3017dc1eb3d68998

Observation dc2b24e6-c43d-4ecb-b453-762a7e6c9de8 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.689734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.689734Z digest=sha256:7f13c5cd89bc6778e21ccb6c5c5af5f8cc542a44468ffb147fdd25495b0869ad

Observation 6b751465-a209-4e61-9e02-4ab9df02e2e7 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.693657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.693657Z digest=sha256:d3d6fa82388ee11fde4eda03c470f3f959e8091186c8bbed4764af4e98b4841b

Observation 18c75d0a-b2f2-4773-a2f8-8fc7c8dd3284 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.697445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.697445Z digest=sha256:5daa4aad77937c66ab63fc2aac0b96b087d50180ebb4064548dcda59f5ac4e0c

Observation 3c3a9258-cdbb-4d5b-92a2-039e3c1d37d9 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.701291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.701291Z digest=sha256:4ab7168ef46c0d2c978de2c67998e72aa3a93094d4aab1169118941aedf54ad1

Observation 4adf3b1b-e2f3-4a84-8a27-eb4d1093642b · outbound

This paper cites K.-p.; Qian, T.; and Wang, S.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation K.-p.; Qian, T.; and Wang, S

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:45.462402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.704823Z digest=sha256:4bc17c61b49abc821c511fdfb2acad50a354b00fb3035867b11f11f5563bb65b

Observation 57a9ad2d-a67b-4079-a171-e7daf591b99f · outbound

This paper cites VSE++: Improving Visual-Semantic Embeddings with Hard Negatives.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation VSE++: Improving Visual-Semantic Embeddings with Hard Negatives

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.708251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.708251Z digest=sha256:476c72a5779ce73bb9feb18b024a1e4260f9e58895f18d44b3986db835274406

Observation b3883893-5b76-4d5c-a712-3e810113d082 · outbound

This paper cites S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation S.; Shlens, J.; Bengio, S.; Dean, J.; Ranzato, M.; and Mikolov, T

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.711987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.711987Z digest=sha256:381bba7695998f74857bc4bac1ce5b0179ffe2ba655ffad5d75b743311f92601

Observation 7821a4e7-f3e3-4e19-926c-50d9ee9197ce · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.715548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.715548Z digest=sha256:3b84d6bc5938a875636afe64ffc82c069aa2e54ed8677785f657546b919d116e

Observation 4a80bfb5-b0b2-4573-94b9-876f2eb89f3b · outbound

This paper cites OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation OphCLIP: Hierarchical Retrieval-Augmented Learning for Ophthalmic Surgical Video-Language Pretraining

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.719020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.719020Z digest=sha256:36931ceec293435119b3eec59755861aa18546740e41e04124146a1167bc763b

Observation bce94d8c-11e7-4d67-beda-510fc5ad6544 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.431894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.726564Z digest=sha256:e35817b629f14d80a53b2764f085b231a4a578ea2a9b2a8ec6367b6e5e9e671e

Observation 34dbc0fd-30fc-44bf-afd8-40fdd24a42f4 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.730495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.730495Z digest=sha256:354b36ff7f69161d6ac46d02aae670d3ffadbe4a3a1be2aea27166e9257198e4

Observation 5d6d876d-f52e-4c50-b6bc-c7d7e7457558 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.736267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.736267Z digest=sha256:5390894903cb6290f672ffaef0143ae4bbefbbeb0215df885ef8e1ff5d76c273

Observation a09af1a3-2b9b-46bd-ac49-81e4258ca51d · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.740539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.740539Z digest=sha256:de971df36c48ed482b0944d11745f58fb8aa1d698916d196605a2333de4f8e3d

Observation 0f706992-c754-4cc4-8ed3-cfa581e8a94e · outbound

This paper cites M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation M$^3$IT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.743967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.743967Z digest=sha256:eef27dbabc706c104cfcc58e6b311c5e6e65618d2c18ecc12bf3913cf25053c6

Observation 3d81ac7a-ff0c-4cff-8b33-21f531c92a70 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.409110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.747739Z digest=sha256:0ae07d31f71b1b4ef5581884fc146c27b29a5d685cef324158a6d8fdedd8397b

Observation 9bc62403-be46-45df-b7b9-ac5c02775dc1 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Improved Baselines with Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.751352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.751352Z digest=sha256:9d45f964a8c1cc5c749f9bb3cae9cc0a6ab5920f41baac1ddecf3ee0ee0f11d0

Observation 620194f9-467b-4a8a-85a8-c0e622fdbf23 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.398877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.755283Z digest=sha256:a44bfbecffbf85910ea675d9c40019093bf67a92f920bf8b376982f29d67ab92

Observation e42b04e2-881a-4cf1-9295-94638494492f · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.758732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.758732Z digest=sha256:996510fb55dea7c7a0b3c7b60162d27040106fe7ddb2951ebf152b5a94a33d9f

Observation 76bb4713-f1db-4249-a7d3-f824a59d7501 · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.762111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.762111Z digest=sha256:2131ff46300fd4c7dad6981793044f2cfd2fac399d57030866c09dc8d665c2eb

Observation d88a1c08-ceea-4557-8d90-ebd642c51052 · outbound

This paper cites J.; and Yan, Y.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation J.; and Yan, Y

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.765864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.765864Z digest=sha256:5a0a2f851e1973839abd3eb6f0d2bd3b6d911c0e25c6c3ddec05d9268d7a2541

Observation 7b70f671-3c40-4dfb-8340-20523a88b89a · outbound

This paper cites IMAGDressing-v1: Customizable Virtual Dressing.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation IMAGDressing-v1: Customizable Virtual Dressing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.769254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.769254Z digest=sha256:70e05dbf3086dccd54d32a00d37a495060569e3af21a885ba400af8a5b2382b0

Observation b5ed5497-9a08-4ec5-b7b1-52e0524d14a6 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.773148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.773148Z digest=sha256:1639e755c866f1ebf38eae4b2de5b6577a35abab5090004687dd4016a14c6a5a

Observation 2a03c5ce-a27c-4224-ad62-437f84aff12a · outbound

This paper cites Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Boosting Consistency in Story Visualization with Rich-Contextual Conditional Diffusion Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.776695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.776695Z digest=sha256:7f37b537a93f389adf35541155c0117010a5a450ef90e7dc9109126254a181ca

Observation d74a7b21-5708-4231-a500-46c57347085b · outbound

This paper cites ???? Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation ???? Advancing Pose-Guided Image Synthesis with Progressive Conditional Diffusion Models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:45:45.368028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.780400Z digest=sha256:896a9e0a3666d416fc86aeb167d36168909939d30856fe1074a0c2a0e01e9ea1

Observation 3e395ca2-e21a-4d8c-a5a8-a44655a5da63 · outbound

This paper cites Parrot: Multilingual Visual Instruction Tuning.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Parrot: Multilingual Visual Instruction Tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.784099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.784099Z digest=sha256:23437a861001455835274f99815444586de449985764c303aebfac6a09b83557

Observation 34a5d01c-c8e4-45f6-aac7-c14682479013 · outbound

This paper cites PILOT: A Pre-Trained Model-Based Continual Learning Toolbox.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation PILOT: A Pre-Trained Model-Based Continual Learning Toolbox

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.788078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.788078Z digest=sha256:0b3939fafe9d39b066b164c82dec3383742c52b269347d59b61e81fd107a2f99

Observation 396e2e59-ce9f-4521-8481-8b8377aa5703 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.356442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.791824Z digest=sha256:6c34f0397d074eb4f98291eec9749797125edd7e6fc7fd6c498d06213d6cad09

Observation 56b8f89a-496f-47a4-b0a1-88e60bc87b48 · outbound

This paper cites TextSquare: Scaling up Text-Centric Visual Instruction Tuning.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation TextSquare: Scaling up Text-Centric Visual Instruction Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.795528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.795528Z digest=sha256:a9038369f33534eb4357b6107e4df0e2331d9503d00c9ca60691400eb99fae41

Observation 642c7b47-686e-4b37-9202-3497554fc21a · outbound

This paper cites MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.799374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.799374Z digest=sha256:0b411e5dc4f1378be9c9e763366193f48e87789d621dae0fbdfa0e4937b4e347

Observation 442acfa7-1bd9-4235-8824-f33b9a04493c · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.345348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.803209Z digest=sha256:097c62a1f17abbe232ea00ac737b6801d6795b6107902442ea96724de6cab952

Observation c3ab9fa4-b95d-4483-8a88-3c089406603b · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.334384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.806264Z digest=sha256:8c2fe58eb9018160981d37ee9a2f7fa5eaf5194578df40a5b449e945f48abe8c

Observation 56a4b260-275a-45f5-b687-9d6ec4439724 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.323492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.809370Z digest=sha256:2a9029a02e00fe52079b7be74e151ee2e821813929084f336828f2c9bdc6c7b0

Observation 83f60d85-ae1a-43ce-bfa9-083fdd2e72ff · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.812686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.812686Z digest=sha256:4b3dab4d5744abc22d12ce9fe14e50f377184d0cc1253b9349585bbcbddc7f48

Observation e68e8702-e122-46c0-9500-451de0b6bcb1 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation CogVLM: Visual Expert for Pretrained Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.816068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.816068Z digest=sha256:1f6d81c15ff97d590897c7e6c9e7f9a9f958036cc6e79c81ffcad1d21921e090

Observation f8382428-2d73-451c-8dae-0309364dc75c · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.819685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.819685Z digest=sha256:07671047a4d365a0e39829a587767e5f7fdcdf3e73696f7d2b74b0c388b5262d

Observation f71c19a6-29b3-4554-89fa-06a4eb34e964 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.823310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.823310Z digest=sha256:8845664104526358f84706470a58b1ee9946efaed80d21933f98447f77090159

Observation fc69dfbd-bff1-487c-adec-93489945bde8 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.290479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.826541Z digest=sha256:1366bf24ee3e7aa1515102026c3557269e146088ec6d3190f60dfdb8ac8af52f

Observation 1b992b59-3a11-4214-8781-0888aec3cbb7 · outbound

This paper cites Instance-adaptive Zero-shot Chain-of-Thought Prompting.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Instance-adaptive Zero-shot Chain-of-Thought Prompting

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.830014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.830014Z digest=sha256:7735d3974fdbe67e8b2e74fc86bee73bd0a9660f3e8174c67e0ae7299d10e8bf

Observation c2a20de0-3f71-458b-b134-5d1d6262bc07 · outbound

This paper cites TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.833743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.833743Z digest=sha256:92ec0d73aa7d07cdf1b9ff7a49239dc2bbf4e307b06e36d3228653fa8043fb48

Observation 5afc7bfb-70f5-45c0-86c0-7b265c26c9ad · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.279234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.837610Z digest=sha256:618224a9e36031cdda890845d4f0e9fbf8294c9bde489324ec07371c92bac594

Observation 0695b377-4d15-41ff-9a21-6d63b07ac2d6 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.840954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.840954Z digest=sha256:0412562cf74e76d06c66e844d32f5f48f5eaa164e654930d7a1755c5cf68a61a

Observation a995798a-d4f3-4859-90c3-06a9585b31c1 · outbound

This paper cites Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Seeing Clearly by Layer Two: Enhancing Attention Heads to Alleviate Hallucination in LVLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.844437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.844437Z digest=sha256:9a920bf39afce054572fe1ada9ddf308e70f49cbffec1db4731ed21620533a93

Observation 74983426-dfa4-4d02-85e0-4281e23a878c · outbound

This paper cites From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation From Redundancy to Relevance: Information Flow in LVLMs Across Reasoning Tasks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.848057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.848057Z digest=sha256:0ac1f8fcf696ec38d4e075e117521e36614013546a18d3abf0386ad0649644b0

Observation 47904241-94ad-457a-85e9-7447f8ee7bd7 · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.852046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.852046Z digest=sha256:20837ae97082fbb73a6aa179f606546020a4691d1c0b0fb4e1e7a8e578772eec

Observation 46d43404-51d8-4228-9e7c-5db68ea67d8e · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.260897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.856179Z digest=sha256:8ff6c288cab02bc5768eba036982cb803a3733833c2ec177c838915ca3d62def

Observation 6ea5cc41-ee29-46f1-ba0f-10ee75153277 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.249757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.859985Z digest=sha256:a0bfe3f13a6b62c038d9d3762593ba99fb4b2e3e44ef6c25f9e5df7fada9fffd

Observation c4bba911-bde8-4b85-a860-85edb8eeb5a4 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.238750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.863687Z digest=sha256:b179a36ef221003f4a47d6011e568d6a72a3cb4706dabdfcaa862d721e0cf96b

Observation 3478655e-384c-403c-8842-4eb1ba3418d7 · outbound

This paper cites an unresolved cited work.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:45:45.227225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T16:45:44.867315Z digest=sha256:ee20cf597ed3601116cc4867a4c350a0c397d9a5819f31e6201d745f6a11e8d8

Observation 29ae8978-1afd-4aa4-9e5f-a58527ee7e6d · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Enhancing Multimodal Large Language Models Complex Reason via Similarity Computation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T16:45:44.870849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:45:44.870849Z digest=sha256:18add3f229f1aac933de092af30044431b016526dfe510639a026d5c1bbcd7fb

Pith citing papers

No inbound Pith citation observations are available.