Pith. sign in

Paper Citation Record · LEDGER

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

As of 14 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2608.11167.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11167 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:57:29.583850Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact9
  • verified fuzzy12
  • unresolved64
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 565c71e8-44ee-40a7-bc9c-ca5e5cff1448 · outbound

This paper cites Visual Instruction Tuning , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Visual Instruction Tuning , booktitle =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.196786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.196786Z digest=sha256:8ccb5a47c6fdd4b64bafbdfcd24ef9cf39f1e00649c8c2a1409ea0bae57e219e

Observation 885c6a52-3474-490d-95a4-3e54da950b4a · outbound

This paper cites LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.205718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.205718Z digest=sha256:ce67b5bbb0a0d5fe507f00716cf5a351d17e83a83c0491ff35118ba3c2a83f1a

Observation 4b2b103e-19f2-4ef5-8ddd-af34039fd017 · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.210383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.210383Z digest=sha256:70382b6b1825a9a77de144e152a28e35c839380a250ab198eb69150eac1c7534

Observation 0bc2528a-21ed-49d6-afd1-97a4894d6f61 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.214140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.214140Z digest=sha256:aaff3f616ab9d9e39fb076620be2a25be28dc239b3ec654146b599c143b72c0d

Observation 754d5abb-09e1-49bb-a305-b31610cae958 · outbound

This paper cites Qwen2.5-VL Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.218561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.218561Z digest=sha256:3bdb2f064f0b6771b060c5904337812491d81d9e4cdf1b55b90400fe929a7de0

Observation a3f5ab58-210a-4f26-be6b-b9e5e7637237 · outbound

This paper cites 2025 , eprint=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2025 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.223012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.223012Z digest=sha256:696773ea90a391e599fa4e9888442d47f7430439371a4e37156eb09b1e5efa16

Observation ec6fb014-cc3f-411c-97ab-3317bc25fe51 · outbound

This paper cites MiMo-VL Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment MiMo-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.230150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.230150Z digest=sha256:55953e8bd42bfcdb3296e2cfec8405ac2bf7a04135fcabb52f4f941b77ded60c

Observation c498bd22-1e36-48e9-8295-3b69342f6ca9 · outbound

This paper cites 2025 , eprint=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2025 , eprint=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.234182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.234182Z digest=sha256:6667f33ecaf8e9722ee26b15c05d43d6153f3c03265daa81612334a4ce9fa9f5

Observation 56cd082d-a781-48ff-a7b9-25ea335c561a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.237647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.237647Z digest=sha256:c92d27de2322b32da64d04315969f0cf4e5a3ca4be6df3974603bd9cab38ccf6

Observation 9700f8ad-d9d6-4c9c-9ab2-e256ead06748 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.241637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.241637Z digest=sha256:8e70c8bd93da19e567edd51f616d5b83bbbf45a0b298c53bc5564d2aa50b5bd7

Observation a98493a4-1f8a-4c78-8c33-d5e957b981c8 · outbound

This paper cites Scalable Vision Language Model Training via High Quality Data Curation , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Scalable Vision Language Model Training via High Quality Data Curation , booktitle =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.245635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.245635Z digest=sha256:02fbb83ec81719d2ab9dfa39b8d90a666a92a208ec10a3b5f5d50400dcc980bc

Observation f4742661-3eac-49a7-96af-d507aed679e0 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Flamingo: a Visual Language Model for Few-Shot Learning , booktitle =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.249132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.249132Z digest=sha256:d7c6396c08da67492d21de1800a4de7bea3c332977df8a587e5c926c66d9cf62

Observation 5e997eb9-1d21-4376-aba1-c8cb9b751f88 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.253141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.253141Z digest=sha256:c43110a3d9b69d1a6a138a7bf73a07ca9e3832703806e0576c09d2bea0412eff

Observation a1ca3212-96f5-4120-8be9-32b9fd38b029 · outbound

This paper cites Seed1.5-VL Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Seed1.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.257406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.257406Z digest=sha256:0efc2de579c95d2f0ae4c11aa924ae20f2c7b4bcb64bf4369e7c3d49e4bfd9e7

Observation 606d72bd-1521-41af-86f8-ed8e16e64fee · outbound

This paper cites CoRR , volume =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment CoRR , volume =

Reference 16

Resolution
verified exact
doi, observed 2026-08-12T04:57:30.180987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.261748Z digest=sha256:ad68efc622909285d0df9b089a2f85db4ef0075420211791fdbcf6902b03189f

Observation ea48c056-da6f-48a1-ab50-01bb588aa137 · outbound

This paper cites Computer Vision -.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Computer Vision -

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.265494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.265494Z digest=sha256:bbf590d4cbf98e41ec46df02f60b72dbb6246117c93f5d620b342e9b4fe22df0

Observation a47e8c63-bfea-4588-8ce7-f1fe15e5d248 · outbound

This paper cites Qwen2.5 Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen2.5 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.269711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.269711Z digest=sha256:3a25f43cc8dc42b4622de6abf09b7acc57b453a48c757e6f1ebba0fca0a99923

Observation de5b4bd7-3978-4b42-83fc-81c5c01d00f2 · outbound

This paper cites Qwen3 Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.273623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.273623Z digest=sha256:490deb6369a5b4a4f69427ff8f2ae0319263d5e22c312c5fbcaed81574a4daa4

Observation 8679e8f7-be7a-4ec3-9f82-ee1df8f835b9 · outbound

This paper cites The Llama 3 Herd of Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.277639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.277639Z digest=sha256:a7f428cce75189fcf40a9caaf2035f63a4b0ae4b40f2902b3ff483a42e8d790c

Observation cd8b26b5-f973-4828-8103-9c4e6686aec1 · outbound

This paper cites Making the.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Making the

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.281596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.281596Z digest=sha256:2cecfc5487dca949a898b09617c2e9fa78f02c46e40e31a55f7a675a77a2b4a2

Observation 7843997f-c70e-4364-835e-31a78e905f08 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models , booktitle =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.285486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.285486Z digest=sha256:76bcdbb326b3a41777d7e534368b4c6001839ffefbecc0102e56d9eb5b0a582d

Observation 8fb8e444-3744-443c-9fa4-582fcb27a03e · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-Around Player? , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment MMBench: Is Your Multi-modal Model an All-Around Player? , booktitle =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.289625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.289625Z digest=sha256:bc2e9b92602e9ba4a72533abb11dcd44a5ee714a5731945d505ca4dbeb49ea94

Observation f478145a-65b0-4b2f-92cb-dec2857d45c9 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.294805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.294805Z digest=sha256:8b2fec16a34fdb25134f0e61a5e463b35b6c5303c7e22d2de4ad46c6eba0f4cd

Observation cc75e00e-fa7a-443b-9502-7c12dd2e591b · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models? , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Are We on the Right Way for Evaluating Large Vision-Language Models? , booktitle =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.299431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.299431Z digest=sha256:e4c69792c1183e42cbf29eb0f12b657cc621c66eff21ac090e26c830dbdaee94

Observation b6f4399b-c5c1-438f-832f-2d47f8d26dc6 · outbound

This paper cites Forty-first International Conference on Machine Learning,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Forty-first International Conference on Machine Learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.301283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.307220Z digest=sha256:22067b5e382a4c5140f041c3929747afc62420fcc7b83732c4cfbf5b98cfce2d

Observation 92af91a1-ab07-4a2f-bbd3-d324f400ce42 · outbound

This paper cites Hudson and Christopher D.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Hudson and Christopher D

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.311156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.311156Z digest=sha256:baf4a56beb3469f2c71997652ba82b632d14118ced043b53e027d065f7e11170

Observation b2889eba-9548-4e16-b9c6-e9ec1d9db9fb · outbound

This paper cites A Diagram is Worth a Dozen Images , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment A Diagram is Worth a Dozen Images , booktitle =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.314674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.314674Z digest=sha256:228e0c037c594292b28be000aedc22814fb3a2db06ca1e3044b7fb63c6273097

Observation 5d21991e-be69-43e5-b0e1-c1de80f27ee0 · outbound

This paper cites Joty and Enamul Hoque , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Joty and Enamul Hoque , editor =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.318875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.318875Z digest=sha256:b8a3447dd70f4bca6ca61e8be62c35e24396443f96eec8226d61793a3b779e7f

Observation 966b00db-cd54-4d65-85b5-0de2840a07ed · outbound

This paper cites Cambrian-1:.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Cambrian-1:

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.289350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.322900Z digest=sha256:180a08f3f5c1271930f2c2b5db84f37bc40ca08682d17d649677793f19f3f99b

Observation e8eb2ba9-63e7-4ac4-aa06-97e6ee4e2dfd · outbound

This paper cites OCRBench: on the hidden mystery of.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment OCRBench: on the hidden mystery of

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.327016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.327016Z digest=sha256:a3f486e502102061c8f87b79ddff941ac5b500cd10c610e4de5a0c57c2364d49

Observation 0bb53f02-c088-4c53-9237-85ab45e99862 · outbound

This paper cites 2019 , url =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2019 , url =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.330804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.330804Z digest=sha256:ed37834f8346f593c2ce2e7bc5abf6d2c25613d80a669414fec9c896849ddd66

Observation 8d61bc53-0d30-4d5b-8979-cb784e26e15d · outbound

This paper cites Berg , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Berg , editor =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.337861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.337861Z digest=sha256:15365f3ef8b8cf6933b23d38f196f974b2392536d8c5e01e62907c3b992cb38d

Observation 438c1bf1-98ec-4d3d-8c16-65f9f746a306 · outbound

This paper cites Yuille and Kevin Murphy , title =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Yuille and Kevin Murphy , title =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.341502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.341502Z digest=sha256:b99f44eefd0581147653cf06f0fd9b9da2af4989f214b81e8e3e67ff28cf2c48

Observation e74cd858-eb1d-44c7-9060-f68719cf461e · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.345672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.345672Z digest=sha256:266d69f6888a2225eed4e6e1ce94def00c7e347bf42128dc5d30af0b18fa6f14

Observation 205dfa8c-497f-424e-a738-adee591aca4e · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models , booktitle =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.349323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.349323Z digest=sha256:80a917cf7b1708babb9647b66b2798357e524f7740dd146d8d0753a6e8a23e5d

Observation 45e93531-70cd-4e47-bbed-9e0b49aca790 · outbound

This paper cites ShareGPT4V: Improving Large Multi-modal Models with Better Captions , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ShareGPT4V: Improving Large Multi-modal Models with Better Captions , booktitle =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.352991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.352991Z digest=sha256:de28f014a52c3f5dd28d57cabff1c397976e283a7240900ac2156439e75023f1

Observation 3cec23b2-d15e-479f-8f48-d39036e9083a · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.356681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.356681Z digest=sha256:39107b599c9ce1c43d18169cbaa152a9c950d960d4e156a8463b1451c39f4cae

Observation cef80f05-23ca-4fd2-b9f9-f0cdad2e25d2 · outbound

This paper cites DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception , booktitle =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.277887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.360560Z digest=sha256:665a7c0b148591e7cf6cb179b007acea5cb29ffb691485a5d152f2ae5164efae

Observation 89047adb-f0dd-4db6-8626-d5c982287beb · outbound

This paper cites DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.366022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.366022Z digest=sha256:cd90e8d188c0d6f194be5e0c08b9cefc0ff21c5778cb76cc11adbec3056f0b2a

Observation 21da1ec2-b606-4317-af1b-4c006199bc6b · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ImageInWords: Unlocking Hyper-Detailed Image Descriptions , booktitle =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.370959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.370959Z digest=sha256:6b1025fdf2aab36f5dd3294fd45e462e1685cbe004f3e1a35efdefc61f49f045

Observation 4e2ded75-9897-483d-bdaa-3244b3cc8133 · outbound

This paper cites Computer Vision -.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Computer Vision -

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.374937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.374937Z digest=sha256:64b5ac77822891286d01a4b30e6922c9a14a9b7915356fbed7120f73146c3d37

Observation 59fce482-b53e-4896-a479-f67ead3f281a · outbound

This paper cites CoRR , volume =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment CoRR , volume =

Reference 45

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.921963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.380107Z digest=sha256:86f42a23209799f783ac916e966579e25b169fff6b2d5400604f8cadf1f1860c

Observation b417e9be-da9b-45fe-ade6-3cc9286b023f · outbound

This paper cites Microsoft.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Microsoft

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.385103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.385103Z digest=sha256:7b3517df01b94da2847b0cd55122e610cc24a8350b366a4bd99b0aa64de30c37

Observation c5ddacd0-7702-4c31-95f8-6ca162d3a25d · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.389041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.389041Z digest=sha256:d1c51548a4b66cc331d5fba9e442dee37328f0ffc491eb7fa0f5de7bd564e193

Observation ce3f8042-9f48-4740-8fa3-884b0755b6c6 · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.396828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.396828Z digest=sha256:2cb3da5176262fb9266b8a10c125fda1ca1c3d7b62678ae93c68a20b0cd46dc2

Observation 9d10c119-6fdb-454c-b553-625a839350ea · outbound

This paper cites 2024 , url =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2024 , url =

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.402678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.402678Z digest=sha256:36b5be554feefe20028134d56bc82ec14e3303972d0a3482152d0687b7c9c853

Observation abbab471-c565-464f-aea3-0451272a2df6 · outbound

This paper cites Segment Anything , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Segment Anything , booktitle =

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.406419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.406419Z digest=sha256:5033d91c45091212b2444a06c29b1f05202a317caf0852b0decd6932662d2bf0

Observation bb5db8b4-3b52-4168-a186-825279c19c33 · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-12T04:57:29.410596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.410596Z digest=sha256:b403611086c60a3dc04b81ba159974eb586d7e0e0f9f6a65e75b763dccc44441

Observation 70fb8413-9ed0-4208-8925-81225e0307ff · outbound

This paper cites 2019 International Conference on Document Analysis and Recognition,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2019 International Conference on Document Analysis and Recognition,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.414823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.414823Z digest=sha256:9e42eaf6db489149c91f90713e235a823ac8788b78560193942c65b3162d0967

Observation 0be0bab2-6ab8-41c5-b7ff-f0acd559e871 · outbound

This paper cites Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations , journal =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations , journal =

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.418656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.418656Z digest=sha256:c03e20933a5f63da3b0d62c6d12c40fdfba4859a6b6cd526f4d112df83ef23c4

Observation 99e3d290-888c-4409-8749-adef6a03714e · outbound

This paper cites Price and Scott Cohen and Christopher Kanan , title =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Price and Scott Cohen and Christopher Kanan , title =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.422714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.422714Z digest=sha256:1120484ed86a69416dd7654e430113cd7ec550415f1032e8d213d52d9578606f

Observation e793620b-01d4-4fc2-8d3d-0b07548dbd58 · outbound

This paper cites OCR-Free Document Understanding Transformer , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment OCR-Free Document Understanding Transformer , booktitle =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.427268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.427268Z digest=sha256:0a1ddcf2db389b2c74815885e550ef93afef08202a8a9efdb3770fbcf79b93c7

Observation d7e7e17a-16cd-40c2-9c9c-6fce36402e1b · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:57:31.266267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.431339Z digest=sha256:809f35602600ed53adccecb46c45a3464c434e4a6de3562ec4713c03831234e9

Observation 2ff79465-f090-4877-939c-03b535a9919f · outbound

This paper cites Grounding.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Grounding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.435187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.435187Z digest=sha256:5341854a8f39bde9f0de3c94a9e22ec85f5607c5b9495b07af28c837aa4b8315

Observation d872d31f-f3b5-4831-a098-3ecb84aa09a2 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment The Thirteenth International Conference on Learning Representations,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.438815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.438815Z digest=sha256:beeb51fc7479e8e3e3bea1d9fc49eea718c32133c29dfe2575a1eac4cc765cbd

Observation 5f8c0848-552d-4c53-b33b-54ac20f4bec2 · outbound

This paper cites Forty-first International Conference on Machine Learning,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Forty-first International Conference on Machine Learning,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.246523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.446143Z digest=sha256:8b7163cade5e5fc02011128665e27c54a73991eeeb405da5d7e731e2a8504711

Observation bdaf6862-6507-4e7c-9463-79024530986e · outbound

This paper cites Hinton , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Hinton , editor =

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.234332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.450985Z digest=sha256:51ad3e66d503d906e445df0ac4ac2dfd2af767aef6b2ff203569b8b2cf0c03b3

Observation 9e0ab83a-63a3-4426-b20c-d62751cdb8e8 · outbound

This paper cites SEA : Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLM s.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment SEA : Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLM s

Reference 61

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.788892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.456006Z digest=sha256:8fa3ea864d6082878cc000725e43a6d024ce2f5e6c4fcaf87d1ff1b8e4738ec8

Observation d6ac425d-f669-4b50-ae94-43c4517bc5f7 · outbound

This paper cites Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.460876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.460876Z digest=sha256:7d77e7bf26ecc9457b7440af34e7ed0d2fc70b269a9225690f3596f916ba9238

Observation 13a8c10a-4f3b-4043-88ba-1996568f87b1 · outbound

This paper cites GroundingGPT: Language Enhanced Multi-modal Grounding Model , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment GroundingGPT: Language Enhanced Multi-modal Grounding Model , booktitle =

Reference 64

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.762427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.469014Z digest=sha256:6ba98d80ee0dc93a732c93cfd061a69547d8482b60830a22d017968c85f5f85b

Observation e4d0ef44-8a33-4824-a5f3-a1f312028ad2 · outbound

This paper cites Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

Reference 65

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.750175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.473156Z digest=sha256:4136179fd73f94f822218e99959e7df0ae3cca044026e474b2886fe5ded2cb9b

Observation d34a3720-006b-476d-9616-28a8ab2b83c5 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.477504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.477504Z digest=sha256:46eaaca4632a3721f175392f6f89771f1254a23fa465759d7ace0570dd190ff3

Observation 771488ca-ee58-44b5-878e-c446e0a4dfe8 · outbound

This paper cites ParGo: Bridging Vision-Language with Partial and Global Views , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ParGo: Bridging Vision-Language with Partial and Global Views , booktitle =

Reference 68

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.726467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.486341Z digest=sha256:b6718cf2f0ae35a17b79b1e62ed1af23080752331d11e86ff888f22e2d09d392

Observation 73485882-dd2d-4a8e-930b-2214a5f87f1f · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.490215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.490215Z digest=sha256:ecda67bce049997912db99d36dec0db067a7273b8dbcc28a591b398819bb2f67

Observation b0b25b03-1f68-4821-9d4a-7fad0d63e5f6 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Ferret: Refer and Ground Anything Anywhere at Any Granularity , booktitle =

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.222667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.494547Z digest=sha256:f02fde4a88733f4c9020c6b425d1ce33d42384c46be4e7f10c00bf325579b427

Observation 2d1c01d8-6b9d-4668-901f-e007c922e92b · outbound

This paper cites ICCV , year=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ICCV , year=

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.210661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.498269Z digest=sha256:0160b6a94df07d3701fc1285c78dc13a5398ef3cb65fb9f6c7433beae2479c81

Observation 2046d4a3-1d46-4e01-8fb0-c5a24677d9fc · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:57:31.197617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.502369Z digest=sha256:d1d77fe8d992a29ed9b7bfcc61cc449407fbb2a7a1eefa72e3761c5a0cbacf7d

Observation 9aa41628-5252-4051-b715-16d7d06c68d3 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.505985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.505985Z digest=sha256:18abf8427cd3bda28641ff3fb142520cf6af6c61e7decc35b962670dcd19edc4

Observation baf50106-d779-424c-acb3-61de245ec292 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Learning Transferable Visual Models From Natural Language Supervision , booktitle =

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.509977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.509977Z digest=sha256:50fb07eeaa19ba5b8c37b80d9ba82b267be36aa9ffa57a3b3c711d49f4ee0e83

Observation 1b9d230d-3562-416b-80aa-8a8b8503e51a · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.513508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.513508Z digest=sha256:95a726f7cf6ef0377b6cc260d86f416790eae8c154a0075e498e267088d13659

Observation 2d5521f7-924d-4bba-b023-f15988c62f71 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest , booktitle =

Reference 76

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.677408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.518355Z digest=sha256:1a7af3664009413324015ac77cd4e4939257334aca7361b04f1b136f03882f70

Observation 02cc1978-67db-47a9-bbb7-fcad49a87637 · outbound

This paper cites Shaker and Salman H.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Shaker and Salman H

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.522179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.522179Z digest=sha256:29ed6de419d1355a434ca6fa08e3fb68391301b9faae2f0678f7c70319ff4094

Observation 70b2912e-eb61-455e-8ada-155eaddd0cbc · outbound

This paper cites NExT-Chat: An.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment NExT-Chat: An

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.176905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.525748Z digest=sha256:656fd2e61c1f2f06a69243b5985b55686c4168a4a294b40a7345bc380f2d4755

Observation e9d00afa-6941-4047-8a5c-b6247aaab4f7 · outbound

This paper cites LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models , booktitle =

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.529634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.529634Z digest=sha256:7fecf250fcf639275b53d9e18e13dba6fbc218c3b66bf174c37b814488ec6638

Observation d660a6f6-af95-4a67-b46f-6defdd0d40d6 · outbound

This paper cites The All-Seeing Project.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment The All-Seeing Project

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.534030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.534030Z digest=sha256:78eefe18c4d7b3c805e93162abbe64c44b1bf01ac2922cc441c6104842153658

Observation 474a49c6-8f61-4d83-9ac5-b84e8e4d9867 · outbound

This paper cites Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training , booktitle =

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.163045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.537581Z digest=sha256:4f9779bbf3238a53f81ed31eca36c5656bf81d80c8ba768990b761ba462c401d

Observation e313d1cf-c9e1-4d48-985b-13a69e1b5aac · outbound

This paper cites PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment , booktitle =

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.541551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.541551Z digest=sha256:1bb472db36f34250258d1023812f7785bf1eea2da2c20897e432217599dc0ec5

Observation 3f70b69b-cdfc-4972-b6fc-dbdce5df0209 · outbound

This paper cites Code-Switching Curriculum Learning for Multilingual Transfer in LLMs , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Code-Switching Curriculum Learning for Multilingual Transfer in LLMs , booktitle =

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.149501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.545608Z digest=sha256:41b2de9a9ed2f08e888cce3b3331d1c2a2cf5a63f45f80b33a41fb7afa2386e2

Observation cd13adba-e00f-494f-af79-67378847e00b · outbound

This paper cites DICT-MLM: Improved Multilingual Pre-Training using Bilingual Dictionaries.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment DICT-MLM: Improved Multilingual Pre-Training using Bilingual Dictionaries

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.549323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.549323Z digest=sha256:97b2e9cf75ba3f442d1255d418333ba288f3386baef33f091419e9624d95a983

Observation d323adac-6712-4058-92e6-6a814d721a39 · outbound

This paper cites CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual

Reference 85

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.642572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.554233Z digest=sha256:74027bbf3166fa1bd44fa2c11b978420f97253ab3a0f68f14f46f36e1c26ede1

Observation c4518eaa-3692-4674-9be4-5fb1b2a65984 · outbound

This paper cites Bilingual Alignment Pre-Training for Zero-Shot Cross-Lingual Transfer.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Bilingual Alignment Pre-Training for Zero-Shot Cross-Lingual Transfer

Reference 86

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.628745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.558383Z digest=sha256:8a880e114f7336c61a2d00f9fabb4e7fa13716155e133bbd32d7d54e91330de0

Observation ed807199-313f-400c-b98a-76c76d9c7e41 · outbound

This paper cites Latino Language and Communicative Behavior , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Latino Language and Communicative Behavior , editor =

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.135163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.562649Z digest=sha256:aeda9a235e1816acca132004015960d779043149650208bcc820f5a0ea1755cf

Observation 3626a8b2-33ad-484e-9e91-26d287c368d9 · outbound

This paper cites Thara and Prabaharan Poornachandran , title =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Thara and Prabaharan Poornachandran , title =

Reference 88

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T04:57:30.364833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.567258Z digest=sha256:4ae8187b20e206cb6b05d60afe0e4097c794a5b63ecbba1dde84a4c2ede459c7

Observation 0f9fbaa1-9520-482c-b3a1-b903594a379a · outbound

This paper cites Alabdulmohsin and Avital Oliver and Piotr Padlewski and Alexey A.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Alabdulmohsin and Avital Oliver and Piotr Padlewski and Alexey A

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.121851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.571385Z digest=sha256:5c30f680c537610fddc5b28e3730a8ce3045cc7354d35ccbf68d6d04810700e2

Observation 0fb93f31-adc7-44a4-9fd3-175496eb9a09 · outbound

This paper cites Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.575127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.575127Z digest=sha256:2d7b7a957e201b1f9067613bdf825060026dd2ae16e38be96ac1185e8346ed84

Observation aa44995f-2e8f-4629-81e1-0efa883a7051 · outbound

This paper cites 2025 , howpublished =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2025 , howpublished =

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.579562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.579562Z digest=sha256:29dbf98db18ec12ebd612527561d64206ecf881f67e7002350fea80366d8be73

Observation e8ec4d31-14bf-4d50-ae36-e862ddc6b724 · outbound

This paper cites 18th International Conference on Pattern Recognition.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 18th International Conference on Pattern Recognition

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.583850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.583850Z digest=sha256:ef8f468f892a5503541d9a15e8e98e8b9a925fbda3fe0e4cff67b8d5ed3840e1

Pith citing papers

No inbound Pith citation observations are available.