Pith. sign in

Paper Citation Record · LEDGER

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment

As of 15 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2608.11167.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11167 v1

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:57:29.583850Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact9
  • verified fuzzy12
  • unresolved64
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 565c71e8-44ee-40a7-bc9c-ca5e5cff1448 · outbound

This paper cites Visual Instruction Tuning , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Visual Instruction Tuning , booktitle =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.196786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.196786Z digest=sha256:28916b77077e926eae61cba6e0e28e958c10a806dc8aa6e1885b340638a08e34

Observation 885c6a52-3474-490d-95a4-3e54da950b4a · outbound

This paper cites LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LLaVA-NeXT: Improved reasoning, OCR, and world knowledge , url=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.205718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.205718Z digest=sha256:b0a5de7806161250f178b48c0e3325c8d85a61b5e7598743a1cdba522b2cdb12

Observation 4b2b103e-19f2-4ef5-8ddd-af34039fd017 · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.210383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.210383Z digest=sha256:2caf3f12fffbf83403266679f4dcb08903b3538ddb641dc052e549dc072a1348

Observation 0bc2528a-21ed-49d6-afd1-97a4894d6f61 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.214140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.214140Z digest=sha256:46e61c96a8665380492b7e15e5c5ab69d6b3a5119234d1c8077114d0c7af0f46

Observation 754d5abb-09e1-49bb-a305-b31610cae958 · outbound

This paper cites Qwen2.5-VL Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen2.5-VL Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.218561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.218561Z digest=sha256:2052db0496cef99ccb58bdb028f72d00de5506fc716892cd315493225de8c7ec

Observation a3f5ab58-210a-4f26-be6b-b9e5e7637237 · outbound

This paper cites 2025 , eprint=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2025 , eprint=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.223012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.223012Z digest=sha256:c9138c339ccc8f2a3221c4d83ab24a5ecc6727af131ab0c6dadef7c4718a891c

Observation ec6fb014-cc3f-411c-97ab-3317bc25fe51 · outbound

This paper cites MiMo-VL Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment MiMo-VL Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.230150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.230150Z digest=sha256:449c2cef480bc3cbd4b80bde2bf541d508c260e00cfe90a8b462adee1c262e22

Observation c498bd22-1e36-48e9-8295-3b69342f6ca9 · outbound

This paper cites 2025 , eprint=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2025 , eprint=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.234182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.234182Z digest=sha256:83fbad74f5b0582d496c575e16e1e3c8bb1d15e98ac84b1fbff44955798d8d9e

Observation 56cd082d-a781-48ff-a7b9-25ea335c561a · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.237647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.237647Z digest=sha256:623a96a8e10c7d5b59520217427b136966f71e1ddd87cba6aab6890add3404cb

Observation 9700f8ad-d9d6-4c9c-9ab2-e256ead06748 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.241637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.241637Z digest=sha256:f000bf13e4d67a914938407bdb39644b067dc30b38280dc768c501a9c30c0cb5

Observation a98493a4-1f8a-4c78-8c33-d5e957b981c8 · outbound

This paper cites Scalable Vision Language Model Training via High Quality Data Curation , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Scalable Vision Language Model Training via High Quality Data Curation , booktitle =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.245635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.245635Z digest=sha256:2763912d1af4308c48d1672ba6d831a3e42160378a769f621819c8a58921bd6b

Observation f4742661-3eac-49a7-96af-d507aed679e0 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Flamingo: a Visual Language Model for Few-Shot Learning , booktitle =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.249132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.249132Z digest=sha256:a8fcd92a6c30bc26b430680f21aa0dce932334f141bbd500a231432731a43ec5

Observation 5e997eb9-1d21-4376-aba1-c8cb9b751f88 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.253141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.253141Z digest=sha256:1fd8ffa25806cfd33e47de59efb3aa7c8913450acd3966a1f359cafa7089055d

Observation a1ca3212-96f5-4120-8be9-32b9fd38b029 · outbound

This paper cites Seed1.5-VL Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Seed1.5-VL Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.257406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.257406Z digest=sha256:badaf114179e5c93cf12caa485fee1873be06d12cc1ed2ddf57b325b7318a942

Observation 606d72bd-1521-41af-86f8-ed8e16e64fee · outbound

This paper cites CoRR , volume =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment CoRR , volume =

Reference 16

Resolution
verified exact
doi, observed 2026-08-12T04:57:30.180987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.261748Z digest=sha256:2dfc4a8db48fbb6c77915a1a285433638e2f1875a0569414756faec7b39e49dd

Observation ea48c056-da6f-48a1-ab50-01bb588aa137 · outbound

This paper cites Computer Vision -.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Computer Vision -

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.265494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.265494Z digest=sha256:b505f1df686e7718a3378b2b783eb06b30404d700704c28bc6ee107a0cab16c4

Observation a47e8c63-bfea-4588-8ce7-f1fe15e5d248 · outbound

This paper cites Qwen2.5 Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen2.5 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.269711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.269711Z digest=sha256:c2af1a52c7f39e0c07db97f1dcb955395c38bbb32dba70e84a8c5ae79523012d

Observation de5b4bd7-3978-4b42-83fc-81c5c01d00f2 · outbound

This paper cites Qwen3 Technical Report.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Qwen3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.273623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.273623Z digest=sha256:fc293dcfcb056a2e5ca9f09fc9a0c5e7858ff718fa720d7632b6e0525db24411

Observation 8679e8f7-be7a-4ec3-9f82-ee1df8f835b9 · outbound

This paper cites The Llama 3 Herd of Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.277639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.277639Z digest=sha256:58bfdb53d73dd49576bcb1d8da121e7c084022248e74df0abf3e726379a26a26

Observation cd8b26b5-f973-4828-8103-9c4e6686aec1 · outbound

This paper cites Making the.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Making the

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.281596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.281596Z digest=sha256:827e81d6eb72ecb736d2508238367b1f9d6f1837f957f3bc3872f655dd75b4e4

Observation 7843997f-c70e-4364-835e-31a78e905f08 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models , booktitle =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.285486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.285486Z digest=sha256:71721e50d4b5313074ef4eb4da98d105764c8d33722e3d65b2d52bac24f7635f

Observation 8fb8e444-3744-443c-9fa4-582fcb27a03e · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-Around Player? , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment MMBench: Is Your Multi-modal Model an All-Around Player? , booktitle =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.289625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.289625Z digest=sha256:df1c1a922bfab4a2a9391fa0f75ab5fcd50200634fe5310d564d5ab3998d5c75

Observation f478145a-65b0-4b2f-92cb-dec2857d45c9 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.294805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.294805Z digest=sha256:a6dcb0540608762de9f387a0cd03746794b7a14c16c74f33898266f3818aaa36

Observation cc75e00e-fa7a-443b-9502-7c12dd2e591b · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models? , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Are We on the Right Way for Evaluating Large Vision-Language Models? , booktitle =

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.299431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.299431Z digest=sha256:0f69db80a064de2f07c2ca3e9991005cf5e1dfd467309ad029cf65c69a50ae43

Observation b6f4399b-c5c1-438f-832f-2d47f8d26dc6 · outbound

This paper cites Forty-first International Conference on Machine Learning,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Forty-first International Conference on Machine Learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.301283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.307220Z digest=sha256:5c7192b9b0b1f55fecc24ff77e48d429ec6ca3fc2b006a7a1b47c3c77d0a235e

Observation 92af91a1-ab07-4a2f-bbd3-d324f400ce42 · outbound

This paper cites Hudson and Christopher D.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Hudson and Christopher D

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.311156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.311156Z digest=sha256:5eace34e1673f3e6152012a37117cead539245a3cb75844b7628c306b7474a8b

Observation b2889eba-9548-4e16-b9c6-e9ec1d9db9fb · outbound

This paper cites A Diagram is Worth a Dozen Images , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment A Diagram is Worth a Dozen Images , booktitle =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.314674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.314674Z digest=sha256:86779132247533b1b749fcf8f283127f76b95858361530f4a2812768fac25f44

Observation 5d21991e-be69-43e5-b0e1-c1de80f27ee0 · outbound

This paper cites Joty and Enamul Hoque , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Joty and Enamul Hoque , editor =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.318875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.318875Z digest=sha256:d92aed6a8406f0acd431a52e0175394bf5efe2af31760d220c75761efe4d8ca6

Observation 966b00db-cd54-4d65-85b5-0de2840a07ed · outbound

This paper cites Cambrian-1:.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Cambrian-1:

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.289350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.322900Z digest=sha256:a863dc29127f9acb371451cf5bb438c5b452700abdabf380f268fe39409b4a2f

Observation e8eb2ba9-63e7-4ac4-aa06-97e6ee4e2dfd · outbound

This paper cites OCRBench: on the hidden mystery of.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment OCRBench: on the hidden mystery of

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.327016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.327016Z digest=sha256:53621079eda1d905847399a52b42b5b5dee4678c55209a6994b0d9cc6fa4dc89

Observation 0bb53f02-c088-4c53-9237-85ab45e99862 · outbound

This paper cites 2019 , url =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2019 , url =

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.330804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.330804Z digest=sha256:ae0f36e4b2be06291046658f4f7e8543cf1a9ac09e8fde510c439fd1af6b885e

Observation 8d61bc53-0d30-4d5b-8979-cb784e26e15d · outbound

This paper cites Berg , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Berg , editor =

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.337861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.337861Z digest=sha256:5f0cd97ed9addd4e28a7cd8fb935806edfc12ace8d3aa127497e21253e9b03c8

Observation 438c1bf1-98ec-4d3d-8c16-65f9f746a306 · outbound

This paper cites Yuille and Kevin Murphy , title =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Yuille and Kevin Murphy , title =

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.341502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.341502Z digest=sha256:565b883908ec92e5ede0e8d3f0ecc2a1af83d9f5008d4eaae6ca31595f566cce

Observation e74cd858-eb1d-44c7-9060-f68719cf461e · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.345672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.345672Z digest=sha256:78016d96064e06ff3c70a2a64625ca3b14936fa23311c70bcf3944f1991e83ca

Observation 205dfa8c-497f-424e-a738-adee591aca4e · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models , booktitle =

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.349323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.349323Z digest=sha256:1480b3a36dd864a9324839d06cebe82b1b012a5d0bf4f057fa27eea244587128

Observation 45e93531-70cd-4e47-bbed-9e0b49aca790 · outbound

This paper cites ShareGPT4V: Improving Large Multi-modal Models with Better Captions , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ShareGPT4V: Improving Large Multi-modal Models with Better Captions , booktitle =

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.352991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.352991Z digest=sha256:0dfb2e27fc8e7a0709e0c01251a132dc090ac1f48be4a0b51e299f988e12abcd

Observation 3cec23b2-d15e-479f-8f48-d39036e9083a · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.356681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.356681Z digest=sha256:37ec7b7f465b62a64aa81ec2a4b590604b292ca0f9dddd25d5ace8dd6fe90544

Observation cef80f05-23ca-4fd2-b9f9-f0cdad2e25d2 · outbound

This paper cites DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception , booktitle =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.277887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.360560Z digest=sha256:238dcad9d378d3f8265515240bd92fdeadb9989b51cd1232a993dac9390c90e2

Observation 89047adb-f0dd-4db6-8626-d5c982287beb · outbound

This paper cites DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.366022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.366022Z digest=sha256:d2affba3647b63dac7430090acad08487781e159ae794ca4facb7bcca9de75b9

Observation 21da1ec2-b606-4317-af1b-4c006199bc6b · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ImageInWords: Unlocking Hyper-Detailed Image Descriptions , booktitle =

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.370959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.370959Z digest=sha256:18b8f4b25271cabfca2181b193aa4eccbb849b8dfe60fee39d905bc317135446

Observation 4e2ded75-9897-483d-bdaa-3244b3cc8133 · outbound

This paper cites Computer Vision -.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Computer Vision -

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.374937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.374937Z digest=sha256:d90a4090f3a9237342e364848b04088496406edc4b0d8b64a823b3af59ee161e

Observation 59fce482-b53e-4896-a479-f67ead3f281a · outbound

This paper cites CoRR , volume =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment CoRR , volume =

Reference 45

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.921963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.380107Z digest=sha256:553d2db771808e7b1214229ff7a113ee9514336b23ef23c5f5b66344be1f9dd0

Observation b417e9be-da9b-45fe-ade6-3cc9286b023f · outbound

This paper cites Microsoft.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Microsoft

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.385103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.385103Z digest=sha256:d43b8588a988dfd75896cdb2e4f021862a1db32715599976f370e03c3fc6a301

Observation c5ddacd0-7702-4c31-95f8-6ca162d3a25d · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.389041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.389041Z digest=sha256:439e43908d30d757e263e96bbf3b2af0616ee6136c9bf1a4c8d16f07a01390e4

Observation ce3f8042-9f48-4740-8fa3-884b0755b6c6 · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.396828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.396828Z digest=sha256:6e5f7e85b5e91790086540b2b22efff4cfefea91c2298de77650f57a84208af2

Observation 9d10c119-6fdb-454c-b553-625a839350ea · outbound

This paper cites 2024 , url =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2024 , url =

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.402678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.402678Z digest=sha256:c5e355ca4de0d28e1bfe0a7d77be4976013aa6a1ab45fb179be3f690f60c9897

Observation abbab471-c565-464f-aea3-0451272a2df6 · outbound

This paper cites Segment Anything , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Segment Anything , booktitle =

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.406419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.406419Z digest=sha256:c9239399fcbf4a220f68b8a7ac58f3a9a3a6ecafcd77e278193c8ad129bc0ea9

Observation bb5db8b4-3b52-4168-a186-825279c19c33 · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 51

Resolution
malformed identifier
no resolver link, observed 2026-08-12T04:57:29.410596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.410596Z digest=sha256:1ebc01433f92c0005e264723e0db9ada5c4ca718d678b159ae0bb7eda9483d77

Observation 70fb8413-9ed0-4208-8925-81225e0307ff · outbound

This paper cites 2019 International Conference on Document Analysis and Recognition,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2019 International Conference on Document Analysis and Recognition,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.414823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.414823Z digest=sha256:0436447eae4692ad4e9989a9a1b5d0f191b1459c9ee24b2fd37c547ac1d27e48

Observation 0be0bab2-6ab8-41c5-b7ff-f0acd559e871 · outbound

This paper cites Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations , journal =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations , journal =

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.418656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.418656Z digest=sha256:1dabe39112183ba7a50df9192755ad816d7aab107c0d783893df7f440e75a94a

Observation 99e3d290-888c-4409-8749-adef6a03714e · outbound

This paper cites Price and Scott Cohen and Christopher Kanan , title =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Price and Scott Cohen and Christopher Kanan , title =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.422714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.422714Z digest=sha256:ab9724d586791acd6dae657af9f724da909e9ee06bdec58245bd9ba37fd4e8e1

Observation e793620b-01d4-4fc2-8d3d-0b07548dbd58 · outbound

This paper cites OCR-Free Document Understanding Transformer , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment OCR-Free Document Understanding Transformer , booktitle =

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.427268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.427268Z digest=sha256:d3d5f6fb2cc867462aae4a89aaf5acdf99a46c0ba2633f06b5880c3c128afe7e

Observation d7e7e17a-16cd-40c2-9c9c-6fce36402e1b · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:57:31.266267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.431339Z digest=sha256:dd98d87d189d17f321c1371d1f8155a87936054472c9ccc4dd0bf5c7f404c42a

Observation 2ff79465-f090-4877-939c-03b535a9919f · outbound

This paper cites Grounding.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Grounding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.435187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.435187Z digest=sha256:13f59dded11799994813f487d54590f1c3f775ffd36628b43fae22aee338ffda

Observation d872d31f-f3b5-4831-a098-3ecb84aa09a2 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment The Thirteenth International Conference on Learning Representations,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.438815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.438815Z digest=sha256:d5c82f5a513c9987538d059770372e24bf33a7e9f90952fc2e25bebe6a9f892e

Observation 5f8c0848-552d-4c53-b33b-54ac20f4bec2 · outbound

This paper cites Forty-first International Conference on Machine Learning,.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Forty-first International Conference on Machine Learning,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.246523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.446143Z digest=sha256:d9322c476fc50993d2dcde9fe73ee4dacde4c18fefda5982b2be639bccdd8654

Observation bdaf6862-6507-4e7c-9463-79024530986e · outbound

This paper cites Hinton , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Hinton , editor =

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.234332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.450985Z digest=sha256:13004968a32c37f8d2eeeb96a20fd4a253520cf0bf943e3c441b63fb9e6594cc

Observation 9e0ab83a-63a3-4426-b20c-d62751cdb8e8 · outbound

This paper cites SEA : Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLM s.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment SEA : Supervised Embedding Alignment for Token-Level Visual-Textual Integration in MLLM s

Reference 61

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.788892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.456006Z digest=sha256:673dc8765ecd9a14ff067d8a6386816e08c5598cdb693f2858721f5cbb9a8ddd

Observation d6ac425d-f669-4b50-ae94-43c4517bc5f7 · outbound

This paper cites Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.460876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.460876Z digest=sha256:be0961b64af67a4d6929e6c6c2faa482372e1b558bde140dcdb4da3522f0cf5d

Observation 13a8c10a-4f3b-4043-88ba-1996568f87b1 · outbound

This paper cites GroundingGPT: Language Enhanced Multi-modal Grounding Model , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment GroundingGPT: Language Enhanced Multi-modal Grounding Model , booktitle =

Reference 64

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.762427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.469014Z digest=sha256:1053412b18a513c3c470a27526e6e6b70234b0f44ea30228a61a3a40869163f6

Observation e4d0ef44-8a33-4824-a5f3-a1f312028ad2 · outbound

This paper cites Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models

Reference 65

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.750175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.473156Z digest=sha256:21cc44eb59fec0ecf6256b1396f8802a28361adbe1b65e7e1fa621e11cfe2063

Observation d34a3720-006b-476d-9616-28a8ab2b83c5 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.477504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.477504Z digest=sha256:5e0a354ee72a95ddf0e41bda3f6004a05e603b59b48f5030def1cc203aa22359

Observation 771488ca-ee58-44b5-878e-c446e0a4dfe8 · outbound

This paper cites ParGo: Bridging Vision-Language with Partial and Global Views , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ParGo: Bridging Vision-Language with Partial and Global Views , booktitle =

Reference 68

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.726467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.486341Z digest=sha256:94139994d1560c5cc3c7f2e9d7366e88cb2ff58529f2feaef71e3e33d13fde39

Observation 73485882-dd2d-4a8e-930b-2214a5f87f1f · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.490215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.490215Z digest=sha256:c30db495703cfc62d1e03fe67d5eb5582faa4db3189d6f53f0d40191dc3371d2

Observation b0b25b03-1f68-4821-9d4a-7fad0d63e5f6 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Ferret: Refer and Ground Anything Anywhere at Any Granularity , booktitle =

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.222667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.494547Z digest=sha256:efc1e1dcd7016c8ea79663bd0acb6d91b66e4826f0e256b3f1eaceaa4e189ee2

Observation 2d1c01d8-6b9d-4668-901f-e007c922e92b · outbound

This paper cites ICCV , year=.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment ICCV , year=

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.210661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.498269Z digest=sha256:45aa67616ffc82d986f247c8418cd692dcccf844e8b1cf615a7ad617b898b66d

Observation 2046d4a3-1d46-4e01-8fb0-c5a24677d9fc · outbound

This paper cites an unresolved cited work.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-12T04:57:31.197617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.502369Z digest=sha256:79868afe0a090eb05c31d5d841ebef9992c88db342563a55601fc470e7567575

Observation 9aa41628-5252-4051-b715-16d7d06c68d3 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.505985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.505985Z digest=sha256:c2c1bd2ec1c38a26e36eba0c61b249e356425c365287c9af4349b2a57e5019c6

Observation baf50106-d779-424c-acb3-61de245ec292 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Learning Transferable Visual Models From Natural Language Supervision , booktitle =

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.509977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.509977Z digest=sha256:60af734e787a742c503246f11a464ea63e57566bb2592da089b557b417159057

Observation 1b9d230d-3562-416b-80aa-8a8b8503e51a · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.513508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.513508Z digest=sha256:9e416fd2c570b0880be03b80b37c54eec15315abd1ed2a2f4c5edc2817bf7142

Observation 2d5521f7-924d-4bba-b023-f15988c62f71 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest , booktitle =

Reference 76

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.677408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.518355Z digest=sha256:eae93854604579739dcdcd0b11f8e1fc8d0d29ff859a45918fa0849580d722f3

Observation 02cc1978-67db-47a9-bbb7-fcad49a87637 · outbound

This paper cites Shaker and Salman H.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Shaker and Salman H

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.522179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.522179Z digest=sha256:48ac2880a0e4640b1ba48a8002f2e30963cd62c468d73902885e93f523b0c34a

Observation 70b2912e-eb61-455e-8ada-155eaddd0cbc · outbound

This paper cites NExT-Chat: An.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment NExT-Chat: An

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.176905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.525748Z digest=sha256:f66b807aed74ec6177fd44e22d06a9c5db1eeb12a4969f850a838984eeb1a4bc

Observation e9d00afa-6941-4047-8a5c-b6247aaab4f7 · outbound

This paper cites LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models , booktitle =

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.529634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.529634Z digest=sha256:f3b6898fa75dca1c2990947995d416694f1ad459a9b8218802e9fc252224c7f3

Observation d660a6f6-af95-4a67-b46f-6defdd0d40d6 · outbound

This paper cites The All-Seeing Project.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment The All-Seeing Project

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.534030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.534030Z digest=sha256:cdc67fd31ac255a7e6eef2aa3d8443b1ecf4030c7135c0cc1c1d40af0305f01c

Observation 474a49c6-8f61-4d83-9ac5-b84e8e4d9867 · outbound

This paper cites Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Investigating and Scaling up Code-Switching for Multilingual Language Model Pre-Training , booktitle =

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.163045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.537581Z digest=sha256:140f7ddaf3592a16c5dd3d225445ec359f363520f847d9474b68e1d8193f7d16

Observation e313d1cf-c9e1-4d48-985b-13a69e1b5aac · outbound

This paper cites PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment PreAlign: Boosting Cross-Lingual Transfer by Early Establishment of Multilingual Alignment , booktitle =

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.541551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.541551Z digest=sha256:473d4695bc2dede8d9bf22f9d43f826d89b115e755c71f941ebbbc8d556080f8

Observation 3f70b69b-cdfc-4972-b6fc-dbdce5df0209 · outbound

This paper cites Code-Switching Curriculum Learning for Multilingual Transfer in LLMs , booktitle =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Code-Switching Curriculum Learning for Multilingual Transfer in LLMs , booktitle =

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.149501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.545608Z digest=sha256:b4110028d098fefa170c7a8ef77f476264147332d56fb864e6cb84d69e6b41be

Observation cd13adba-e00f-494f-af79-67378847e00b · outbound

This paper cites DICT-MLM: Improved Multilingual Pre-Training using Bilingual Dictionaries.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment DICT-MLM: Improved Multilingual Pre-Training using Bilingual Dictionaries

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.549323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.549323Z digest=sha256:463531ede8abd656898200da377fcca07a90701ee59b96686050146179ba793f

Observation d323adac-6712-4058-92e6-6a814d721a39 · outbound

This paper cites CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment CoSDA-ML: Multi-Lingual Code-Switching Data Augmentation for Zero-Shot Cross-Lingual

Reference 85

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.642572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.554233Z digest=sha256:5a43bc73030aeb031d44ee8e88f2a3d37ff22eb291f5bd4ea1af8647b1aa40fe

Observation c4518eaa-3692-4674-9be4-5fb1b2a65984 · outbound

This paper cites Bilingual Alignment Pre-Training for Zero-Shot Cross-Lingual Transfer.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Bilingual Alignment Pre-Training for Zero-Shot Cross-Lingual Transfer

Reference 86

Resolution
verified exact
doi, observed 2026-08-12T04:57:29.628745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.558383Z digest=sha256:169d0b9b71fc2a495bfcb963560a4b5a31c836e47f415277b09d48637a820d50

Observation ed807199-313f-400c-b98a-76c76d9c7e41 · outbound

This paper cites Latino Language and Communicative Behavior , editor =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Latino Language and Communicative Behavior , editor =

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.135163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.562649Z digest=sha256:f1a8bc5c3cf89cfc16cbe5bf05e71106673f7884a18924312f50610f8d7ae6e3

Observation 3626a8b2-33ad-484e-9e91-26d287c368d9 · outbound

This paper cites Thara and Prabaharan Poornachandran , title =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Thara and Prabaharan Poornachandran , title =

Reference 88

Resolution
metadata mismatch
raw_fallback, observed 2026-08-12T04:57:30.364833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.567258Z digest=sha256:0045f6d8b38098fb0c570efb6a300ff59ba58a624e0bb834c4efb65078760724

Observation 0f9fbaa1-9520-482c-b3a1-b903594a379a · outbound

This paper cites Alabdulmohsin and Avital Oliver and Piotr Padlewski and Alexey A.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Alabdulmohsin and Avital Oliver and Piotr Padlewski and Alexey A

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:57:31.121851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T04:57:29.571385Z digest=sha256:0ee18593fed3d73dfbeb70e5435496671f0f338c4b0f8c4121d8aefa632317a6

Observation 0fb93f31-adc7-44a4-9fd3-175496eb9a09 · outbound

This paper cites Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment Hu and Yelong Shen and Phillip Wallis and Zeyuan Allen

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.575127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.575127Z digest=sha256:7cc53e42c18f1c771b7d2ba11f7c50c344fc7b789cdb0939fb695735134fd4fb

Observation aa44995f-2e8f-4629-81e1-0efa883a7051 · outbound

This paper cites 2025 , howpublished =.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2025 , howpublished =

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.579562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.579562Z digest=sha256:07a8829fee20b33567a38e69dcd17b9ffa6cb2490a72b1af712c4e3dc415f1fc

Observation e8ec4d31-14bf-4d50-ae36-e862ddc6b724 · outbound

This paper cites 18th International Conference on Pattern Recognition.

MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 18th International Conference on Pattern Recognition

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T04:57:29.583850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:57:29.583850Z digest=sha256:1f61cb3fb4a830101feed5620be32a4ca61f17503c01def427caddc12689edb5

Pith citing papers

No inbound Pith citation observations are available.