Pith. sign in

Paper Citation Record · LEDGER

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction

As of 16 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2608.12209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.12209 v1

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:17:33.970322Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact2
  • verified fuzzy39
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e27eee67-71f9-4100-8bcb-157b5ef51e24 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LLaVA-OneVision: Easy Visual Task Transfer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.494370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.494370Z digest=sha256:da9b578c77f583825b87f4153759c3cee1c8477bd29cddba40820928c8c24ac7

Observation 73daf702-45dc-4ba5-9b8c-c00c0f426859 · outbound

This paper cites Qwen3-VL Technical Report.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen3-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.499723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.499723Z digest=sha256:f2792bdbd3e003367992c36c6167813007f63b2bf86edc5ce268557c41e86307

Observation 581d4fef-c86b-4679-892c-e199b0988514 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.503957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.503957Z digest=sha256:7df551c27cf44c6243cbd68a2950bf660cc880d46ff27b9ec2be94cea059d5ec

Observation 690036a4-7e49-4628-8a91-54aefacf29ac · outbound

This paper cites Visual Instruction Tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visual Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.508431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.508431Z digest=sha256:bf0ae59712b25e5b00a2e6bf67f392394a2ca85ab6040b8bed156a535738e3a0

Observation 073cf827-4cbb-4837-ab8f-b899bd3e0298 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Learning transferable visual models from natural language supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.513081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.513081Z digest=sha256:e208f0931905de76f52f865316f8ca72e2475190b723d90b9817bf029c18724c

Observation 66e1153a-e3fe-4ab7-a84d-c0a1ec0e8aba · outbound

This paper cites Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visual cot: Advancing multi-modal language models with a comprehensive dataset and benchmark for chain-of-thought reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.517459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.517459Z digest=sha256:5b51d7a78909d6abe7e2fb8a3bbd54f0f2d064099ee5a90cdb94eb35812e271d

Observation 48223edc-b902-4f01-9d60-bb73da0595d2 · outbound

This paper cites An empirical anal- ysis on spatial reasoning capabilities of large multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction An empirical anal- ysis on spatial reasoning capabilities of large multimodal models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.522320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.522320Z digest=sha256:c9667398f4359fc7ca78b54bbed29e80f9e48cc81a2c603a2a4f4021d5f3b6af

Observation 9dcd2796-7c6f-470c-b83f-a84dce81d3b9 · outbound

This paper cites Why are visually-grounded language models bad at image classification? Advances in Neural Information Processing Systems, 37:51727–51753, 2024.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Why are visually-grounded language models bad at image classification? Advances in Neural Information Processing Systems, 37:51727–51753, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.537516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.526218Z digest=sha256:9b904d2db25aa5025b0716593b8cab2cd7067ace1b0fb706e2d1e8e1c3f215a5

Observation 56769107-c26a-4d69-8ea9-0a62b49efece · outbound

This paper cites Question aware vision transformer for multimodal reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Question aware vision transformer for multimodal reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.521970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.529994Z digest=sha256:52bd1baa8d9d0da67a9e9887c46952ac6068b6354f78d164b10d0967abd7ad5e

Observation d29ae9be-8ff3-4050-9b9a-60a890622529 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.534077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.534077Z digest=sha256:594013928baddade0d39a13e62142ebf5a7a1454718731a2b972b58addacbd1e

Observation 8111b51c-97ea-491e-bf8a-5ac7e5ac0c30 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Show-o: One single transformer to unify multimodal understanding and generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.505930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.539092Z digest=sha256:cf2e338f77dab979a6fa14f4643325feedee731861f93d74f9f4d5f7a2b0ad84

Observation 1b064072-c6b8-46a6-a37c-2e788fcd5d72 · outbound

This paper cites Transfusion: Predict the next token and diffuse images with one multi-modal model.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Transfusion: Predict the next token and diffuse images with one multi-modal model

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.490249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.544262Z digest=sha256:28ac5608201ce8c7e4a26a470e7e8f55361dbaa890083ae9eaa13088d03d2af9

Observation 5357a682-5d79-41d5-9b67-aa6be7d681e1 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Emerging Properties in Unified Multimodal Pretraining

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.553581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.553581Z digest=sha256:029c0e6d7b5180f6d928422323c17149160faa6e9656d27660c2a13a7e7797ea

Observation 3f0b369f-d2e4-4afc-9563-01b08ee21b37 · outbound

This paper cites Mammothmoda2: A unified ar-diffusion framework for multimodal understanding and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mammothmoda2: A unified ar-diffusion framework for multimodal understanding and generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.558435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.558435Z digest=sha256:b1199ecbdc6c7e4d2e881eff2aa5b4b7fc0574ccf9661ea1936411572430ce55

Observation 569f7263-26fc-45b3-8025-46d8129b94ac · outbound

This paper cites Lance: Unified Multimodal Modeling by Multi-Task Synergy.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Lance: Unified Multimodal Modeling by Multi-Task Synergy

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.562911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.562911Z digest=sha256:8fd2af3396a39e0c74577e51f708bd494e2e75121fdd39def87d6c01ffa8b023

Observation 63003f27-2aa1-4654-9c49-363a64cad2c1 · outbound

This paper cites Multimodal learning with next-token prediction for large multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Multimodal learning with next-token prediction for large multimodal models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.567839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.567839Z digest=sha256:fb6bf9b0f2884d1ad8a44a4af2196d4f0620027c00786387de13ed37f41cf883

Observation 84770594-92dd-4d11-8867-6f6eccdd453c · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.572476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.572476Z digest=sha256:9cb17dbc205aa1e732126ea10c391497bc383e5047571c19c0cd187e11652a45

Observation 5da4b6b3-7fb0-4f5a-befe-3a99babd2c83 · outbound

This paper cites Mmada: Multimodal large diffusion language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mmada: Multimodal large diffusion language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.447949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.577175Z digest=sha256:cf6fb5c3457e2371ea525da359240be52aa76670adc4f15bcc0db6d9d1c2ea08

Observation 75ccea02-af2c-4f5b-a6b5-5904a0f2c0b2 · outbound

This paper cites Longcat-next: Lexicalizing modalities as discrete tokens.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Longcat-next: Lexicalizing modalities as discrete tokens

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.581554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.581554Z digest=sha256:93517f8d246a91d8274e22e34088cd35b4f499b75f7ce54686d64caf9ed95b51

Observation 64c0448f-797f-402e-b382-d4c47a23a458 · outbound

This paper cites Next-embedding prediction makes strong vision learners.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Next-embedding prediction makes strong vision learners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.585968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.585968Z digest=sha256:0cc1049b032f72d2f8d06ad0414f530c37bbb29034ef9c210fdb136a6cb7a211

Observation 37c21e1c-d592-4b6d-ab4d-9da81a50ac3f · outbound

This paper cites Unihetero: Could generation enhance understanding for vision-language-model at large data scale? arXiv preprint arXiv:2512.23512, 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unihetero: Could generation enhance understanding for vision-language-model at large data scale? arXiv preprint arXiv:2512.23512, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.590538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.590538Z digest=sha256:5df35cfeeccd7bda2977b0f3cc67295d0256c6da8eef0ae0e17e54433744311a

Observation 942b4361-6ee5-44ce-ba8e-a87358e57908 · outbound

This paper cites UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:17:34.447253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.594876Z digest=sha256:fb307c40b9bec51049b0f74e89aa2ac169d4cccef90298a62aa347070f9a25d9

Observation 0c30f67c-517b-4927-81ba-45a7ce43b7b8 · outbound

This paper cites Uni-x: Mitigating modality conflict with a two-end-separated architecture for unified multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Uni-x: Mitigating modality conflict with a two-end-separated architecture for unified multimodal models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.599748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.599748Z digest=sha256:bd1a3b5f553fd90f994e4f656d714e85b2c525483611cbf1d6f61f1c87512607

Observation 6985abff-97f7-4ef9-9bb6-0f31241ac41d · outbound

This paper cites Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mixture-of-transformers: A sparse and scalable architecture for multi-modal foundation models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.431635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.604319Z digest=sha256:b87932706d59ef410e7a7e5bf65b0307445a704967a2ebe1916ec5a9926d79ee

Observation 1dfc456c-42b0-4591-ae33-cf1c8703e850 · outbound

This paper cites Qwen3 Technical Report.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.608967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.608967Z digest=sha256:9beadef04eb8b1951747b9dfb472dbd301579227d75fdac3f0714a08a7cb82dc

Observation 7ef9f9ac-56f1-4129-9351-ddfa7291dc70 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Adam: A Method for Stochastic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.613869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.613869Z digest=sha256:8668f6c5154091feb413ee2a92fad649806b9fd8ff19e6c921fc95c9e47927eb

Observation d32e69c6-17cd-44c4-abf9-8a2e7f63c81b · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Zero: Memory optimizations toward training trillion parameter models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.415500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.618475Z digest=sha256:203932e6a9ae72e8e3f38e1300fc2a6702ec1b80db760f24d97f6c6069c1c40c

Observation 41ae2848-0ec8-47b1-90a8-5f6a04382004 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.399745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.622968Z digest=sha256:257a519b3b555faf08b4fd46b96ebc9d169eaf6b0b1c70e81105954fac847fb4

Observation 183cc7c9-3aab-49f4-9be4-b059d9ce5cfd · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.384039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.627474Z digest=sha256:ef89202b2c9539841c7086b83864afc30a9d2daaddaa7309e78c2b6ae8622bf1

Observation f66b54ef-2f8e-4fa1-b366-dae59ce38d49 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Blink: Multimodal large language models can see but not perceive

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.632306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.632306Z digest=sha256:1a1a0b386b0fe4dece71ec823b5905badfd35b5b7a3c63cb736bd8fda5284f81

Observation d1930387-009c-43db-90c4-b86a88be5003 · outbound

This paper cites Charxiv: Charting gaps in realistic chart understanding in multimodal llms.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Charxiv: Charting gaps in realistic chart understanding in multimodal llms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.636732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.636732Z digest=sha256:a0254a1424104a98dc1593781caaffffe5be6665f02b18d8f7b899a81be8e368

Observation 39e60b28-eb04-4718-9a89-3abff8768d58 · outbound

This paper cites Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Dynamath: A dynamic visual benchmark for evaluating mathematical reasoning robustness of vision language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.641183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.641183Z digest=sha256:098d680b79060ee9fae50a0217e9e8aa7e822cf5d7f0f99f3263db942135934a

Observation a74c4a37-3aca-4d5d-94d0-28c8a50ffbdb · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Measuring multimodal mathematical reasoning with math-vision dataset

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.645477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.645477Z digest=sha256:55c4ce2de629e49b9736f024f4adb62ee9589c20823e659d87e0d5d252c0967f

Observation 320d399e-1b4b-4016-ac2b-82b8f1efb5be · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.326990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.650019Z digest=sha256:f6f152299b270a9b365c308447eedfe7ff245f34b07e23718bfd693a2e466578

Observation f06d5710-f9c5-4b5f-a98d-e20707970c57 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.654244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.654244Z digest=sha256:2ecacfe571e7ce2c8f581542e8f2bdda75a9e15511f0b6f3784d1d0f2e3402b3

Observation 02f09eb4-7996-4dfb-af6e-2b535747e291 · outbound

This paper cites Visulogic: A benchmark for evalu- ating visual reasoning in multi-modal large language models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Visulogic: A benchmark for evalu- ating visual reasoning in multi-modal large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.311942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.658494Z digest=sha256:f4b41fcc870d1f3b6870e2279a73859c046d895c4011a9e65eac95a6bdefcc00

Observation 26c84fe8-d53e-442a-a4e8-85a94ad09106 · outbound

This paper cites Teaching CLIP to Count to Ten.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Teaching CLIP to Count to Ten

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.662073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.662073Z digest=sha256:5cd589c0565453f38569373f5ed949d3046922dfe7e65616221189456da5554f

Observation 1804ee86-6b67-4b94-947f-7787318f7452 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction PaliGemma: A versatile 3B VLM for transfer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.665902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.665902Z digest=sha256:6f8a793a3c8d54596e2cd0931bb60a1a27d5d0df4edd6b86d2e15331330053be

Observation def81ece-2147-47a0-bae5-2907a8f1f1a4 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.296526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.670137Z digest=sha256:09e746c4824fc177045005f70aad855c36f4174800b000b227c68b4a44d41a02

Observation 40f3269a-697b-44c9-a810-275fa14c1ae3 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.280336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.674031Z digest=sha256:bcfa54deb5d8e66ee2218fbd672ebef92abae99fd2fedaa0acf03b6f6991ee64

Observation 314f5c6e-7cb5-418f-b2df-3c51bcb1741d · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.677995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.677995Z digest=sha256:fe019dcb954b14e3b870989b444749d5222013570bf55b653879bad842085759

Observation 2cbd7d36-a227-4539-97e0-6a8033974b9a · outbound

This paper cites VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.682022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.682022Z digest=sha256:b89ec24732b8ef4d3cbdbd68d8abb71bd453687c72548ae382db01b68e1a76b0

Observation cd0019bf-cc06-4fde-a865-57ac6f7ea711 · outbound

This paper cites Thyme: Think beyond images.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Thyme: Think beyond images

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.254442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.686813Z digest=sha256:2079c137a2babe30b54684f5c32c54fd27674bd39529f012dcc158413c80c06e

Observation 523eabbf-f825-472c-b160-5d5f89a617c0 · outbound

This paper cites Improved baselines with visual instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Improved baselines with visual instruction tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.691642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.691642Z digest=sha256:799a361ba901f08105511885207b92547286740c286b2f9be17b4d90192a4605

Observation 83854f49-131d-44c5-bbf6-29f72bce05df · outbound

This paper cites Qwen2.5-VL Technical Report, February 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Qwen2.5-VL Technical Report, February 2025

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.229128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.695959Z digest=sha256:7dea8580454303bb9153b0ee52aca8de373c9f83dbbd7dd5e609dc04faaedfc3

Observation e64aa5f8-a4cc-45e2-8d69-c4de94a0fee5 · outbound

This paper cites Metamorph: Multimodal understanding and generation via instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Metamorph: Multimodal understanding and generation via instruction tuning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.212857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.700366Z digest=sha256:3ee7d1dc51811e73e17b2ae686af15406b72355c1e2eb860cfabd691d876de08

Observation 6a154796-6cdf-4c11-a87a-e2d00e90aabb · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.704691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.704691Z digest=sha256:db0ee9bb41305924e55dc344614e83e27274a4bf9ae0836c42f5eb8713fb68c1

Observation c858fa4b-9750-475f-82ff-832482d94d18 · outbound

This paper cites Show-o2: Improved native unified multimodal models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Show-o2: Improved native unified multimodal models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.196673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.709277Z digest=sha256:3dc33dcea035f94ff72317ffe54f73fb2475a4389e3cd9995232536254f63fe6

Observation afd85e1c-5dd4-4d8f-8ef3-6c2d91e10f02 · outbound

This paper cites Cheers: Decoupling patch details from semantic representations enables unified multimodal comprehension and generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Cheers: Decoupling patch details from semantic representations enables unified multimodal comprehension and generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.713410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.713410Z digest=sha256:023ddac5c18cfe823dbdb685cc61b57499138c1667401e79d2472382a27ff59a

Observation 50dfeed4-2f57-42eb-a738-4a3adbb969c6 · outbound

This paper cites Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Coyo-700m: Image-text pair dataset.https://github.com/kakaobrain/coyo-dataset, 2022

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.718062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.718062Z digest=sha256:08744d5a6f64189123ab8e56b4374a1425fabad67e9fb6505163b74eca905328

Observation 48670c1e-7796-4337-84f5-53a46388c1c1 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Imagenet: A large-scale hierarchical image database

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.722979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.722979Z digest=sha256:45d5c243e4a71ce5f1cea06de78e9c82616e6278bf351cdcd9bb25112d32114b

Observation 5ecf47df-c7d9-4053-b9f6-444cd7a69bcc · outbound

This paper cites Gen- eration and comprehension of unambiguous object descriptions.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Gen- eration and comprehension of unambiguous object descriptions

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.161417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.727985Z digest=sha256:59ed896d962bb2c826340bb72ae9a6e65876faa6d567be5b5bab4db4e7cebd82

Observation 3091c69c-ee9c-4e96-a108-d3f6d8b20ad8 · outbound

This paper cites Referitgame: Referring to objects 23 in photographs of natural scenes.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Referitgame: Referring to objects 23 in photographs of natural scenes

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.145538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.732783Z digest=sha256:87799b779434c47c8f51ab6c1db3e1c04187439b98ca23cd206cbe1906f95214

Observation 26584f7a-e5c5-430d-8e1e-9bb1d93aa0af · outbound

This paper cites ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.737344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.737344Z digest=sha256:3809022d13ba78099d6f1c9bd131e7ed09225dce52aaec7a9de49df05abed8fb

Observation c70cbff7-5db0-4c6d-b223-3480478740f4 · outbound

This paper cites an unresolved cited work.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.742241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.742241Z digest=sha256:9d93fac51a95c0c271575d309dd1fa016b3a8fcec27a36d9ac86c2fc3e370a97

Observation 905e03c1-049e-4f64-a7bc-241d4ddf8505 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Flamingo: a visual language model for few-shot learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.747386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.747386Z digest=sha256:67010385975ec855531d370ae6203050ef97cdbbac8427486fe364fc6fd9d3d2

Observation 9268cac0-a2b7-4868-b137-d12d81434bcb · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.111944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.752571Z digest=sha256:b2b1fb1f183721a598e86e2b9bdb2b88bc2e3e0d57ea36f09c0cc09e789d10b5

Observation 84d506e9-63f8-4616-8f32-61a9a384bdfe · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.758009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.758009Z digest=sha256:f73d4ef580d2064115983927581486e6aa1acb81c9d7933f5f10d705c3451e76

Observation 806d6f5a-3a59-438d-b1ec-c1e99e70adca · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.762555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.762555Z digest=sha256:ca4971e4741c857127ee536ba011ff74efb9d1bb46ef4c3fd37f9c7c68cc4dd6

Observation bd887cbb-d5da-4d59-9661-9de455213db5 · outbound

This paper cites gpt-5-system-card, 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction gpt-5-system-card, 2025

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.097932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.767368Z digest=sha256:51fa6f4e2ca8d4f6cccda9cf3737ffe494676b60ffd656a8674069f032ec81f1

Observation aacebb07-2ed8-446c-91cb-8ecf55737892 · outbound

This paper cites URL https://blog.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction URL https://blog

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.082321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.771718Z digest=sha256:6bfd585c10db787c43cb7b05f19e7a4e4d97c848fadf4fc9231dcf1e2346979a

Observation 53b7ce97-9450-476f-be32-e5a9b92933bf · outbound

This paper cites Gemini 3 flash: frontier intelligence built for speed, 2025.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Gemini 3 flash: frontier intelligence built for speed, 2025

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.067638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.776922Z digest=sha256:3590dbee8d60ff9c2251ff9750023cf979d0ab2d1fe84f3791b2dfce0a616c93

Observation fd929ecf-1804-4e14-a388-44982ba0013b · outbound

This paper cites VGR: Visual grounded reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction VGR: Visual grounded reasoning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.052990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.781312Z digest=sha256:9adb1181d195a9ab70177ff2c3872aabaff7a9152e5e049144a6e69e9e6d6595

Observation e58da99f-15e6-4f75-ae40-28dbe6609a75 · outbound

This paper cites Explain Before You Answer: A Survey on Compositional Visual Reasoning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Explain Before You Answer: A Survey on Compositional Visual Reasoning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.785540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.785540Z digest=sha256:0f77a044129f4a73960b4495fb2eb8b4da175f6f442fb04d0bd889312e7686c0

Observation 8f707088-def0-4a1a-9541-a31fde45abf9 · outbound

This paper cites Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-XL-Pro: Reconstructive Token Compression for Extremely Long Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.789852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.789852Z digest=sha256:28c763e926be74b7a052a7f7a7867b5aa78ff0692d4cc89a17fcfa28ce17f49a

Observation ac6385c7-46e4-4ff2-ae91-b502cbe09e08 · outbound

This paper cites Video-xl: Extra-long vision language model for hour-scale video understanding.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Video-xl: Extra-long vision language model for hour-scale video understanding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.036455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.892662Z digest=sha256:c2f575967ba098142c24e876aa6abfa808af491796598b3840c87f463cdaeb41

Observation a929b0c3-f6b3-4cd3-b19b-0091995455d4 · outbound

This paper cites Videochat-flash: Hierarchical compression for long-context video modeling.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Videochat-flash: Hierarchical compression for long-context video modeling

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.020711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.897426Z digest=sha256:030646f7875747921561df50f38e1c4c50eadd70560697a11456685f99137c7a

Observation a4a286f7-2de5-4dce-8344-2dd178faa383 · outbound

This paper cites Reconstructive visual instruction tuning.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Reconstructive visual instruction tuning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:35.004875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.902010Z digest=sha256:9ecf5767e7e6e3870e535e49fcd3a53db28ec9334cc71ffdea628b37031f9f0d

Observation e779dc83-cfa0-4390-b4fb-f2934a771cda · outbound

This paper cites Autoregressive semantic visual reconstruction helps vlms understand better.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Autoregressive semantic visual reconstruction helps vlms understand better

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.989707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.906676Z digest=sha256:f93838104f3a42275f51fd99c99e82f084c88597fd7061cbd3087372da43c0af

Observation 0dae9be8-18b4-462f-aedf-d247e85fa217 · outbound

This paper cites Generation enhances understanding in unified multimodal models via multi-representation generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Generation enhances understanding in unified multimodal models via multi-representation generation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.975035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.911508Z digest=sha256:604ff4bd3bbb79b79407603f5cf30c9c9d02db3036f0965daf065783178d8f59

Observation b8ab97a0-4659-420f-bbd2-40479f921270 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-16T00:17:33.916989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:17:33.916989Z digest=sha256:bbacce98aefdc07c72372640052fd5e984c8629f884ee5241913e26166aaebaf

Observation f820c2a0-3031-4f0e-87b8-b377bd0734ae · outbound

This paper cites LMFusion: Adapting pretrained language models for multimodal generation.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction LMFusion: Adapting pretrained language models for multimodal generation

Reference 72

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:17:34.087176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.921910Z digest=sha256:f3a62495a0db6c5ddb07fcf9b9038f716ad3939779e946447816af12572dcb2b

Observation 95c572ab-905d-4dd6-bc6f-629fc15d2719 · outbound

This paper cites - Row 1, Fig.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction - Row 1, Fig

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.960789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.926801Z digest=sha256:fbaadf39dd35aea46b8395d05d3f9f47888ba29ce8374c376f4bd6f729eb3467

Observation 3b16e616-1bc7-4869-aa16-07fc00ab01c8 · outbound

This paper cites All figures share the same two-shape overlapping composition with consistent contour and inner-line connectivity.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction All figures share the same two-shape overlapping composition with consistent contour and inner-line connectivity

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.947379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.931751Z digest=sha256:f550a2342bcc48c45af0d41459fd0703877cdf3f182535a6845b32345961f822

Observation 2c55ad56-d9a3-48ba-9d36-a7298d51ff94 · outbound

This paper cites Option A is a triangle pair with a mismatched shape; Option B is a triangle pair with inconsistent line count; Option C is a quadrilateral but with a different stacking style.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Option A is a triangle pair with a mismatched shape; Option B is a triangle pair with inconsistent line count; Option C is a quadrilateral but with a different stacking style

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.931691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.935730Z digest=sha256:bbcfb96fadfef25c8645ab45c06beba32beea7a5986c2858c898b34d15429dbb

Observation 64359b35-023e-4c33-8ad6-d23ce8ab9a83 · outbound

This paper cites A bipartite graph is a graph whose vertices can be divided into two disjoint sets such that every edge connects a vertex in one set to a vertex in the other set.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction A bipartite graph is a graph whose vertices can be divided into two disjoint sets such that every edge connects a vertex in one set to a vertex in the other set

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.915799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.940312Z digest=sha256:a97ac08febdedb9f104f5ea2ba51a37c04ee927d82fc78c0661d3a04241fc3ca

Observation ace91903-90da-48b4-8a18-a0dfcaaadbf3 · outbound

This paper cites If there are no odd-length cycles, the graph is bipartite.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction If there are no odd-length cycles, the graph is bipartite

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.899750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.946065Z digest=sha256:edbec487bafba99a112d63847a62fbf1890a7f3d7d897105b99a1d87f85d72ba

Observation fed92e39-9f71-4eb9-b408-a516c7389120 · outbound

This paper cites This is because we can color one set of vertices with one color and the other set with a different color, ensuring that no two adjacent vertices share the same color.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction This is because we can color one set of vertices with one color and the other set with a different color, ensuring that no two adjacent vertices share the same color

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.883082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.950502Z digest=sha256:ac5f39df689c2779b1c0b1784ee0e9223059d737d9fbe038cb459c426ba92434

Observation 7a933e72-25a3-440a-84da-7f5ef6be6931 · outbound

This paper cites </think> The final answer is2.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction </think> The final answer is2

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.866845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.955074Z digest=sha256:700d21848027ecd95c267390307c147a177ab21c616ec32f929ef0fe211c154c

Observation 1cb9d5ad-aa2a-48ce-adbc-690eaec86790 · outbound

This paper cites There are two visible wooden poles supporting the tree.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction There are two visible wooden poles supporting the tree

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.850621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.959994Z digest=sha256:f1648c886bbc8d763073f4f82eb01a38a20ef467fe255202f462d18a40aabccc

Observation 349c59f0-13ea-4c7c-ab7d-5e1d5209b7af · outbound

This paper cites There are no additional wooden poles visible in the background or elsewhere in the image.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction There are no additional wooden poles visible in the background or elsewhere in the image

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.833597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.964690Z digest=sha256:f57684377f7828495d4b05f19b463e152b91a94f77b950719f9707040fc57b53

Observation efe688c9-dee5-47cf-8bef-abf21d9a764e · outbound

This paper cites Given this analysis, the correct answer is: **C.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Given this analysis, the correct answer is: **C

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:17:34.817360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.970322Z digest=sha256:a0969e3938313d0fb4dbe2e859ad154261500a72cfc298c24168c53c8fccb10e

Observation ff1816c0-e15b-4c2c-99e6-b95510212187 · outbound

This paper cites an unresolved cited work.

Generation as Auxiliary Supervision: Enhancing Visual Understanding at Zero Inference Overhead via Decoupled Embedding Prediction Unresolved cited work

Reference 2025

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:17:35.474150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T00:17:33.548992Z digest=sha256:01a47471355aa3523cce81d12c2141436dfc03d42710835e4d1526e44b1055fe

Pith citing papers

No inbound Pith citation observations are available.