Pith. sign in

Paper Citation Record · LEDGER

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2507.21741.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21741 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:31:26.112925Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b290fad-7789-49b4-81c4-471ea9f4f717 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.450246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.013359Z digest=sha256:d9d9f197a9b13ae2366a3b8c919323eb9df8f68f843f24d028a4b04b0d165143

Observation d6158cf1-a03d-4fa7-82d0-0b2f80fb5002 · outbound

This paper cites Uniter: Universal image-text representa- tion learning.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Uniter: Universal image-text representa- tion learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.414687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.029256Z digest=sha256:f4591014a8ce22086a033c03a7ed5926cfc656b5451ab6d784b750a5b0a91713

Observation 19b30b09-1918-4e97-9c92-8b68e719fc18 · outbound

This paper cites GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.032238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.032238Z digest=sha256:c8e586cd7a7b86c8c1e442cc1a97d9c553aee4682456f517c43399cb11e3ad0a

Observation f5cbb349-b375-4e07-979b-b13e43915fe4 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Gonzalez, Ion Stoica, and Eric P

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.038126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.038126Z digest=sha256:5e1038ebef94a03ef67de312120fc88968cd4fb2395745d695f809273d2bbefb

Observation c1c8e38f-9a7a-46f5-88c9-4a4c482968e3 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with in- struction tuning,.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Instructblip: Towards general-purpose vision-language models with in- struction tuning,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.399980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.040867Z digest=sha256:da19d01e86a8c9bf6039bea804202436250b8db755d30aea4d9f4ffc79f9e18a

Observation 20f951cd-4963-4261-a512-14d0829a679b · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces NVLM: Open Frontier-Class Multimodal LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.043505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.043505Z digest=sha256:7209f7b52a01b8bda207495d1f94e4358f2532a60e410209b61cbecff6e49e6b

Observation 68acd8d9-cf1d-4eb4-b295-f68536314c2a · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multi- modal large language models,.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Mme: A comprehensive evaluation benchmark for multi- modal large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.391906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.046621Z digest=sha256:0afd98dfe6a868301a9efcb46265069985b103d600ce0a188bea3d0ad902c075

Observation a444735a-947f-4d70-b860-00624c37a323 · outbound

This paper cites MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.049030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.049030Z digest=sha256:47628e961fac2ded0bbcd120e1e656d2000329851685a92e2993ec3bea966bba

Observation 8b4015ee-c20f-46d4-ab7a-afa9f0be0a91 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces LoRA: Low-Rank Adaptation of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.051888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.051888Z digest=sha256:851718c7a46784fc1d81437d89ebff86ef0a34d1d48c07b853847f48ba5cfcca

Observation b2704b91-d00a-4a38-96a5-37ccec7f7ff1 · outbound

This paper cites Dvqa: Understanding data visu- alizations via question answering.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Dvqa: Understanding data visu- alizations via question answering

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.383131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.054721Z digest=sha256:e86fcfa4ccfc5d18d4eb598fd2f1247c8633af93b9d0876f48858b7fbe75d1c0

Observation 7ba075e9-3b99-40b1-8010-230270d6c175 · outbound

This paper cites VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces VALUE: A Multi-Task Benchmark for Video-and-Language Understanding Evaluation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.059865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.059865Z digest=sha256:d74bf1a25d0a93c381ea509fa84307a526482e2165f8ea36ee551ad96603f956

Observation 1c167ae1-5c2a-43d5-825b-1f490f60785a · outbound

This paper cites Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.366203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.062802Z digest=sha256:9e12708f2cb099f80a00526b68b21a7a590a3b432e5eddb7672593590660dea3

Observation 1ba7ca47-eae9-479e-a316-ec72aeb044c8 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.065302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.065302Z digest=sha256:482deb3561ab3a62b8d06af06587b5ba1513d24b28149efcc205e8508d78a083

Observation 76aa291a-bc34-46bf-a010-0976aa03728c · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Evaluating Object Hallucination in Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.068491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.068491Z digest=sha256:3eb9dfbb84245020321301707d5c5d7c55b1da5416a355b4d1ccac53e83b76da

Observation ede1ea7f-a126-498b-8996-33326c899b63 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.071331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.071331Z digest=sha256:9bc9f005845a65438ce38011e07724e32709f367c3d1e170396d18f7ff75b64c

Observation b04a3492-ffc0-4989-92c4-271b99e7bb8a · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.074772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.074772Z digest=sha256:5af9dfd7d701d745f8853cdf9ba6aeae9dcb544a4d0d7575e22bc2b8b76fecad

Observation 1dd6caea-b43d-4aef-9ce0-1164a14fa382 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Docvqa: A dataset for vqa on document images

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.356567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.080676Z digest=sha256:81f268513f1c9f4cffa6a963f5ab144fc5c769c836e75d8487cd721e562ea557

Observation 0249a39e-5b92-43f2-8dc0-9867770aaaa5 · outbound

This paper cites Hello gpt-4o.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Hello gpt-4o

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.347968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.083312Z digest=sha256:009294501ff14dc8d7c1d574c5dfa454ccb3e1b92da06168d6544fab9033125e

Observation d1136b45-0eb7-4423-b9cf-45ab13f36844 · outbound

This paper cites [Radford et al., 2021a] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces [Radford et al., 2021a] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.338397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.085780Z digest=sha256:5a79f1af45e88b05c40fe68ed5ddc3bcd9f6fb21e10e47f35255c47ee8af5d3d

Observation 36f9e785-7302-4ec7-a97b-9e954a01647e · outbound

This paper cites Hug- ginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Hug- ginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.329223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.088243Z digest=sha256:66b0e9538f8e34df3ee7c470def4253b0f639e0915ac301a388759b95a3faefe

Observation 3fab29e7-4af2-4062-99ec-f3d863f35600 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Gemini: A Family of Highly Capable Multimodal Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.091282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.091282Z digest=sha256:3404288aa9953b4511324529090c4e19a9707fe63a86579f03b4ee4810a9e664

Observation 3a37acac-9ece-45ac-b4fd-db999bb3fd8b · outbound

This paper cites Attention is all you need.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Attention is all you need

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.320574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.094018Z digest=sha256:81b630ef9381f47bd50490e0d7b9df1c38f22b0ea68dc994700f997bbf546a16

Observation af5e91bc-0260-4800-a078-d8b76e320927 · outbound

This paper cites Introduction to convolutional neural networks.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Introduction to convolutional neural networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.311833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.099172Z digest=sha256:13ff3d3dfe5aa97a2e2bef523af24f508289f1866d7e298f32069fca07a9826b

Observation 755eb3ef-c83d-46cc-b854-83e06da943dc · outbound

This paper cites DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.101813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.101813Z digest=sha256:de88c439954911a12d79e1d46bbf9a4a862df1a3d57dbfc244d9e43b353c1455

Observation 601549d3-cebe-4ac9-940d-289cc83354ca · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.104757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.104757Z digest=sha256:36660e3a45187227d6887b5f128c7c8200e04b9ad3219f4a11ffcdd4cddcf7a6

Observation 45b1656b-993a-4acb-a062-d5813e053d15 · outbound

This paper cites Florence: A New Foundation Model for Computer Vision.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Florence: A New Foundation Model for Computer Vision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.107503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.107503Z digest=sha256:f3c6217f3ecce87a689fd7c31360b0b4470a1badb14ac5a5fc874058d793382f

Observation 41f0bfe4-120a-48ef-aa2e-b2bde586493c · outbound

This paper cites Easygen: Easing multimodal generation with bidiffuser and llms.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Easygen: Easing multimodal generation with bidiffuser and llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.303148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.110181Z digest=sha256:8da7423fecffd012da544818d27437d3118e5018b36d2d7a18840a5b8e04a831

Observation 832fbd14-48b9-4ce9-a888-c694d51f440d · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems , 36:46595– 46623, 2023.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in Neural Information Processing Systems , 36:46595– 46623, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.293335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.112925Z digest=sha256:edf9ae256610edfc3e397905051c00e95c687318d977f8c08dca4ffad7f477c9

Observation 67244016-c8d3-4d9f-9475-48c2196fe6d2 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.096481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.096481Z digest=sha256:fd1f165136685d872d12918b17d0acb3f75f843f3d0b6ccd46d8cd5a8656cc27

Observation f39d8168-5047-4033-9f8d-a1275fdd40f3 · outbound

This paper cites Ocr-free document understanding trans- former.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Ocr-free document understanding trans- former

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.374882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.057257Z digest=sha256:47d185f4beac3cb9d0f7b2cf0fd164118befe92797a57621f4b012c8f6ebb1b0

Observation e640e00d-ee1a-49ad-9e67-2782314a0c7c · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Honeybee: Locality-enhanced projector for multimodal llm

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.423521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.026185Z digest=sha256:0ca4b16356bf63254354d69c225d9e1005625831c9fa1101f2589840e00eac78

Observation bc5c1853-1c5a-4793-94e1-21bd4be5d566 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.035221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.035221Z digest=sha256:da7cce0f4698cb130ca98187ff3fcede9373b18bf8d8d94e1474e17398f10b23

Observation 2a606dd0-01a3-457e-987e-de605d458991 · outbound

This paper cites Claude 3.5 sonnet.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Claude 3.5 sonnet

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.441144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.016801Z digest=sha256:300e59a8db988b7fbb3bbe89dd6f241413de49769cfff630983f795abdd99b85

Observation 9781c94c-d3a5-4783-b2cd-f10bc6e367d7 · outbound

This paper cites Language models are few-shot learners.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Language models are few-shot learners

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:31:26.432442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T12:31:26.023150Z digest=sha256:5db240c5a51ffd878bad93b37349dcff817964407840d67d796905989518365c

Observation fabbac7f-b0b7-49b4-a84c-845b2a2754cc · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.019701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.019701Z digest=sha256:79cb407973a2ed7e0a2d590c2ce28c74c89af6461836bd2fe4c1cbd1dc513cdb

Observation a858a163-883c-4807-a054-80b7481666a0 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T12:31:26.077842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:31:26.077842Z digest=sha256:8ea2aff5132387e35b448666bab0f1c3cf291b6b83b557c6c464d3f2cccb6e24

Pith citing papers

No inbound Pith citation observations are available.