Pith. sign in

Paper Citation Record · LEDGER

Olympus: A Universal Task Router for Computer Vision Tasks

As of 15 August 2026, this Paper Citation Record lists 100 of 102 outbound references and 1 inbound Pith citation observation for arXiv:2412.09612.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09612 v3

Coverage vector

measured 100 of 102 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:57:06.818547Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:14:42.975508Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:14:45.125038Z

Reference resolution

100 of 102 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved75
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f5f8ebd4-2298-4343-9a45-0bc80dba677f · outbound

This paper cites Jointly training large autoregressive multi- modal models.

Olympus: A Universal Task Router for Computer Vision Tasks Jointly training large autoregressive multi- modal models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.739052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.739052Z digest=sha256:52ed75142118fab68971fd82009b1f79d445187029964f507c732dda6a3f4cd7

Observation f8074de5-f924-495e-b8be-8ed07ecbfc3a · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Olympus: A Universal Task Router for Computer Vision Tasks Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.750471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.750471Z digest=sha256:22a1ad563d6987ae3e6301f18a0dd25664d83282c3d699eb03ece3f2de02c222

Observation 5e0662cb-420b-4460-b502-37702a660500 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Olympus: A Universal Task Router for Computer Vision Tasks Gemini: A Family of Highly Capable Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.763199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.763199Z digest=sha256:8703a110ca886a7015e00c07daec1361742b5701197d580602393ea79a7ef9da

Observation cf49f528-b834-4f9d-bae9-a9f96b913ef6 · outbound

This paper cites Vqa: Visual question answering.

Olympus: A Universal Task Router for Computer Vision Tasks Vqa: Visual question answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.779086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.779086Z digest=sha256:bc3f8e66f73165730e922640c8354e575dd6b0737f0ea1797a800d2fe39fab8a

Observation 24f30fd7-eba8-41e7-a57f-d4f27ac30262 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Olympus: A Universal Task Router for Computer Vision Tasks Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.788903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.788903Z digest=sha256:5ba148791121c6e342a28266c0cab95cdc989f722109932ec9974478f2a18e61

Observation f6827a3a-35a7-439c-8758-eab2f98db29e · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Olympus: A Universal Task Router for Computer Vision Tasks In- structpix2pix: Learning to follow image editing instructions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.801317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.801317Z digest=sha256:a96f57b7b48fc4c90a983254a0cb0a291c7257621cb278c9c6a496cda2c1da4e

Observation 5ff84088-9dc4-45ce-8106-a5d639bcf826 · outbound

This paper cites Language models are few-shot learners.

Olympus: A Universal Task Router for Computer Vision Tasks Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.818171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.818171Z digest=sha256:7949e6ee6558266ae62bec0cdec93c7f937f20b1dd4bba241b711b622c6f4bd8

Observation 63253333-4349-4526-831f-2ef8f38e3e98 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Olympus: A Universal Task Router for Computer Vision Tasks Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.827444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.827444Z digest=sha256:d33db4125dc50910d5d25e06deccd6203ff9d32a1937ec240cb8f6336cda78da

Observation 4c43580b-f786-4bee-b863-cd8448fc4d7c · outbound

This paper cites Text-to-3d using gaussian splatting.

Olympus: A Universal Task Router for Computer Vision Tasks Text-to-3d using gaussian splatting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.842477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.842477Z digest=sha256:3471484100afd72eaed543fed81387f4b9f51aec618b4d62be5a8b336085c0c8

Observation 6180be68-b226-4889-9aae-24618a048fc1 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Olympus: A Universal Task Router for Computer Vision Tasks Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.848790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.848790Z digest=sha256:2327dddd5a1c8bf47c38b196dbe73575b79e711331f9e49889dfe0eb85642740

Observation d39d4adc-2360-4716-9ed6-535f32f1af5b · outbound

This paper cites Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C.

Olympus: A Universal Task Router for Computer Vision Tasks Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Benton C

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.858906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.858906Z digest=sha256:a5c2af412e874b493aa3ca8cc091d312b0a7537b711b84508bcdc084d58220cc

Observation 413c2f35-be99-454d-a4c1-e40a2e70ac73 · outbound

This paper cites MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices.

Olympus: A Universal Task Router for Computer Vision Tasks MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.869030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.869030Z digest=sha256:b9544af9601a87ac9ad410b2fd4761a7fb517dee4f5490475d77892998aff3fd

Observation b5953200-717a-4380-abbc-8d18c8a9ba63 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

Olympus: A Universal Task Router for Computer Vision Tasks MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.877440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.877440Z digest=sha256:cada524ae85e5771970f237eb8fac92a0f414d2087a9b2c6748f5ea1e4dccd0d

Observation 28498705-c197-4dd1-b2f2-438994bccc90 · outbound

This paper cites Swin2sr: Swinv2 transformer for compressed im- age super-resolution and restoration.

Olympus: A Universal Task Router for Computer Vision Tasks Swin2sr: Swinv2 transformer for compressed im- age super-resolution and restoration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.891307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.891307Z digest=sha256:443253901fa9ac8a3d23ac4dd7286c48368a2db0925c4c6037b1ba110458598e

Observation a0dcd795-c6c6-4542-b704-bb72b06f01ef · outbound

This paper cites InstructIR: High-Quality Image Restoration Following Human Instructions.

Olympus: A Universal Task Router for Computer Vision Tasks InstructIR: High-Quality Image Restoration Following Human Instructions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.904435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.904435Z digest=sha256:ba6da19f6654dcd75885a647947cae63d7de15c406e11d49772b16d3a44097e2

Observation cfb28bee-8a3c-4b6e-858f-c404232f58ec · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

Olympus: A Universal Task Router for Computer Vision Tasks Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.914435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.914435Z digest=sha256:3c948b829506966cedb9d695b76de4e01c651c9c53dea56150a8dd31026b962e

Observation 09647163-bcc2-475c-8f2d-4282bde20cb3 · outbound

This paper cites DreamLLM: Synergistic multimodal com- prehension and creation.

Olympus: A Universal Task Router for Computer Vision Tasks DreamLLM: Synergistic multimodal com- prehension and creation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.938658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.938658Z digest=sha256:52462af1c9c2057c90df22d36e6210a3f192c34a2f6806bd861fcc876716ee3e

Observation 7e92694b-5216-49a0-a326-13524f07d6e9 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Olympus: A Universal Task Router for Computer Vision Tasks An image is worth 16x16 words: Transformers for image recognition at scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.945583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.945583Z digest=sha256:80a152f810935a95904d5a28781438144967e6ee173d13d0419d30ec3c546f97

Observation 4f1cb4f0-01c1-4e75-baa1-9dd39be6f903 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Olympus: A Universal Task Router for Computer Vision Tasks Scaling rectified flow transformers for high-resolution image synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.961591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.961591Z digest=sha256:f767f1d68acc4edbf946d30802b3b8494f76edace6996f63f7673480a72522b7

Observation 52bd0f10-12b4-467c-82ef-eee4156a0c75 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Olympus: A Universal Task Router for Computer Vision Tasks MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.977798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.977798Z digest=sha256:208a47aa915744f21095aacd3842eca80dc7307d17b2b51e260a4a1f33921ca5

Observation ed89863c-9687-4265-b18f-38da981a60d2 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Olympus: A Universal Task Router for Computer Vision Tasks SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.985650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.985650Z digest=sha256:d14527b5535d25f0f3be28de09451245b84044068aea69b92c2f0f92102c04b5

Observation 879700aa-4510-4ca0-92c6-83b078cfd6c3 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Olympus: A Universal Task Router for Computer Vision Tasks Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:05.994989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:05.994989Z digest=sha256:0cdf1705154da4adc9b280f96a2c85110c3a46144e0887bb633810531bfeb398

Observation 63ce18e2-ca37-4576-af68-722e1a4b4aed · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

Olympus: A Universal Task Router for Computer Vision Tasks Visual program- ming: Compositional visual reasoning without training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.000710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.000710Z digest=sha256:bf0565b7c6929ba267ee8958a6bfbff3dd6f5fd8b04be98f0e7aba21155d0da9

Observation be1e08b0-c69c-4162-9aad-b68a774727c9 · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Olympus: A Universal Task Router for Computer Vision Tasks Vizwiz grand challenge: Answering visual questions from blind people

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.014009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.014009Z digest=sha256:6b88090432765b7e5f0b76a1986175fce7414bf500c50c75a46e1e54d6a14af2

Observation f57083ff-43df-484f-8131-e01d1f431a9f · outbound

This paper cites Efficient Multimodal Learning from Data-centric Perspective.

Olympus: A Universal Task Router for Computer Vision Tasks Efficient Multimodal Learning from Data-centric Perspective

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.024131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.024131Z digest=sha256:1ad590a1ac2bd249d0494fbc763de77b85aead0e7e22b8ace6bf7c963d20ee89

Observation 4d547d06-d94f-4d01-90f1-f41824c16405 · outbound

This paper cites Denoising diffu- sion probabilistic models.

Olympus: A Universal Task Router for Computer Vision Tasks Denoising diffu- sion probabilistic models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.030297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.030297Z digest=sha256:f3ba3a332b78453a1b2c41e51caa9e0b743e16155325666d380271a764345223

Observation 8fa00956-1704-4db1-aac4-a396a27972a1 · outbound

This paper cites Video diffu- sion models.

Olympus: A Universal Task Router for Computer Vision Tasks Video diffu- sion models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.061769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.061769Z digest=sha256:23cde909388b44a95a496345ce7ac97373fa761feefb52e301d59aba864656ee

Observation 79d767ff-aed9-4263-9f97-9e0324c62875 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Olympus: A Universal Task Router for Computer Vision Tasks Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.071077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.071077Z digest=sha256:527ac5a5187f3d37ef7fff0d020ee3462774ab53f8fc20bca2f57f7cf4f4f084

Observation 2f58b597-9f1e-490e-924b-a61d7e4b5213 · outbound

This paper cites GPT-4o System Card.

Olympus: A Universal Task Router for Computer Vision Tasks GPT-4o System Card

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.081115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.081115Z digest=sha256:0b92e9715de26ffbebaf685fb41317c9ca56d76266c6e52cc016a70aff653d09

Observation 153f4d05-3b25-4b49-9c1a-c6ae14fa6f57 · outbound

This paper cites Text2video-zero: Text-to- image diffusion models are zero-shot video generators.

Olympus: A Universal Task Router for Computer Vision Tasks Text2video-zero: Text-to- image diffusion models are zero-shot video generators

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.087395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.087395Z digest=sha256:a3520d3712bd2bda6393ad9820d13494ba7ebf7181dd3778b1544ebf0c11996d

Observation 074f8014-4c4d-444c-83e6-11db2b7a5f67 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Olympus: A Universal Task Router for Computer Vision Tasks Adam: A Method for Stochastic Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.095406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.095406Z digest=sha256:9a773b27edaa1d705672eaed9aebee6ec6ec5fa2f48121eb04c0acdc888c62c6

Observation b815a305-1770-44aa-8764-c50cabcfb95e · outbound

This paper cites Auto-Encoding Variational Bayes.

Olympus: A Universal Task Router for Computer Vision Tasks Auto-Encoding Variational Bayes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.112192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.112192Z digest=sha256:02cf9275a8b1924d26cf1325c390b0de5e608df795262154a8be173a9cbe7964

Observation af63466e-3bd7-43eb-9d8b-98eef79a04bc · outbound

This paper cites Obelics: An open web-scale filtered dataset of interleaved image-text documents.

Olympus: A Universal Task Router for Computer Vision Tasks Obelics: An open web-scale filtered dataset of interleaved image-text documents

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.123830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.123830Z digest=sha256:2ab81a68e01e892b08429f8e3912aec6de37c0a18e413cc41a2488113cb4a9dc

Observation 596f06dc-b71e-4528-bc44-eec994744491 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Olympus: A Universal Task Router for Computer Vision Tasks LLaVA-OneVision: Easy Visual Task Transfer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.130954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.130954Z digest=sha256:da86aca10c7d642418a93bd145d513e998e1de1d57412888eba6f908476944a5

Observation 68b6491e-5845-4741-9ddc-b9f54858f562 · outbound

This paper cites an unresolved cited work.

Olympus: A Universal Task Router for Computer Vision Tasks Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.140440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.140440Z digest=sha256:3a65168d8c23a6f6075655dce867ebc97e130338c763572f76bf6e16277914d0

Observation 2aa77f1f-837c-4504-88fc-a72a8e0eac97 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Olympus: A Universal Task Router for Computer Vision Tasks Evaluating Object Hallucination in Large Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.150477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.150477Z digest=sha256:add67bdbff0444ded5f2ab529a25d77256a9d34598b324717e2f28891428948b

Observation e388f73d-46fd-4485-affd-7bf6cf84da69 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

Olympus: A Universal Task Router for Computer Vision Tasks Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.161470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.161470Z digest=sha256:56b9930cdd400be8639d89775dae30a642aa4d464670f7b393f5d5f8680c0003

Observation 9b99bae4-854b-4e3c-aabd-43cedaedfd3c · outbound

This paper cites Taskmatrix.

Olympus: A Universal Task Router for Computer Vision Tasks Taskmatrix

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:10.101330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.173060Z digest=sha256:3a9508175e682c16e03613cf0095194516e3e29e2180ac769ac21311623f250d

Observation 2dbd13d8-b596-4c2c-abbe-156399e844fe · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

Olympus: A Universal Task Router for Computer Vision Tasks MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.181897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.181897Z digest=sha256:7f7a5a08523433731d9db45620e8cdb1a41260f84e5b18e027988ef234d86b31

Observation c8b5ac1b-688c-4cae-82b0-20efe3db61c4 · outbound

This paper cites Revive: Regional visual represen- tation matters in knowledge-based visual question answer- ing.

Olympus: A Universal Task Router for Computer Vision Tasks Revive: Regional visual represen- tation matters in knowledge-based visual question answer- ing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:10.044415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.198388Z digest=sha256:35819b73d97b2a4bc2341ac674e3122f97e17db223c2bbbf652b5f78d5a2e68c

Observation 8b114b7f-a2d6-41a5-aea4-6e8b00ce3bed · outbound

This paper cites Smaug: Sparse masked autoencoder for effi- cient video-language pre-training.

Olympus: A Universal Task Router for Computer Vision Tasks Smaug: Sparse masked autoencoder for effi- cient video-language pre-training

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:10.012964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.212095Z digest=sha256:46610d100a812a30211b6fc4e4fb335d6a92ce7daa3bdc147453b63b7b4e9992

Observation b57ec712-b42d-485c-bdce-99aa453cdc52 · outbound

This paper cites Text-driven image editing via learnable regions.

Olympus: A Universal Task Router for Computer Vision Tasks Text-driven image editing via learnable regions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.989331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.221520Z digest=sha256:bc7f2371ce24ebd3753a0c0a7c216fdfed0eb470a264a8774f18bf9e35f1bde9

Observation 3dc1c883-4108-4578-90c9-ceb11602342e · outbound

This paper cites DreamPolisher: Towards High-Quality Text-to-3D Generation via Geometric Diffusion.

Olympus: A Universal Task Router for Computer Vision Tasks DreamPolisher: Towards High-Quality Text-to-3D Generation via Geometric Diffusion

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.235045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.235045Z digest=sha256:17ff697e982223a21dc934deb1371efbcf59a0fd429c57afb93b6ba1ee1070c6

Observation 39cf98f6-871e-4caf-959c-2d746af44546 · outbound

This paper cites Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge.

Olympus: A Universal Task Router for Computer Vision Tasks Rethinking Visual Prompting for Multimodal Large Language Models with External Knowledge

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-11T16:57:08.262214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.241776Z digest=sha256:4fcd1d2e834793ee772623f10cacd16cf332b2aad2edc695dcb55158d67cc883

Observation f2f0e5d6-a6aa-4faf-a735-c88385a36711 · outbound

This paper cites Improved baselines with visual instruction tuning.

Olympus: A Universal Task Router for Computer Vision Tasks Improved baselines with visual instruction tuning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.950437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.253395Z digest=sha256:b0e7c2939af08bf12fd857d70c0b18062e6fc8bb1e1af30cc824e502fa7e1831

Observation 8ad5e083-1d51-457c-a4dd-18e90383590f · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Olympus: A Universal Task Router for Computer Vision Tasks Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.267493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.267493Z digest=sha256:0abbbd1365df6a04b5ccc70357dd398d55560c30c60ca9b65067b816062eb4cb

Observation eccf3dc0-527b-4d4d-a5fa-1f14ea540741 · outbound

This paper cites Visual instruction tuning.

Olympus: A Universal Task Router for Computer Vision Tasks Visual instruction tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.277100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.277100Z digest=sha256:16a2c3e2c4824f5b0cd1fa90d29d021c1d82991c8622d40c1fd88de904467ef6

Observation 2d632f72-8bea-4ab6-a33b-29a9a7a58912 · outbound

This paper cites Visual instruction tuning.

Olympus: A Universal Task Router for Computer Vision Tasks Visual instruction tuning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.874943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.284472Z digest=sha256:9293a9153ec54f742208b22fbd932ae2b73e3c42b7d50f4b3d3c05a6bf5c2ba8

Observation 173af06e-3ddb-4933-b580-fd213a232042 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Olympus: A Universal Task Router for Computer Vision Tasks Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.291226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.291226Z digest=sha256:0c54691804fb9f09cae2a44f9f6865492dbf98925cd1bc082848ca6fe9502db1

Observation 1f65522e-41b5-4bae-ba4f-26e62e743e0b · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Olympus: A Universal Task Router for Computer Vision Tasks MMBench: Is Your Multi-modal Model an All-around Player?

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.299895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.299895Z digest=sha256:3f3471c264eb3dc137502406b4f03f3e2bfa44b170c05bf5e8456637846cff2c

Observation 0bbbcf05-3ef4-4f87-b9ed-1f73b3e7eeb9 · outbound

This paper cites Wonder3d: Single image to 3d using cross-domain diffusion.

Olympus: A Universal Task Router for Computer Vision Tasks Wonder3d: Single image to 3d using cross-domain diffusion

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.845097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.313191Z digest=sha256:d52b3f0527d20b5542826e2b965784c8023299b89ec350e719983183a4f97882

Observation eed922a2-6872-4a12-83b7-623b300def17 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Olympus: A Universal Task Router for Computer Vision Tasks Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.322005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.322005Z digest=sha256:05eff35df57632b9d07c7902155493c877194f955dd9de65ea17fd9a0aeb04c1

Observation 3103a04d-74eb-4f3a-950a-5e9604070d88 · outbound

This paper cites Computation of normal- ized edit distance and applications.

Olympus: A Universal Task Router for Computer Vision Tasks Computation of normal- ized edit distance and applications

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.789078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.327652Z digest=sha256:fe7d1eb8be8173748829f2c2c3d1b523408f93d53ac362a490b3dd6848238024

Observation be23d132-43b9-4498-bd91-628c2264475f · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Olympus: A Universal Task Router for Computer Vision Tasks MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.339114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.339114Z digest=sha256:3e1a78a8826bbe47f8a670fc4d2a41fba03bce2c346cbb822937f81daecc0f3f

Observation 1ff440b1-5ec6-4a83-92ea-619618a455c3 · outbound

This paper cites Phi-2: The surprising power of small language models, 2023.

Olympus: A Universal Task Router for Computer Vision Tasks Phi-2: The surprising power of small language models, 2023

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.754906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.347560Z digest=sha256:2fe35fd77e88d5839cf2ca559657f2ea500d002699655be3a5c91a3a47fa42d4

Observation 1a14f775-3761-4d4d-9a73-bf91356bcd9a · outbound

This paper cites Improved denoising diffusion probabilistic models.

Olympus: A Universal Task Router for Computer Vision Tasks Improved denoising diffusion probabilistic models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.358713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.358713Z digest=sha256:3cba39b550a177c6023f0537a7fb9de746095be2c21752a90c285c53f0ed2933

Observation 332e90d5-d41f-41ea-b1c6-ddd79b0bd221 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Olympus: A Universal Task Router for Computer Vision Tasks Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.374889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.374889Z digest=sha256:bd6958d1f3e85c0ee36917dc387e607f629534a3ed1b3d4b8074ab466dee0461

Observation 22fc9ff3-1f3d-412b-bc46-db09e861eacd · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Olympus: A Universal Task Router for Computer Vision Tasks SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.388795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.388795Z digest=sha256:af8fe03ace4f63343ee4732126686e5a72f7791476a4c4365f965d1c7662ef96

Observation 64e75bec-70cf-49d2-a2ff-840ef6d6489b · outbound

This paper cites Tool learning with foundation models, 2023.

Olympus: A Universal Task Router for Computer Vision Tasks Tool learning with foundation models, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.702326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.396571Z digest=sha256:31407d51eebd99a8b09a72e3cf645f05be7b9516fcf79a46f1ee2017ec68f359

Observation 8881dc30-2730-419d-a7bd-6c8dc9845864 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Olympus: A Universal Task Router for Computer Vision Tasks Photorealistic text-to-image diffusion models with deep language understanding

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.676668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.403942Z digest=sha256:16a4f5a5f829018ee530843708ae8ab45e4cc8d4fa34165bd2f014bf0f7e7c81

Observation 6b3b3e7c-87b0-49df-889b-5a39d2f1593c · outbound

This paper cites Toolformer: Lan- guage models can teach themselves to use tools.

Olympus: A Universal Task Router for Computer Vision Tasks Toolformer: Lan- guage models can teach themselves to use tools

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.635889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.414930Z digest=sha256:fcad5991f2a83f7d0da514ecbaa3dcb69d269e5a6cb3a97ed6ff9efb6b6aebfb

Observation 98a23545-0ad7-4122-a6e7-a66dfdcc0628 · outbound

This paper cites Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face.

Olympus: A Universal Task Router for Computer Vision Tasks Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.596101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.425612Z digest=sha256:dbbee3e835bb12d9431cf9789c67fc25f200e98d57db13201d98902a510a6c13

Observation 702ae479-9a6d-4899-9919-2aa79d520fd7 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Olympus: A Universal Task Router for Computer Vision Tasks Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.434821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.434821Z digest=sha256:9851e4ad425b8f20a9bf168ecd792b8f20df4e10f2fe528e12f20f8052d7d6fa

Observation f615300b-e68c-45cb-995f-055759bdab95 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Olympus: A Universal Task Router for Computer Vision Tasks Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.445814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.445814Z digest=sha256:3a0d89fde06b38842eaafd1beca7169f3cf68b7835eb3aeafdc45a258e6a5df1

Observation 75ce509b-ebb7-45a5-bf2c-39432f6673c6 · outbound

This paper cites Towards vqa models that can read.

Olympus: A Universal Task Router for Computer Vision Tasks Towards vqa models that can read

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.569932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.453326Z digest=sha256:a33a191a7dffe1e176cec58ab82d746fe2d4996d7bc08b828f3b87ea4279f07a

Observation 3f77945a-dd15-4250-ac02-b931a53b9615 · outbound

This paper cites Emu: Generative pretraining in multimodality.

Olympus: A Universal Task Router for Computer Vision Tasks Emu: Generative pretraining in multimodality

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.461329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.461329Z digest=sha256:1c4c3f5bcffdb88ad86c3f12b000f0ebacdf7d12625e0d680635c6073e73ddd2

Observation 41f6e9fa-8a80-42ab-a5f7-bfababe1f242 · outbound

This paper cites An Empirical Study of Multimodal Model Merging.

Olympus: A Universal Task Router for Computer Vision Tasks An Empirical Study of Multimodal Model Merging

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.469711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.469711Z digest=sha256:5fe0ac3b155c361a9c49c998eb40b53a9220f401b796bc69da83ef8e0ed24da1

Observation 110fa0c3-e474-48ec-9aeb-4eb775b2cdd3 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Olympus: A Universal Task Router for Computer Vision Tasks Vipergpt: Visual inference via python execution for reasoning

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.510466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.479853Z digest=sha256:3ddff4149ff253ed68950b5febed54b225029cfa4acc63e37fe4352043f9fe2e

Observation 7c368a9a-2525-41ef-aaa4-f75eac08d3f1 · outbound

This paper cites Lgm: Large multi-view gaussian model for high-resolution 3d content creation.

Olympus: A Universal Task Router for Computer Vision Tasks Lgm: Large multi-view gaussian model for high-resolution 3d content creation

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.467959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.488168Z digest=sha256:88258721237cf2af251d1266803dd42a68da960671d2acee006f056fdc94a199

Observation b2fa4058-4ff7-44b7-a808-4996554c662b · outbound

This paper cites Any-to-any generation via composable diffusion.

Olympus: A Universal Task Router for Computer Vision Tasks Any-to-any generation via composable diffusion

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.428213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.496849Z digest=sha256:1f05adf5033b922f81580f15ebcbe28820ce23d5a2048da429c9deaca40720f8

Observation a52d5fc2-798f-4c27-ae20-2e37aebb3129 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Olympus: A Universal Task Router for Computer Vision Tasks Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.508803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.508803Z digest=sha256:f735f61c12954528fe8d8b9993e45757893c4c5bc84fd7d09de4f096b38095b3

Observation 009075ca-5b90-47f6-a29a-0565ad916360 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Olympus: A Universal Task Router for Computer Vision Tasks Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.525636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.525636Z digest=sha256:94dc8a165102d99fcef2ee8e635e62aeb6e6ecb8be0e0d264b8560aab17aea73

Observation ecbce76f-6c7a-4a3a-a62a-ffc2e2c7e6d0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Olympus: A Universal Task Router for Computer Vision Tasks LLaMA: Open and Efficient Foundation Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.534914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.534914Z digest=sha256:163c5b5ed346d0e0d8c1dc1c13dd07a5aa7dda7acee0ebaff3dedbc490ffc44f

Observation 9b6f357e-f1e0-45aa-95da-a20d48baec2f · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Olympus: A Universal Task Router for Computer Vision Tasks Emu3: Next-Token Prediction is All You Need

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.545433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.545433Z digest=sha256:cc8e87efdcc2af282e684efb1884625ffd7199c20ea39bfd54c93d228525c2c5

Observation 9414bdaf-b781-47d4-b560-a7194aa76a8d · outbound

This paper cites Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models.

Olympus: A Universal Task Router for Computer Vision Tasks Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.555393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.555393Z digest=sha256:053c63d358459e7be6c4a997c7b1f059408bfec7db1e5e114f803c97e3f817b6

Observation 92615464-8fb5-479b-8100-b317c8dbd6f4 · outbound

This paper cites General object foundation model for images and videos at scale.

Olympus: A Universal Task Router for Computer Vision Tasks General object foundation model for images and videos at scale

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.392821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.567868Z digest=sha256:c6c62e82f30f5a6ce40b1980835fbca93a2c3a6f4b155631f97d1ef88afc9231

Observation 474b456c-df7a-4e04-9822-dcc99e0cd337 · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

Olympus: A Universal Task Router for Computer Vision Tasks Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.577043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.577043Z digest=sha256:26f60107297331bb5fc05e9033b745f6e34d74bc310d2597a00570a84cb99cbf

Observation 4573a69c-cef9-4e6e-9e57-06640bc20b12 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

Olympus: A Universal Task Router for Computer Vision Tasks NExT-GPT: Any-to-Any Multimodal LLM

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.586644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.586644Z digest=sha256:98b5f7a7d7364ea257f60a56c6d19b53f1b072fbfc3b514c7983c6914dff4e3d

Observation 56f113ee-d064-4a3e-a554-bc032c86245c · outbound

This paper cites OmniGen: Unified Image Generation.

Olympus: A Universal Task Router for Computer Vision Tasks OmniGen: Unified Image Generation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.594700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.594700Z digest=sha256:8d8e9e39ab8f9b43382483f2d43ba13570185f0874516665726b410895fcf0b8

Observation e7480453-182b-4561-b16e-e5e76cd5032a · outbound

This paper cites Segformer: Simple and efficient design for semantic segmentation with transform- ers.

Olympus: A Universal Task Router for Computer Vision Tasks Segformer: Simple and efficient design for semantic segmentation with transform- ers

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.333404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.604714Z digest=sha256:15f529d1b04b54837eacf2a0ca0e7436c3b70b1e1c6f7ed0c14ffceed1f29674

Observation 2cee957c-1db8-4a2d-94ca-6f3448285f19 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Olympus: A Universal Task Router for Computer Vision Tasks Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.613698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.613698Z digest=sha256:155de7af968820d48d66a1c67a19abc63305225f3eca55afaaa1a05527f333b1

Observation 591bf398-99fc-43d8-824d-8d4b9385402b · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

Olympus: A Universal Task Router for Computer Vision Tasks xgen-mm (blip-3): A family of open large multimodal models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.620645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.620645Z digest=sha256:2acbfe08d8625a465d680777caf11ddcf1f3d0ced6189c23081941265b61391a

Observation 00c3e428-45e3-4abf-8006-aa6ae9a18da3 · outbound

This paper cites Depth Anything V2.

Olympus: A Universal Task Router for Computer Vision Tasks Depth Anything V2

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.627632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.627632Z digest=sha256:ae2565eed0d8bcf5ab7d81fa69d239454ae54ba9c0882b6e61e11a96578f0f30

Observation 1c9b2f8d-ec45-4230-8951-73f294e3a390 · outbound

This paper cites SEED-Story: Multimodal Long Story Generation with Large Language Model.

Olympus: A Universal Task Router for Computer Vision Tasks SEED-Story: Multimodal Long Story Generation with Large Language Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.647447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.647447Z digest=sha256:310da024aacc63dbdcaf744bf4b35667c53e3e5830d0a36bda939c07aef11118

Observation ba5958f9-32c9-448c-8339-8248cb331ac5 · outbound

This paper cites Effec- tive whole-body pose estimation with two-stages distillation.

Olympus: A Universal Task Router for Computer Vision Tasks Effec- tive whole-body pose estimation with two-stages distillation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.288670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.656672Z digest=sha256:74be6c670134af5123e9e885522daf7b0045e0eb158d100047404795154573ea

Observation 96e2ee73-1641-4f64-a244-6b4bb204ba61 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Olympus: A Universal Task Router for Computer Vision Tasks CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.670118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.670118Z digest=sha256:1b41ed8ab337eb66ab73d965ebf4198f0578568c0334a2830e5f96904bbce6e7

Observation 4bc95560-121a-4c73-8631-2b8e258f4c55 · outbound

This paper cites X-VILA: Cross-Modality Alignment for Large Language Model.

Olympus: A Universal Task Router for Computer Vision Tasks X-VILA: Cross-Modality Alignment for Large Language Model

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.678270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.678270Z digest=sha256:36e1a66f9c390d3f3d3275cf23d18e1b3063c960b55cb1b2b93447efe1d34e3f

Observation c727b0cb-7c99-47b2-b93c-2987120b4c66 · outbound

This paper cites mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration.

Olympus: A Universal Task Router for Computer Vision Tasks mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.685194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.685194Z digest=sha256:3cd98f65ab66358dfa9f61bdd6d3b70f4e1b84b9ad4265eb4eab60f9976cd972

Observation 5b4ec436-78ae-4f69-9a6f-668f022d1bff · outbound

This paper cites Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning.

Olympus: A Universal Task Router for Computer Vision Tasks Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.697783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.697783Z digest=sha256:5f9a49ee4f234b930d3c1163d5c162057f1cfeb958a53394247d0a9bcff51ac9

Observation 43e2af3f-6910-4e58-b436-2b6b308b6e5c · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Olympus: A Universal Task Router for Computer Vision Tasks MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.713700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.713700Z digest=sha256:805bbcdf272e4b8b9404723e6e46c2490ae034bce4f2a22c60f93b28a43407b9

Observation ca5f6c16-ada8-4ff5-a221-193265025458 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Olympus: A Universal Task Router for Computer Vision Tasks Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.250717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.721779Z digest=sha256:54ee6485495c4fa2d6209d8dd6548ad9f316777087c886b1dea9160fe11c37bf

Observation a4504037-b643-4c5e-8e24-f40cf451b5dc · outbound

This paper cites Sigmoid loss for language image pre-training.

Olympus: A Universal Task Router for Computer Vision Tasks Sigmoid loss for language image pre-training

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.203355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.730548Z digest=sha256:61ddd4c0b03471a3f6157de34cd14008ac47959d2660c238122402849ef81609

Observation ae212811-3589-4a7f-8dc9-b8d64013523e · outbound

This paper cites AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling.

Olympus: A Universal Task Router for Computer Vision Tasks AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.741831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.741831Z digest=sha256:826e58cf024a5366819dead14ce991e3e33a99796c625e4c0255887942554e03

Observation 9ca8dcb7-4c0f-45fd-b97e-9d6e7669307f · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Olympus: A Universal Task Router for Computer Vision Tasks Adding conditional control to text-to-image diffusion models

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.179127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.752753Z digest=sha256:f35dbfa762b29489ab36c6e033ec9ff9767e1dc26208eef0a2875495c5f15f17

Observation 0adb6467-9173-4dfc-9e8b-4fa8e873a8c0 · outbound

This paper cites TinyLLaVA: A Framework of Small-scale Large Multimodal Models.

Olympus: A Universal Task Router for Computer Vision Tasks TinyLLaVA: A Framework of Small-scale Large Multimodal Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.775315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.775315Z digest=sha256:a02adbe3bef3f0fcfa0728fe25eb352e4da93ddc886a4058ba71841fa6d41715

Observation 0fd87351-a4e7-4ca9-8008-d1b6529d15d1 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Olympus: A Universal Task Router for Computer Vision Tasks Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.786978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.786978Z digest=sha256:a1bd089a299075634ca39c7d1d198efe61cfb9447990646a05f393c340e70a35

Observation 4d93e074-3274-4cd9-97fc-75fc5893ee8b · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Olympus: A Universal Task Router for Computer Vision Tasks MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.795666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.795666Z digest=sha256:9a95a937f5f595220df74559184b52680a1dc0577a348f5259c54024991428cd

Observation 2020f042-2401-46d7-aa31-dfdfa02ce3c7 · outbound

This paper cites VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation.

Olympus: A Universal Task Router for Computer Vision Tasks VL-GPT: A Generative Pre-trained Transformer for Vision and Language Understanding and Generation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.802539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.802539Z digest=sha256:81cebd3d2c544b39ebcc9ca2a6b60cc43353fd2365fa3df1433c78457cd9c849

Observation cc37c1ed-4bb5-499f-82f0-b8bfdaffeef7 · outbound

This paper cites Mipha: A comprehensive overhaul of multimodal assistant with small language models.

Olympus: A Universal Task Router for Computer Vision Tasks Mipha: A comprehensive overhaul of multimodal assistant with small language models

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:57:09.122407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T16:57:06.809177Z digest=sha256:16a94f5cd1b948b5699405f4abd4478703cf4e5205a065765851d8d619c6c39a

Observation 01a51fb4-2b19-467e-95e9-18ad1872f3f2 · outbound

This paper cites LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model.

Olympus: A Universal Task Router for Computer Vision Tasks LLaVA-Phi: Efficient Multi-Modal Assistant with Small Language Model

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T16:57:06.818547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:57:06.818547Z digest=sha256:c86109dfe289b15a274adcf4a374172bc51c9d2c67fe013b75390c6d27b34ae9

Pith citing papers

Observation 5107b992-66d4-4588-b9e9-4437c1efa220 · inbound

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation cites this paper.

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation Olympus: A Universal Task Router for Computer Vision Tasks

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:14:45.153165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T11:14:42.975508Z digest=sha256:18fc3b3ff17854f053a52a44ce2c882e74b1f3fe11d82f844eef28bd6782f771