Pith. sign in

Paper Citation Record · LEDGER

Cosmos 3: Omnimodal World Models for Physical AI

As of 10 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 52 inbound Pith citation observations for arXiv:2606.02800.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.02800 v4

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T15:08:33.957835Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 52 of 52 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T05:06:51.699806Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch9

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 90ab2315-cd92-4d4c-b723-4373ad1efe21 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Cosmos 3: Omnimodal World Models for Physical AI PaliGemma: A versatile 3B VLM for transfer

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.409103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:50f6e2e0bc754341d4ac0a67a2f88e6ca8b01b2b86a8808e60e6abde6b33c838

Observation 92d4c903-e3f6-41ab-b4b9-d406b0253b4e · outbound

This paper cites Internvla-a1: Unifying understanding, generation and action for robotic manipulation.

Cosmos 3: Omnimodal World Models for Physical AI Internvla-a1: Unifying understanding, generation and action for robotic manipulation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.364550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:940da35c48716ec69b75ef6394aaf91ac93db61d28228753b9adef1705d41171

Observation e3df3b5b-4aa8-4deb-9ede-690dc6f95bab · outbound

This paper cites GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation.

Cosmos 3: Omnimodal World Models for Physical AI GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.382387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:e28e2fc21f5fc65776b6a51564c11145632c5e547545394a498a753ed1257704

Observation 7fc89fb4-9072-4458-b829-3c7d59a981c5 · outbound

This paper cites Out of time: Automated lip sync in the wild.

Cosmos 3: Omnimodal World Models for Physical AI Out of time: Automated lip sync in the wild

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-28T15:08:33.957835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:b40ea2895bae78d2f54a3dc1467de25abe34a75975a91c0f43acea39cb79bb0d

Observation 82577fa1-45ae-4aa8-ba34-435ca39c7be7 · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

Cosmos 3: Omnimodal World Models for Physical AI NVLM: Open Frontier-Class Multimodal LLMs

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.387412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:57263f7515eb8cb80d4ff35c9fea0e73dc399620c8246ec258c0455c9023d3e4

Observation b7c293f8-d6ef-4633-9cef-b54d65c05d07 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Cosmos 3: Omnimodal World Models for Physical AI Emerging Properties in Unified Multimodal Pretraining

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.392456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:61e524c9ab2c9f59aea2c07018b9a691eabfbb0b91abc4881799ba4966667f68

Observation 26399458-3208-4e5f-ad80-20b554940130 · outbound

This paper cites VLMEvalKit: An open-source toolkit for evaluating large multi-modality models.

Cosmos 3: Omnimodal World Models for Physical AI VLMEvalKit: An open-source toolkit for evaluating large multi-modality models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T15:08:33.957835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:9c89fd8a63dd74494e3a19e36a62e0fe19f5014f17bc53c7be6ddd81a5e685d4

Observation a1e8a19a-7040-462b-b9c1-5a55972fb968 · outbound

This paper cites CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models.

Cosmos 3: Omnimodal World Models for Physical AI CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:36:18.020144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:4f8623a75560dbc17a2240df487dbfc0a2ded2bc23cac3e6420a41ac4507b8a5

Observation b47562f0-b06b-412d-a181-48b19ee21455 · outbound

This paper cites CameraCtrl: Enabling Camera Control for Text-to-Video Generation.

Cosmos 3: Omnimodal World Models for Physical AI CameraCtrl: Enabling Camera Control for Text-to-Video Generation

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:36:18.034862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:40ecb00893bf62d76d6432f0bb43c786e2bb262e61a9d1562aceb9f234d91445

Observation 3131c469-4361-4749-a281-04a1258d019f · outbound

This paper cites MolmoAct: Action Reasoning Models that can Reason in Space.

Cosmos 3: Omnimodal World Models for Physical AI MolmoAct: Action Reasoning Models that can Reason in Space

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.358146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:e37642f45a2acb5964172c848ecc109b96530db25ca4122c8d80a91f2483bc4e

Observation 56a83cb6-2c8d-42b5-86c2-232b9e46d46c · outbound

This paper cites SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes.

Cosmos 3: Omnimodal World Models for Physical AI SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.396613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:ae4e6aa6bb5318ab623b726875f354b5992f73ff7f1c222b41f33481f253d9eb

Observation 3b55fcdc-a366-48ec-bb80-095cba4609f7 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Cosmos 3: Omnimodal World Models for Physical AI GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.405483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:1d573314b8324d87ab1321a3a5b8addcef97358e4adeedfb98342ab90c6d880b

Observation 29b28eea-60da-4b26-8ade-f90d4a326b3a · outbound

This paper cites Learning to Act without Actions.

Cosmos 3: Omnimodal World Models for Physical AI Learning to Act without Actions

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T22:46:18.377940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:37a68ada8084111b22e256076463f7758a79e854668b66f72289aae918dbf922

Observation 0ed2161f-86dc-48bf-80ec-fcf968de0120 · outbound

This paper cites Video models are zero-shot learners and reasoners.

Cosmos 3: Omnimodal World Models for Physical AI Video models are zero-shot learners and reasoners

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:46:18.400971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:875a44ce9eb3bf7412380011a6ed3153ac05bdd4ed7df03dd4796e1ba5f47a56

Observation 61e6f52a-62a4-41f7-bfa8-7662d99a7d2d · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

Cosmos 3: Omnimodal World Models for Physical AI Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T22:46:18.371943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:08:33.957835Z digest=sha256:8defd27a46fae44c351b628c136a97f8cb4690cb8104d1fbaf9c07d5a1fbedf7

Pith citing papers

Observation b840df6a-923a-4426-b29f-9a4cd8b76843 · inbound

Critique of World Model cites this paper.

Critique of World Model Cosmos 3: Omnimodal World Models for Physical AI

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:42:27.116832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:42:27.116832Z digest=sha256:1ca0c844f4afcdd28683a16799ae141d31983baecc6787f86a8bc64228044ed2

Observation 29cba8cc-5629-4e9f-953f-62913f2fb8f9 · inbound

What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory cites this paper.

What Spatial Memory Must Store: Occlusion as the Test for Language-Agent Memory Cosmos 3: Omnimodal World Models for Physical AI

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:37:37.213686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T13:46:29.718269Z digest=sha256:e41cf6e8dad8bb6f2fdaf0569f2b845bf35a7e49cea9c576e3e7390f612e9c78

Observation 91adc934-0bf2-40d7-944c-8c6ddbb0fb3a · inbound

ActWorld: From Explorable to Interactive World Model via Action-Aware Memory cites this paper.

ActWorld: From Explorable to Interactive World Model via Action-Aware Memory Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:28:55.397948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T01:22:12.771098Z digest=sha256:bbd7176d637c933bb4b4ddc5000f0379fe6c4a40c822df2840181df188ca31e2

Observation 6ca8f406-2a4b-422f-9777-76ce08f88f85 · inbound

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation cites this paper.

PAIWorld: A 3D-Consistent World Foundation Model for Robotic Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T21:38:58.754179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T00:19:33.170645Z digest=sha256:20f2670e3ec98f6d29893643a489006a6330468c2cab1839de6819ef4874e760

Observation fa8e1f51-2122-4199-8976-eac1433f015c · inbound

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation cites this paper.

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:19:13.362817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T21:18:07.494794Z digest=sha256:4688e83a0dd99a2ddac4b9d39029452f4a676577f9c3a718b4a3ef495ac49fcf

Observation b6c19015-8863-419a-a00d-ae261b110023 · inbound

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation cites this paper.

SC3-Eval: Evaluating Robot Foundation Models via Self-Consistent Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:51.389791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T05:06:30.594399Z digest=sha256:b40d478fe5ba92511c36b740b32d53ac217b536f9f854ccac6c808d2874f1f46

Observation 4a382069-ca46-4061-8322-4b9cea202198 · inbound

Physics-IQ Verified cites this paper.

Physics-IQ Verified Cosmos 3: Omnimodal World Models for Physical AI

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:29:15.556417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T21:13:36.568700Z digest=sha256:c21c4e31a4090c889ea2d22e6c5ecdf2f13420dbb14f4bb1f288c19e0682a11d

Observation 891ebd45-3d91-4e59-b077-7546724672b3 · inbound

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? cites this paper.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Cosmos 3: Omnimodal World Models for Physical AI

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:39:17.503977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:0a27fa7ee596ef3c83ce2858b4e9b467dc3793a649183486d62652c05e84d7a6

Observation dc43e339-80bc-457f-acb7-8a04d4428f56 · inbound

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation cites this paper.

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:49:41.558859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T11:05:44.972128Z digest=sha256:a0c78d992db7b0b20d814495876ea25690b9b33437bcb5e6c1c1a465236e71e7

Observation 4fb51621-8fce-4a31-9f5a-9e337858f7a3 · inbound

Critique of Agent Model cites this paper.

Critique of Agent Model Cosmos 3: Omnimodal World Models for Physical AI

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-04T11:29:51.076095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T07:57:20.830605Z digest=sha256:ce9125899adc2af6f371d0ac0ebd78183a1e877649473146183c2480ef62345a

Observation 2d8ece93-2c98-486b-b0a7-44a037d7cca0 · inbound

DiffusionBench: On Holistic Evaluation of Diffusion Transformers cites this paper.

DiffusionBench: On Holistic Evaluation of Diffusion Transformers Cosmos 3: Omnimodal World Models for Physical AI

Reference 212

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T16:59:58.195744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T00:06:11.951205Z digest=sha256:5dd6ff866ad5996a4d47118df0b3bc156a80e17274495478e67b9636adaf5023

Observation 65459639-2a72-4e6e-b9b7-6ba7fbd8258e · inbound

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models cites this paper.

Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models Cosmos 3: Omnimodal World Models for Physical AI

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:50:11.445672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T20:57:30.765802Z digest=sha256:7b534b8b10435312a694312c7ad30361b127a43a23d7404d9efec984e7f5982c

Observation e1402191-1b5d-44fd-97dc-807539f2af53 · inbound

Learning Action Priors for Cross-embodiment Robot Manipulation cites this paper.

Learning Action Priors for Cross-embodiment Robot Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-04T21:00:09.820221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-25T19:09:56.409766Z digest=sha256:cf43f8c4c0d7fdef6f2f86c77c5609e946f576c8d2010fee8fbbd98a855b122c

Observation 0707065f-fc4d-47f5-9996-73e35bf64de4 · inbound

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation cites this paper.

PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T20:03:56.948658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T04:34:38.286863Z digest=sha256:05a7b64604580bd2b4b88eaa2fc1fffb959967baa9aed1c1d4d19435ca7535e6

Observation 42699cb3-ed34-4566-a74e-45d7db3f8368 · inbound

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis cites this paper.

RoboGaze: Evaluating Robot World Models via Structured Vision-Language Analysis Cosmos 3: Omnimodal World Models for Physical AI

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T10:54:35.850307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T10:49:58.019238Z digest=sha256:45ba5fb230675a2e9b9c2928117724f76ad167c4b053a39859b0e0aab8d0123f

Observation 8df2960d-1060-4c18-84e1-97ccce87eeca · inbound

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers cites this paper.

Mural: Transferring LLM knowledge to image generation via Mixture-of-Transformers Cosmos 3: Omnimodal World Models for Physical AI

Reference 20

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T09:24:32.131834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T09:23:59.752793Z digest=sha256:ed6d585dc1c8e6d9edef5f0a3ceb4766cb6e87c0f4ff9398898fe52f1eba34bd

Observation 8df1cee3-7063-430e-9ec5-4fda6ccc8145 · inbound

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration cites this paper.

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration Cosmos 3: Omnimodal World Models for Physical AI

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:25:41.706604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T05:31:43.678762Z digest=sha256:99fe6e5cf6e34bd7f1ca1a1d9962d1ceba25f6e34aab1f5025e1cb69e937a1a9

Observation 5aae7d9a-c867-495b-95f1-9df375d77f2f · inbound

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration cites this paper.

World Narrative Model for Highly Controllable Video Generation: A Paradigm Shift from Pixel Sampling to Physical World Orchestration Cosmos 3: Omnimodal World Models for Physical AI

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-15T10:20:59.147440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:20:59.147440Z digest=sha256:c7cde7c11a2b9134ba712baaf4a11806c715fd0ce736309e27412ebf0b6f6561

Observation 595f0b64-ad7b-46dc-a55b-0de06bfff587 · inbound

ROSA: A Robotics Foundation Model Serving System for Robot Factories cites this paper.

ROSA: A Robotics Foundation Model Serving System for Robot Factories Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-02T11:16:52.666121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T11:11:39.584755Z digest=sha256:2c7ca5765bf325123c03505b3e1acef4f46c85339ca45e9802ab229dbdd759e3

Observation 6e6bf47c-766d-4292-bb0c-2aeecf7e3122 · inbound

DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation cites this paper.

DynaWM: A Base-VLA-Guided World Foundation Model for Moving-Object Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T09:10:40.344868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:10:40.344868Z digest=sha256:73162d506f0a065c9aa54c7c3474dc523b26ff03a826fb54ebbb9e0246fb86e0

Observation e815add8-2379-4d1f-97f4-1b2bbe946ad3 · inbound

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation cites this paper.

GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T08:04:48.963890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:04:48.963890Z digest=sha256:ac7e359de1acc11782483c106e09b6f4938a9aa188dadf5a49f418a426591f77

Observation 2e3ebbcd-57e3-4dbe-8636-e49b96634871 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI Cosmos 3: Omnimodal World Models for Physical AI

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:2754d38d3d9aa08e9ea6dcfe5c56a72c02005000ada86f7afcd81164fd55209e

Observation 340b9926-bd9a-4491-b57a-94fc9fb807a6 · inbound

A Definition and Roadmap for World Models cites this paper.

A Definition and Roadmap for World Models Cosmos 3: Omnimodal World Models for Physical AI

Reference 128

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:44.336323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-08T07:10:33.826140Z digest=sha256:2d006295c1202f9308797bdce0e3a86b7db32466304eafa1b8af5bc7a136902f

Observation 5cfc9951-4d25-4d42-a0ba-7eb2e0803ad2 · inbound

From Foundation to Application: Improving VLA Models in Practice cites this paper.

From Foundation to Application: Improving VLA Models in Practice Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:14:45.485926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-08T07:06:13.473493Z digest=sha256:4dbd9cdd8231c83e5c6293a27abe978a9b97a714db78ab26e54a0bbc746b9859

Observation fea50311-ff49-4e5f-9c6d-776b88c0059e · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time Cosmos 3: Omnimodal World Models for Physical AI

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-13T06:48:14.554799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:48:14.554799Z digest=sha256:4e3b6c9478328beccba5202705cbd832cd292347dcd354dcdd2a529ea52ca446

Observation d70e5411-8c99-4851-8678-3fdf7bf4fa3f · inbound

Infinite Worlds with Versatile Interactions cites this paper.

Infinite Worlds with Versatile Interactions Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-09T07:56:04.717628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-09T07:51:36.802801Z digest=sha256:de7bb12713b2afe462d5c0db0e1395a81c404ad2969208641ffa481dd05bba84

Observation ad2ace2e-2a21-4fd2-b5d8-120e307a1492 · inbound

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence cites this paper.

Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Cosmos 3: Omnimodal World Models for Physical AI

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-09T03:05:55.157813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-09T02:55:58.018234Z digest=sha256:2edb5b938bd0e223a1e91cbd3fbe193a287aa30b95ce4f13592db6ec3011e394

Observation 25d1cb50-f7ab-4983-8cba-96e8942e7e39 · inbound

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model cites this paper.

Xiaomi-Robotics-U0: Unified Embodied Synthesis with World Foundation Model Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T04:10:14.360463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:10:14.360463Z digest=sha256:82b81259998b73f065d20edd6c3b1a218a45e0291a58b08775eda82e3e7690e4

Observation 1d25b611-2925-4a85-b2b2-b76883d1024e · inbound

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch cites this paper.

GigaWorld-Policy-0.5: A Faster and Stronger WAM Empowered by AutoResearch Cosmos 3: Omnimodal World Models for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:51.556914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:51.556914Z digest=sha256:01a89a8e04557f536ec63b646934140fdb563d9d83c7d8a23d742633c64fa6d1

Observation 1483f19f-b5a1-418b-8d45-1f748282ee72 · inbound

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination cites this paper.

RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T03:16:52.065229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T03:16:52.065229Z digest=sha256:9dc69ee7e20d849629bf0258ec43cef30d86b31dca3d948ad5c359e545c75a8c

Observation 605d0126-85e7-43ba-99c7-d5c89b8e032a · inbound

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories cites this paper.

Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Cosmos 3: Omnimodal World Models for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T00:03:59.225496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:03:59.225496Z digest=sha256:9c640c8e21f69ff0c2a90a1c997941c39ff5e611c0eb8b700fb578c27d37c0f5

Observation e68063ac-04dc-4b5f-9631-cedf5a97779f · inbound

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model cites this paper.

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Cosmos 3: Omnimodal World Models for Physical AI

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T16:32:50.585508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:32:50.585508Z digest=sha256:608aae877a09668248d1b15fc35cc9879d20a1c88e67fd89747fd4a93c06510a

Observation 4271b470-c198-47e8-85e7-0c3390a4ece4 · inbound

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model cites this paper.

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Cosmos 3: Omnimodal World Models for Physical AI

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T01:57:32.117952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:57:32.117952Z digest=sha256:03d9b2b967b39bbeff105f38199e2da047e0dfe086ce824bf7e9e0097880a7a9

Observation 1b1e1ac5-9340-47c9-9240-c51e842ff138 · inbound

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation cites this paper.

SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T07:09:04.165440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:09:04.165440Z digest=sha256:4a1ba58df90f59c372c074e286f64d04513b68bb2a0af9b121efbdf6d87577a7

Observation 0b6f689c-cfce-48c6-a633-3a01b8441ea5 · inbound

Data Pyramid for Embodied Manipulation cites this paper.

Data Pyramid for Embodied Manipulation Cosmos 3: Omnimodal World Models for Physical AI

Reference 269

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.948646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.948646Z digest=sha256:4ce56547562c12a6941109b56db74f1732ea40013bfad23c3b731be3d36a5bdf

Observation 88035d6b-db89-4b45-b581-abb7c189b321 · inbound

Visual prompt engineering for video models cites this paper.

Visual prompt engineering for video models Cosmos 3: Omnimodal World Models for Physical AI

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T02:13:12.652594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:13:12.652594Z digest=sha256:66041e6187b590a761dc5167fae9dd809c896f1b93fb9b34b41f79006b4afd6f

Observation 21c106ba-8c33-44da-9b27-8b625c2c64fa · inbound

Parallel Decoding Distillation for Fast Image and Video Generation cites this paper.

Parallel Decoding Distillation for Fast Image and Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T00:57:41.124597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:57:41.124597Z digest=sha256:0bf732ac2a8c770a13456ad986f25261a1299cf2e4d45d0e1454a14f6bf739a0

Observation 43e8959f-649f-43b1-9511-c109584963c0 · inbound

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control cites this paper.

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control Cosmos 3: Omnimodal World Models for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T11:27:05.644975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:27:05.644975Z digest=sha256:9535bfea1ed463ae7ff13af9039dc12a0bf8fcc4e6c8dc8c8a4d665022af016c

Observation a1cce6e4-7e97-4b7a-bcde-1d513273eb59 · inbound

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control cites this paper.

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control Cosmos 3: Omnimodal World Models for Physical AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T03:23:43.951163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T03:23:43.951163Z digest=sha256:519c999733ac9178d85715031cef2b2a65934a58ff72094df9647c1e8573afa6

Observation 7a7ef47b-6603-49b2-a56f-8cfc0b606d6f · inbound

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation cites this paper.

Security of World-Model-Based Embodied AI: A Lifecycle of Threats, Defenses, and Evaluation Cosmos 3: Omnimodal World Models for Physical AI

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-31T14:03:34.637523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:03:34.637523Z digest=sha256:7ca2998b260b5959517c4c8970414e805f45b6b7b34a9015f2ecb1be7c56ee9f

Observation dcd5bcfc-b7d5-4d95-863a-b365b362a3c1 · inbound

PhiZero: A World Model Built Around Physical Language cites this paper.

PhiZero: A World Model Built Around Physical Language Cosmos 3: Omnimodal World Models for Physical AI

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T01:50:30.578485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:50:30.578485Z digest=sha256:7db6f380c111706921f2c212e1a9d79628eff9475e8ce67aac152180a99add3b

Observation 0e74a7d6-7e9c-4f72-aa44-f800f65e07b4 · inbound

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation cites this paper.

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation Cosmos 3: Omnimodal World Models for Physical AI

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T04:58:51.666200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:58:51.666200Z digest=sha256:c9a97f4bc138a5967b0772aaf5e5ed6ac719fc4dddfc4877863fdb5731ef8359

Observation e1f898e5-30c6-47cd-9228-20be453563e1 · inbound

Test-Time Curriculum for Open-Set AIGC Detection cites this paper.

Test-Time Curriculum for Open-Set AIGC Detection Cosmos 3: Omnimodal World Models for Physical AI

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:23.328864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:46:23.328864Z digest=sha256:c7c1eb506cffe4375b7aeae94e0c1496eabcab440eb02ff4aed93677a29297e5

Observation 77f3dfff-8583-4639-8069-a4090c0be9b1 · inbound

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud cites this paper.

PhyAI: Real-Time Physical AI at the Edge, Scalable Rollouts in the Cloud Cosmos 3: Omnimodal World Models for Physical AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:52:35.403325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:52:35.403325Z digest=sha256:e2146cb9af261a77af9f7c3af73ab4216bccf429e5f8c8010cb7cd4c13bb352d

Observation 5f3d1e73-e7c2-49a9-ae03-56bc78ba8200 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Cosmos 3: Omnimodal World Models for Physical AI

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.091892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.091892Z digest=sha256:e94c9c6383a71a8845a07600311aa58fb0ea1a773654e8b04ce98814e1b55fe6

Observation 069a6161-a3b9-4834-8189-1a38a73cddac · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Cosmos 3: Omnimodal World Models for Physical AI

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:55.960160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:55.960160Z digest=sha256:aede41204ecd07a85dee90bc9976432866eadaacf0b41959269b628adec708b6

Observation 49e6124c-77a7-49d7-acd3-efdae777dd89 · inbound

HelloWorld: Enabling Socially Interactive Characters in Video World Models cites this paper.

HelloWorld: Enabling Socially Interactive Characters in Video World Models Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T06:07:34.380165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:07:34.380165Z digest=sha256:86f8f0f7baffbb5e0c3f16a7cafd9cec19648628a1537f071963785d323ae5c8

Observation d356f5bb-d61f-47db-95dc-b5ccc07656b0 · inbound

Uncertainty-Aware World Model for Aerial Image-Goal Navigation cites this paper.

Uncertainty-Aware World Model for Aerial Image-Goal Navigation Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T05:52:17.210476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:52:17.210476Z digest=sha256:914a2ff0c3f9579508e505c28fcc70e27c6a215c121f9165e8609167cd0a8cdc

Observation 27895903-1322-4270-a88c-ac4e21329396 · inbound

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments? cites this paper.

XEWorld: Can Action-Conditioned World Models Generalize to Unseen Robot Embodiments? Cosmos 3: Omnimodal World Models for Physical AI

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T23:21:53.436914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:21:53.436914Z digest=sha256:edfea7b2d33b82a4f63db6d0b0c500d0bf0224ce10388f5ea2b7a7e2b2d8e4d4

Observation 7638aba4-46c1-4a6a-95f2-3390386f3d08 · inbound

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models cites this paper.

GAUGE: A Measurement-Grounded Benchmark for Physical Fidelity in Simulation Engines and Video World Models Cosmos 3: Omnimodal World Models for Physical AI

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T20:36:25.642450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:36:25.642450Z digest=sha256:a9aa1128859ab9e0d5581107ef3afd7145246b1dc52940b85d4dd6e55f2e5013

Observation f93d8bf6-d09e-438f-9968-603c8bb60f13 · inbound

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction cites this paper.

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction Cosmos 3: Omnimodal World Models for Physical AI

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T19:27:25.020729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:27:25.020729Z digest=sha256:7dfc74af82e560901662120f9bda55091e5495baa8c8bf9d43cdf039a363fe08

Observation 91dddf21-8cd7-473a-a03a-66e3ae65cbcf · inbound

Addressable Memory for Video World Models cites this paper.

Addressable Memory for Video World Models Cosmos 3: Omnimodal World Models for Physical AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T05:06:51.699806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T05:06:51.699806Z digest=sha256:b4e87cc499ad53aa27170c6760e0c418e444b32d6a17af0aa61fdd3d8b3446ce