Pith. sign in

Paper Citation Record · LEDGER

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines

As of 17 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2507.07818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07818 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:37:46.441294Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:02:52.244884Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T08:26:48.315943Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dd9b177f-34be-43b8-89c5-4f0e9559916c · outbound

This paper cites write newline.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.006424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.006424Z digest=sha256:109bd4db4ec67778dcc1b5a4c5a3e086d6cb258a82ba92d182cd71441636a5bd

Observation 41e8ba27-e195-4415-a321-bae0fcf75039 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.131103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.131103Z digest=sha256:4f8cae5980ea814c88b71578c9f3a1e81f5b898846239e860153964538a36edf

Observation 226d1cf1-0197-4c11-93c2-7c9e1ece8f0c · outbound

This paper cites Stable LM 2 1.6B Technical Report.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Stable LM 2 1.6B Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.188523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.188523Z digest=sha256:67bbab6b84c288aae9453f56e5da9ebd728888e6bf4bf97cd7b4ff6b445e1d68

Observation 035bce93-1f69-4de4-9c7f-0245ceb4ee3c · outbound

This paper cites nuScenes: A multimodal dataset for autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines nuScenes: A multimodal dataset for autonomous driving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.254429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.254429Z digest=sha256:a995ddf4842da1394f28fd44f1e19a3e29b5d1075dee53bb38b470ad72f1e39f

Observation f46c68ad-b279-4d63-8b96-9937854540a5 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.370613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.370613Z digest=sha256:ed72d4267a2acf5637489854cc6c2ee2bb945c03923f079cc3f420d7a4fd98c4

Observation f75bcb01-631b-40ea-904d-0508db5d7e0d · outbound

This paper cites Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.421592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.421592Z digest=sha256:8c5e5df6cd33eef24ebd8dcec59961983938f5ff765b9c12a37901ba52c44bbe

Observation 37f2feec-768b-4037-bf41-91debb223555 · outbound

This paper cites Driving with llms: Fusing object-level vector modality for explainable autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Driving with llms: Fusing object-level vector modality for explainable autonomous driving

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:50.409725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:43.493213Z digest=sha256:980b02a576bf1c3656acab23e2d3a94c21f4a57064119937bdb8e4e052e0197c

Observation 17a83839-8f01-40a8-ac31-e33322cd34e4 · outbound

This paper cites AdaMV-MoE : Adaptive multi-task vision mixture-of-experts.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines AdaMV-MoE : Adaptive multi-task vision mixture-of-experts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:50.091903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:43.577182Z digest=sha256:225cfb052d4b6d83321dc5bfb05632a26b85cc46a3d91514cbeb6711d4fd3b93

Observation 3387ea54-5524-4ebf-be5e-7f6bcbaaea40 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.654779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.654779Z digest=sha256:c90fe510519b84f6417782ce435c6a27e5f8c948070a8bbbc3d396db6c5c9f0d

Observation bcd5ea30-12e8-45de-8fbc-928b722b9736 · outbound

This paper cites Qwen-vl-max: A high-performance vision-language model, 2024.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen-vl-max: A high-performance vision-language model, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:49.760319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:43.746929Z digest=sha256:831122f98e055a5a4adcdef85de80a7d249b63241f382a69d122e82e5f9e7de0

Observation ceb72f98-e359-465e-8e4e-5268d0b94383 · outbound

This paper cites an unresolved cited work.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.824536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.824536Z digest=sha256:96e95f23c5c88b81b270fc85b871e01dd16df725ebcea0ba50c5c7581d67a088

Observation 13687786-8b02-4c1c-9bed-2131bd069e15 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.912623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.912623Z digest=sha256:85dccb919807ad1423ec2d86888da16402a786b9c115665edb6bf91d8dc0423c

Observation f1817b2e-acbc-4aea-97be-fdec8af77a8f · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.005943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.005943Z digest=sha256:0e75746051f2a8822dbd20ef43f82301bdd2b74847ec204b840f588194277c8b

Observation 1fbd3099-efba-4e77-aba1-8223580de194 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.073982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.073982Z digest=sha256:d0782becaa7aa6c996637b7dd3e22e17c1c09b9884756ad91e02f93e3f1b03dc

Observation 37dff330-5c94-4656-9c29-e46c03b36b6c · outbound

This paper cites From regional to general: A vision-language model-based framework for corner cases comprehension in autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines From regional to general: A vision-language model-based framework for corner cases comprehension in autonomous driving

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:49.466464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.175293Z digest=sha256:489c1752ffc4f80bc2915823aa78dc8b6056b6cbdf704aec59311ecb94492b59

Observation c39b43f7-553d-4fa8-b882-a8ce5af221a1 · outbound

This paper cites Planning-oriented autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Planning-oriented autonomous driving

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:48.849077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.227270Z digest=sha256:a1c7a3f05506db71d55255bf27fcd9075de5f0f43f448a6c9c4a2c6071ce3ede

Observation 208520fc-ef92-4241-8d61-1fd602e8cc42 · outbound

This paper cites RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.330575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.330575Z digest=sha256:d62592cfb4f04911963dc34beef4ba555d8094cbe093943f49ba0ebf252d867b

Observation a50447ba-ee14-44df-8cae-aa3de68d0a6a · outbound

This paper cites Making large language models better planners with reasoning-decision alignment, 2024 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Making large language models better planners with reasoning-decision alignment, 2024 b

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:48.360150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.391895Z digest=sha256:17d750bb9085d62d67ae1676f0d05511bc97b9e4098ae4f3d55c79278b2737c2

Observation 705c16b4-2634-4fa7-834e-4c878261629f · outbound

This paper cites an unresolved cited work.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:37:48.176300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.429974Z digest=sha256:ecd89a20e240eaecec9667e0a0b3d0f5e2719289e94d4846f69c2782742dac24

Observation 08425c3b-cea7-4a62-89f6-060f4486ea08 · outbound

This paper cites Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.507398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.507398Z digest=sha256:9525a2883590af602e4757d0f5cb4785e5305f3a6373e0513c725fec1fa6b6ae

Observation 15c83ebb-9597-477f-9ab8-140bf0a5e143 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.567764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.567764Z digest=sha256:d9c0c876baeeb63a4dc970343b8f1fc9205895103985f1c7dcadd0d58d2c2fe9

Observation f3024295-eb79-4046-bdc1-e521673d8a7e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.668987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.668987Z digest=sha256:ea539d0650f177d9b74c556d8d6806fa11c1b10ef1772bb50c029b5d854e7de6

Observation 08b1fab0-830a-4e8d-b47b-88599f0f41cb · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:48.030031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.737427Z digest=sha256:68db8dedee05112e083e1e2fbf9fbbcc1ce31a7236344b2bc031071b25cf5e31

Observation a49f957b-7b22-46b0-8cdf-7bf838e74df8 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.808524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.808524Z digest=sha256:b508023456e84e284a4c664622c3458d01739a14d7bfb755a8609f81340405a7

Observation 2c2e9f44-4840-4713-a710-4540a1339876 · outbound

This paper cites Moma: Efficient early-fusion pre-training with mixture of modality-aware experts, 2024 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Moma: Efficient early-fusion pre-training with mixture of modality-aware experts, 2024 b

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.808005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.865125Z digest=sha256:b55b8ed778f488e7a5e34ce3192c498335cfb87860f8d20d4027dfbbe28992b3

Observation e38584df-0703-4f92-bb3a-649120cbda23 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023 a.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Improved baselines with visual instruction tuning, 2023 a

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.972759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.972759Z digest=sha256:bfad0de14d7ab646695fe5ba2834fb11b4e448db0a6fbe4cf9d17461f71d1787

Observation 8613ffa9-7c0a-45e5-856e-b692bbb3df22 · outbound

This paper cites Visual instruction tuning, 2023 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Visual instruction tuning, 2023 b

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.037569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.037569Z digest=sha256:effd48d4f61281fc38b69afc76a799c09c21a78f0fb9785c2a22847a222c417f

Observation 54b403ca-5d3b-4914-8e15-9e01f8d2f9bd · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.119217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.119217Z digest=sha256:085e3a7aeae1197c6729bc0fd89476e52f31ca57dda3d73bda28abf22ce4b4ce

Observation 7525c862-a7e6-4949-b811-0db14d56237f · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines GPT-Driver: Learning to Drive with GPT

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.201566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.201566Z digest=sha256:60bf21f44ad7840e111179610250d6402da2f31993c28561299071c095f25e2b

Observation c8f031f5-16fd-4c46-9e18-a20906993e08 · outbound

This paper cites Gpt-4v system card.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Gpt-4v system card

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.608696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:45.275342Z digest=sha256:d9a3d9f399345c024a6967ee289285815df4b9868979573a2bc74cd523b92384

Observation b4a901a6-6a8b-4ad1-896b-39bd0ebdaf35 · outbound

This paper cites LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.330042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.330042Z digest=sha256:f4b5d8efd596f3389d016e37ef6047551877124a4048be20d7bfba3199533ae3

Observation 8e9f60ed-40dd-46b7-a0a6-c7f113b717c8 · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Lmdrive: Closed-loop end-to-end driving with large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.428519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:45.416140Z digest=sha256:e3da20973a845789ec4246d2c74aaa1a2e865129a578ed14f4ae133240d2da06

Observation bf62d979-6ff7-4c40-8fc3-0f5d9bc1b76f · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.491881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.491881Z digest=sha256:d799b6a3dd5b626140554658b1be6ee039406c2cc922d2268d15154237bf1ed6

Observation 21e183e7-d9bf-463b-b53d-feccc26bf41e · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines DriveLM: Driving with Graph Visual Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.568121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.568121Z digest=sha256:496cb9f500d47f3cbe5e2accbe4d4e5afef1fe2282a3c02677c48c67c63a9775

Observation 23b76a4e-41e6-4102-909b-d05e0d7b7417 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.673966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.673966Z digest=sha256:d14909fcf28ae2b6610d44464300b6416fdd79463f4d70cd4486c2dd83c6e86c

Observation b07d8080-d6a9-4313-9b60-0fe51ef14d6f · outbound

This paper cites OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.752864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.752864Z digest=sha256:2f33b69755992dd744d7d47963462e17a32c146c000cb37abaf9bff569fa2c1f

Observation ab119678-ae2b-4583-be42-ab3add7e1a52 · outbound

This paper cites On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.837334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.837334Z digest=sha256:c7bfba07810a270fbc53ab8c64e43e0d3237db82d8924f74bace35617dd3fc0b

Observation d789d401-3624-4b30-a015-00980ab687e3 · outbound

This paper cites Two-stage lvlm system: 1st place solution for eccv 2024 corner case scene understanding challenge.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Two-stage lvlm system: 1st place solution for eccv 2024 corner case scene understanding challenge

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.244661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:45.918065Z digest=sha256:35046c7dcede4a51a9794d6d4fc3afbff93f781d9945e2477ca71615328bc38b

Observation 6d778d01-a12e-4015-a18a-b2fe06d7fc4e · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.972426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.972426Z digest=sha256:6fdbdb004469b3127ba00673e37fee4de9dce476118153bbd96a6c693c7a3f9c

Observation b2be8034-3639-4772-9c79-34d983dd8a63 · outbound

This paper cites Flex-moe: Modeling arbitrary modality combination via the flexible mixture-of-experts.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Flex-moe: Modeling arbitrary modality combination via the flexible mixture-of-experts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.083633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:46.073957Z digest=sha256:0bcb6684637e3b3303104ac2ab9726f78fea6c8ae306cb9d6514e33aaad4e3a9

Observation e0d118b6-0f14-48c4-b74d-3d9dad7c8798 · outbound

This paper cites Sigmoid loss for language image pre-training.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Sigmoid loss for language image pre-training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:46.195794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:46.195794Z digest=sha256:749078176487e5ec560ac26344cf40cb22e5beec06c6422d13464ac7cbae7c83

Observation 5f813ce0-2054-427d-b11b-70dab2f5dd13 · outbound

This paper cites MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:46.269978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:46.269978Z digest=sha256:d86bbb89336222ef4fbfa36aee1ce923abdf70ccdff2b2e69f785112ff510fc6

Observation e04c8bf4-71a5-4603-8a24-c4349f920eec · outbound

This paper cites Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts, 2024 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts, 2024 b

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:46.894735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T18:37:46.338934Z digest=sha256:e0cd715a00d8044b5a5c594eb194f545ec989b38b6afd0aeffae6b6ce6cafdd4

Observation c3dd1e9b-0d51-4115-8e98-174d7f77f9a8 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:46.441294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:46.441294Z digest=sha256:dfb3b83df56013eb76fff8b4d92ad9229a3b2e8764547c6c99b15d4763b48eda

Pith citing papers

Observation 6c0db8a3-3f58-49a1-8d04-192b6c2c2481 · inbound

D$^3$-MoE:Dual Disentangled Diffusion Mixture-of-Experts for Style-Controllable End-to-End Autonomous Driving cites this paper.

D$^3$-MoE:Dual Disentangled Diffusion Mixture-of-Experts for Style-Controllable End-to-End Autonomous Driving MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:26:48.317451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T06:02:52.244884Z digest=sha256:942225bf4bb2f8f70bc2c9427051f47364cba7224e1097d746a96cba26dd8439