Pith. sign in

Paper Citation Record · LEDGER

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2507.07818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07818 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:37:46.441294Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T06:02:52.244884Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T08:26:48.315943Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dd9b177f-34be-43b8-89c5-4f0e9559916c · outbound

This paper cites write newline.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.006424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.006424Z digest=sha256:41b1cb51c95efb4e41c66f20950a81083f6e9eabc173f9dbf8aebca555a08a58

Observation 41e8ba27-e195-4415-a321-bae0fcf75039 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.131103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.131103Z digest=sha256:8f0e152673048550b480f5dd31f62850b5fb74c93a299b1ff260c7876de8d455

Observation 226d1cf1-0197-4c11-93c2-7c9e1ece8f0c · outbound

This paper cites Stable LM 2 1.6B Technical Report.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Stable LM 2 1.6B Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.188523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.188523Z digest=sha256:f665f5470d9a51d06951b1b6ea76fdc78173178788916d06340ffa758382380b

Observation 035bce93-1f69-4de4-9c7f-0245ceb4ee3c · outbound

This paper cites nuScenes: A multimodal dataset for autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines nuScenes: A multimodal dataset for autonomous driving

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.254429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.254429Z digest=sha256:4d8cbebe45adf42b86248ee337236de71ead0f9ae10c056853fbde5495be9b43

Observation f46c68ad-b279-4d63-8b96-9937854540a5 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.370613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.370613Z digest=sha256:79ea186ecf2f55f3e16a605ec76f43ca07b92f6a9169ec65e9ed130d6be3767f

Observation f75bcb01-631b-40ea-904d-0508db5d7e0d · outbound

This paper cites Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.421592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.421592Z digest=sha256:8233e386bf234e23f0ba6c94288e82034e6ef5f87e03a62ee6c4dce846a0f4f5

Observation 37f2feec-768b-4037-bf41-91debb223555 · outbound

This paper cites Driving with llms: Fusing object-level vector modality for explainable autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Driving with llms: Fusing object-level vector modality for explainable autonomous driving

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:50.409725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:43.493213Z digest=sha256:0e6ddb8492f23fa7e73c09154627c66e81a65edd104a7d5def562592450dc43c

Observation 17a83839-8f01-40a8-ac31-e33322cd34e4 · outbound

This paper cites AdaMV-MoE : Adaptive multi-task vision mixture-of-experts.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines AdaMV-MoE : Adaptive multi-task vision mixture-of-experts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:50.091903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:43.577182Z digest=sha256:08327f6217ad6cce4b7b1943c2596d32055ae3986b4cc4f452822ff7562a9ca8

Observation 3387ea54-5524-4ebf-be5e-7f6bcbaaea40 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.654779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.654779Z digest=sha256:eccb02a924db66e0fb23de07a594cff4e58730d3f5743b570aaa83b26ac232f4

Observation bcd5ea30-12e8-45de-8fbc-928b722b9736 · outbound

This paper cites Qwen-vl-max: A high-performance vision-language model, 2024.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen-vl-max: A high-performance vision-language model, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:49.760319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:43.746929Z digest=sha256:1159be323c5b24a9c3fcbe7d3da79e72535b1675737178adc56e6337d4aa3c4d

Observation ceb72f98-e359-465e-8e4e-5268d0b94383 · outbound

This paper cites an unresolved cited work.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.824536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.824536Z digest=sha256:9383c356b8d195ebc918e68c243d4988aa11dc78d0e0e9a37f8a35c76d22454f

Observation 13687786-8b02-4c1c-9bed-2131bd069e15 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:43.912623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:43.912623Z digest=sha256:9a3f7191e22958f5ffddc4d29e5c97cce595b90b7ac2122f9c1b8ac09ab4e23f

Observation f1817b2e-acbc-4aea-97be-fdec8af77a8f · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.005943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.005943Z digest=sha256:01bcbaf6e13bcbd458104580ba37dd17a33ceb3a55233e913f5267b0aea7699d

Observation 1fbd3099-efba-4e77-aba1-8223580de194 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.073982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.073982Z digest=sha256:6da7e3098f391ad0c9cb19eb2bdcdb8d83d2fdd8558e560cb55599f33aa9d2a7

Observation 37dff330-5c94-4656-9c29-e46c03b36b6c · outbound

This paper cites From regional to general: A vision-language model-based framework for corner cases comprehension in autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines From regional to general: A vision-language model-based framework for corner cases comprehension in autonomous driving

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:49.466464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.175293Z digest=sha256:71bc68658ea7f7da44c639917d87020a7c2a8fcf566793bcc0c90a2955d8f305

Observation c39b43f7-553d-4fa8-b882-a8ce5af221a1 · outbound

This paper cites Planning-oriented autonomous driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Planning-oriented autonomous driving

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:48.849077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.227270Z digest=sha256:0aaadc9ed8a63c94e40a761a758d771f639b260a2f9fd8c062c0c7890f3c2903

Observation 208520fc-ef92-4241-8d61-1fd602e8cc42 · outbound

This paper cites RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.330575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.330575Z digest=sha256:27289f465ff7e2aa9ced88c3c614f8fd21f1170049ebb164376949b54ebf3092

Observation a50447ba-ee14-44df-8cae-aa3de68d0a6a · outbound

This paper cites Making large language models better planners with reasoning-decision alignment, 2024 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Making large language models better planners with reasoning-decision alignment, 2024 b

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:48.360150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.391895Z digest=sha256:1e48485af2c3bd3934c860d4e472abef889adddd37cdd5ddcb0f4b650edd39e7

Observation 705c16b4-2634-4fa7-834e-4c878261629f · outbound

This paper cites an unresolved cited work.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T18:37:48.176300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.429974Z digest=sha256:958788e9421926cf5cfc8ad8c36ae8cfc7dd97c4a31d4ae70dbd14fe20e1164d

Observation 08425c3b-cea7-4a62-89f6-060f4486ea08 · outbound

This paper cites Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Med-MoE: Mixture of Domain-Specific Experts for Lightweight Medical Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.507398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.507398Z digest=sha256:b24fb3d7b88af1e50252068c132def76465295f0e42f3e805f6223e7ff770173

Observation 15c83ebb-9597-477f-9ab8-140bf0a5e143 · outbound

This paper cites GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.567764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.567764Z digest=sha256:1172ee619e6703ababdc7834dc0346bd295c4aedee78ba4518d3630014e62981

Observation f3024295-eb79-4046-bdc1-e521673d8a7e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.668987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.668987Z digest=sha256:fb673fdd7feb4a542d7b10be02e1b85f131ba836ef0e92bf136b270311263c5a

Observation 08b1fab0-830a-4e8d-b47b-88599f0f41cb · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:48.030031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.737427Z digest=sha256:0934dbec041a9a182caf11441d41d4aa370adf1c423fb2af6274c96554204eec

Observation a49f957b-7b22-46b0-8cdf-7bf838e74df8 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.808524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.808524Z digest=sha256:2209cce5e7909cf4d46c2ba65ff55797fcdbac8b38dadef6793e48c7a35b903c

Observation 2c2e9f44-4840-4713-a710-4540a1339876 · outbound

This paper cites Moma: Efficient early-fusion pre-training with mixture of modality-aware experts, 2024 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Moma: Efficient early-fusion pre-training with mixture of modality-aware experts, 2024 b

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.808005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:44.865125Z digest=sha256:1069a37c4ca204899839f0159fdec68b3969434b88d047132a05dd68bfead52b

Observation e38584df-0703-4f92-bb3a-649120cbda23 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023 a.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Improved baselines with visual instruction tuning, 2023 a

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:44.972759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:44.972759Z digest=sha256:7e389c2b22e2d84a7f96cc41d72614de4e3e46c45ee94cefea2cb55aaa5eb3c7

Observation 8613ffa9-7c0a-45e5-856e-b692bbb3df22 · outbound

This paper cites Visual instruction tuning, 2023 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Visual instruction tuning, 2023 b

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.037569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.037569Z digest=sha256:339b160721672993a1a6effcb22d06bdb2594ea53a9aeb06af41c410b07a1917

Observation 54b403ca-5d3b-4914-8e15-9e01f8d2f9bd · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.119217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.119217Z digest=sha256:2a96a4f9e594cdacb589af1077b987818002c807df6273c011691cfe33cce852

Observation 7525c862-a7e6-4949-b811-0db14d56237f · outbound

This paper cites GPT-Driver: Learning to Drive with GPT.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines GPT-Driver: Learning to Drive with GPT

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.201566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.201566Z digest=sha256:80db599e2184a3a67186a7384d59fd28ae9f1e326d2a9a4a9f2923c20da9a5f2

Observation c8f031f5-16fd-4c46-9e18-a20906993e08 · outbound

This paper cites Gpt-4v system card.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Gpt-4v system card

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.608696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:45.275342Z digest=sha256:49c2f4538eff8b69fd72b85a5fda80a2e349ab7bdb0caf51fc9b02f3fff07337

Observation b4a901a6-6a8b-4ad1-896b-39bd0ebdaf35 · outbound

This paper cites LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines LanguageMPC: Large Language Models as Decision Makers for Autonomous Driving

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.330042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.330042Z digest=sha256:c70ef198979915482dfbccdb64b5e2ded2263aecdb57b2eb88ae7d6f874d8b50

Observation 8e9f60ed-40dd-46b7-a0a6-c7f113b717c8 · outbound

This paper cites Lmdrive: Closed-loop end-to-end driving with large language models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Lmdrive: Closed-loop end-to-end driving with large language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.428519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:45.416140Z digest=sha256:8d3ed3f9025b65d160712918f0c9342c22dde4dd56c5138c542fc6bc8e38e193

Observation bf62d979-6ff7-4c40-8fc3-0f5d9bc1b76f · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.491881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.491881Z digest=sha256:7bed6d2ef27f48ee6defff379ca300e64be45120f4fc1f12d7f8c1d688fbcf2f

Observation 21e183e7-d9bf-463b-b53d-feccc26bf41e · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines DriveLM: Driving with Graph Visual Question Answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.568121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.568121Z digest=sha256:8087f6be9c803697c2c1a2967e60713f9f55e574b7d59bff422574680c7e7879

Observation 23b76a4e-41e6-4102-909b-d05e0d7b7417 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.673966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.673966Z digest=sha256:bd2168bfab04ee0dc0803172aa375507f33ec93d3ff4bdae03734afbb736f9dc

Observation b07d8080-d6a9-4313-9b60-0fe51ef14d6f · outbound

This paper cites OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.752864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.752864Z digest=sha256:c5c2b08991f2d66dab6650f3bfc02b6ac9c396f42157657cc9ec99544fddf72d

Observation ab119678-ae2b-4583-be42-ab3add7e1a52 · outbound

This paper cites On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines On the Road with GPT-4V(ision): Early Explorations of Visual-Language Model on Autonomous Driving

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.837334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.837334Z digest=sha256:490718b33c88579d8b9976ad4d4e9935f87d2b591195400d64060e679fbb43af

Observation d789d401-3624-4b30-a015-00980ab687e3 · outbound

This paper cites Two-stage lvlm system: 1st place solution for eccv 2024 corner case scene understanding challenge.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Two-stage lvlm system: 1st place solution for eccv 2024 corner case scene understanding challenge

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.244661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:45.918065Z digest=sha256:26609258fb988fd0758abc8e224e77216585a5fde16cd7812714b45ad74b5e21

Observation 6d778d01-a12e-4015-a18a-b2fe06d7fc4e · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:45.972426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:45.972426Z digest=sha256:98b4629d39d5f8f2d013e8099b81a7c0118b1eea738166127af6dddfb1c8d4c0

Observation b2be8034-3639-4772-9c79-34d983dd8a63 · outbound

This paper cites Flex-moe: Modeling arbitrary modality combination via the flexible mixture-of-experts.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Flex-moe: Modeling arbitrary modality combination via the flexible mixture-of-experts

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:47.083633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:46.073957Z digest=sha256:7e575e7a2087c1d3c8964e774e15ff54b6fc7411954ac8f164903321b3fe2934

Observation e0d118b6-0f14-48c4-b74d-3d9dad7c8798 · outbound

This paper cites Sigmoid loss for language image pre-training.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Sigmoid loss for language image pre-training

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:46.195794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:46.195794Z digest=sha256:ccedf6c86f7aff60f2ccf7eb072ee53d238194dca8fae3052e0e86146f6483f4

Observation 5f813ce0-2054-427d-b11b-70dab2f5dd13 · outbound

This paper cites MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines MiniDrive: More Efficient Vision-Language Models with Multi-Level 2D Features as Text Tokens for Autonomous Driving

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:46.269978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:46.269978Z digest=sha256:7de0da89e0c66ec21fa72796ed9d4d6220ddfe503ba87b6574158e9c48654ef6

Observation e04c8bf4-71a5-4603-8a24-c4349f920eec · outbound

This paper cites Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts, 2024 b.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines Diversifying the expert knowledge for task-agnostic pruning in sparse mixture-of-experts, 2024 b

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:37:46.894735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T18:37:46.338934Z digest=sha256:1af8a8ef1dfa1f915885a5f4a27454fff289494bd9781c962754c59834c25935

Observation c3dd1e9b-0d51-4115-8e98-174d7f77f9a8 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:37:46.441294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:37:46.441294Z digest=sha256:ce375dd170d85a7cdf8cf0314552b55a13a49fc5f0827eaa2e80df2199ceffab

Pith citing papers

Observation 6c0db8a3-3f58-49a1-8d04-192b6c2c2481 · inbound

D$^3$-MoE:Dual Disentangled Diffusion Mixture-of-Experts for Style-Controllable End-to-End Autonomous Driving cites this paper.

D$^3$-MoE:Dual Disentangled Diffusion Mixture-of-Experts for Style-Controllable End-to-End Autonomous Driving MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:26:48.317451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T06:02:52.244884Z digest=sha256:a84c8050ec8203ea21f65dbffe826dd6f3aad03676a26d616ab3e14586d6b212