Pith. sign in

Paper Citation Record · LEDGER

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts

As of 8 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 1 inbound Pith citation observation for arXiv:2505.18451.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18451 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:13.994884Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:31:01.804061Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:07:26.477842Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee876670-2330-4ec4-836b-459be42943bf · outbound

This paper cites write newline.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.787474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.787474Z digest=sha256:7a14345ca14d10e8f5ee076fdb2e8d2b11c11502fe9c0d2991df18237c207d73

Observation f949c38f-395b-4b26-97f6-3d4c5b2efa6b · outbound

This paper cites GPT-4 Technical Report.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.791615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.791615Z digest=sha256:75014948f08a1b42489b0c578e225241a017e18d5884e94f047732377febffdd

Observation 46c4f29f-ac7b-42d1-9786-c20fe52bd580 · outbound

This paper cites and Frey, B.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts and Frey, B

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.444940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.794683Z digest=sha256:8c101efd8374ec3e1847d36f53e25438b7232c8796d2c7fcb81f8d7e0870c19b

Observation eef45565-0a43-42b4-9a08-abf148bccbae · outbound

This paper cites Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.798057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.798057Z digest=sha256:73638a5cba6ba8e489f836a549302527dd38c96549a85a02aa7b41a230de2c57

Observation 9d3130fa-c900-4122-8c3b-b58019f6b5e4 · outbound

This paper cites SparseLLM: Towards Global Pruning for Pre-trained Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts SparseLLM: Towards Global Pruning for Pre-trained Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.801474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.801474Z digest=sha256:70ba59a752dc332d68e8ca4b1e8cfc450272712ab01091e5a29e6ae8725b8651

Observation 8b8ad919-6076-4717-8ca9-34265405e27f · outbound

This paper cites Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.805695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.805695Z digest=sha256:73aec058a965b8f4516a8bfe240ac261aab633ef1586b2bead32cbd0469c2a84

Observation bdeb2113-6e95-425d-92ec-531e2e5e5e0f · outbound

This paper cites LoTR: Low Tensor Rank Weight Adaptation.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts LoTR: Low Tensor Rank Weight Adaptation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.808665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.808665Z digest=sha256:899e422c390a224b65580c5aa29885f0abf3660b9696c252e13b56391b304105

Observation f7f71b19-f93c-47d1-9c1d-abc0c40fc676 · outbound

This paper cites J., Frankle, J., and Guttag, J.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts J., Frankle, J., and Guttag, J

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.812664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.812664Z digest=sha256:ac31be933d3344fb3bc67a5cea13f40fff6e6aa45b401c7066fa667317950337

Observation 17da0630-eb63-4526-9f7e-bced5784154f · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.815632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.815632Z digest=sha256:20b7ed59f6b32e0bd364d455fa7e35b225eb0bcd714010a877ad1156ed423c36

Observation 49fa680f-6d77-4547-a8d2-557d02095262 · outbound

This paper cites an unresolved cited work.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:34:14.430130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.818973Z digest=sha256:8f8538db6fdf706ce585fc674d1b83cd78bed69d86201202a4e99634b9e2b5cc

Observation 5dd4b55e-e924-4d72-afe0-f14f479134ea · outbound

This paper cites Self-adaptive network pruning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Self-adaptive network pruning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.421042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.822074Z digest=sha256:b9a1474873a2e58d1c636bf411974db03dcd1cb5cf13398487d8afdc50ccd26a

Observation 170aabea-943f-4017-b918-40a89dbc05e1 · outbound

This paper cites SuperLoRA : Parameter-efficient unified adaptation for large vision models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts SuperLoRA : Parameter-efficient unified adaptation for large vision models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.411224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.824784Z digest=sha256:0151c999d484655ecbdb10a780d8aee8fb7e824c02b81e84c21335a0a99d2336

Observation 83cd6933-ec06-4df0-b1eb-f65ae61f62c2 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.827834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.827834Z digest=sha256:61008fc04498bf236278d4cd3eac0b66451867d2b31e9125465b0fe1191aa8a3

Observation 4bea1d77-3b4d-4ca1-9140-a9f5f1c369c1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.830800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.830800Z digest=sha256:80ddc6de8de2478424d97999d991486c6d5a1cfce271042e03b64dbbccadc6a8

Observation 26e789da-3908-4f1e-bffc-69b277cff1cc · outbound

This paper cites Learning to prune deep neural networks via layer-wise optimal brain surgeon.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Learning to prune deep neural networks via layer-wise optimal brain surgeon

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.401276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.833811Z digest=sha256:3ecbe9da6493d2137f549a46a49c47f4f83202b3555b7e56035c5e667545fe41

Observation 1f21c389-ea7c-4695-b1af-3e62eca7df98 · outbound

This paper cites KronA: Parameter Efficient Tuning with Kronecker Adapter.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts KronA: Parameter Efficient Tuning with Kronecker Adapter

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.836356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.836356Z digest=sha256:a00ddff0cdb1c1d561b5f7a8e6a5325e8c341348cbb7bca3aec19173797e3504

Observation ef101800-b5c3-4bef-a36b-b560552873b4 · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.839374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.839374Z digest=sha256:d590445ddc4686224bd41ece2ef3dadf2859dd3cf60f28abb20f33586aa03d46

Observation 1a50bdf6-daaa-4108-b0d4-5b427fbc4a18 · outbound

This paper cites and Alistarh, D.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts and Alistarh, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.392422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.842767Z digest=sha256:cbfea7b59518ed96ff6187c438c8d04cda1c2426e20f44f98c79929c4d07832c

Observation 50985bf4-0104-480d-9ccf-95369343f996 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.846159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.846159Z digest=sha256:62e08586ef35848cb74cce67793285c3bcf61071fd0af07937270349ef63797e

Observation 3c91d95b-5209-43c1-b8a8-b80365ffc148 · outbound

This paper cites Dynamic Channel Pruning: Feature Boosting and Suppression.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Dynamic Channel Pruning: Feature Boosting and Suppression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.849123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.849123Z digest=sha256:e2a1623e13168980b8021bbcb5a7f9384e8d37bffc97410106dfa9b8e9892f19

Observation 0752d8cf-1b0a-44a9-8ccf-c2f6b0c09bd4 · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.852329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.852329Z digest=sha256:2f1b0c07830b4a984ca6207ac81ce783e5f1f923390ce955fe1a1ccdb5b0eee3

Observation 5cb217d8-b2e4-47b1-8841-66d15dd0b820 · outbound

This paper cites Optimal brain surgeon: Extensions and performance comparisons.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Optimal brain surgeon: Extensions and performance comparisons

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.382469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.855376Z digest=sha256:87755e1ae0e1b470a936821216070a47ab72b9a2e735014033acf07898888bd8

Observation 39f65531-9169-446a-8f27-0c4fd5e8ebda · outbound

This paper cites Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.858580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.858580Z digest=sha256:e0dfe7e8cdf7ba47f0f97a36dcb4b5c554faa52fa0f3e0a147e5ed01842f4acb

Observation 28fc2637-ca29-40d8-b830-58b9af7f2d71 · outbound

This paper cites J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.373761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.861450Z digest=sha256:503d72a6813c168c4c9063f316bc13b8353f23083cdbd327137c00c4aa89c5b6

Observation a6cdd1bb-35e6-48ee-b2f9-4bd6f3ca3117 · outbound

This paper cites M., Zhang, Z., and Suh, G.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts M., Zhang, Z., and Suh, G

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.364521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.864543Z digest=sha256:f3bb85aaef5741c7d942b0120e460b4b06dbe84676881b4be4a3a4a4a7a3b280

Observation d9013bcf-e0e6-4bfa-a425-d17696d0c489 · outbound

This paper cites PC-LoRA: Low-Rank Adaptation for Progressive Model Compression with Knowledge Distillation.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts PC-LoRA: Low-Rank Adaptation for Progressive Model Compression with Knowledge Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.867971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.867971Z digest=sha256:e761aad721dad879c41c2a159aec64cb3b5498569f7b012d61f5dd0a4cae34be

Observation eb6f0547-a548-44c9-b1de-e76b457560b2 · outbound

This paper cites Mixtral of Experts.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Mixtral of Experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.871657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.871657Z digest=sha256:26f44dbf71074e5a26c0d21a750a755b97b4c8c0ca19b3e3a7bc9d56da4b1d66

Observation dd945c07-2587-4011-8445-152c43f3944b · outbound

This paper cites M., Bommarito, M.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts M., Bommarito, M

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.356067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.874682Z digest=sha256:68c3a8272d3a748a0825673d3d34db939afabb3612a77311ed2d0e3e36d27bf1

Observation 5adfa46e-d7cc-4f9c-8c71-6f722ba268b6 · outbound

This paper cites Quantum-PEFT: Ultra parameter-efficient fine-tuning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Quantum-PEFT: Ultra parameter-efficient fine-tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.877813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.877813Z digest=sha256:f89f46d9446de3b492339f39430ed7ccd16cc6fc8ae25b426a7a5316cfe1d661

Observation 6c3b664d-2096-4764-b95b-9e2381eac861 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Scaling Laws for Fine-Grained Mixture of Experts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.881134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.881134Z digest=sha256:30b29b614ed291a64b9c8c4a246130fd418c871f987c6d64549f517efd048817

Observation 0cfdc33f-14ff-4ffd-a6b9-8c4e6bea281e · outbound

This paper cites Optimal brain damage.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Optimal brain damage

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.884160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.884160Z digest=sha256:a4e832cd2856bf8ede98060878948bc8edfe5232c5bf93ead3fd849d03894d5a

Observation f33e0535-9d70-4269-8099-058e709249d7 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.887072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.887072Z digest=sha256:5c4bbfe77cabdeec43ac91bd99aaf3024f712e2f3d8ecdd23ca8abace3d4dbd4

Observation 26512cfe-c1cc-4579-9fc1-f63942fc3ef0 · outbound

This paper cites Runtime neural pruning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Runtime neural pruning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.340848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.890370Z digest=sha256:b833f899e2b17ce568436f020f204c587f9b4306436492dccdddefc3233f537d

Observation 76e9a70a-242e-40ce-a59f-4595f3a4f6c0 · outbound

This paper cites AWQ : Activation-aware weight quantization for on-device LLM compression and acceleration.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts AWQ : Activation-aware weight quantization for on-device LLM compression and acceleration

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.333035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.893509Z digest=sha256:74562bde939e89e0c339bbaebd3c5a54f49452dd9a5e4d68b6641b68c10a55b3

Observation 8927a97a-e79a-4fcb-a8d7-c3d2e3f3ece9 · outbound

This paper cites DeepSeek-V3 Technical Report.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts DeepSeek-V3 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.896670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.896670Z digest=sha256:8c0caf457dca32d00b29e06315bd212eb98bc19f388c04acabda30c549a7c98f

Observation 03eda01b-3e32-41be-8467-27aba183f195 · outbound

This paper cites an unresolved cited work.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:34:14.324017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.899793Z digest=sha256:7d51cf06dfd5e71a527288a5f7d110c267e43710ba5f99abe64ab039cca5e12c

Observation 3bab24ca-7da2-4940-82a7-61158f45f569 · outbound

This paper cites LoDA : Low-dimensional adaptation of large language models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts LoDA : Low-dimensional adaptation of large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.316279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.902989Z digest=sha256:d1f311cbf7d148736a47f1a1501dd3785299621a65f6e0a979a709672025c282

Observation 61563575-7d27-413b-8563-453d8dc7da78 · outbound

This paper cites and Deng, J.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts and Deng, J

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.307513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.910492Z digest=sha256:99751c79f0693b8ee6931ba9de8a8ccabaa18016c03a1e4e5395a65ab49a7d95

Observation d32545e3-f1a2-4292-b7cf-8afdb7c15a47 · outbound

This paper cites Deja vu: Contextual sparsity for efficient LLMs at inference time.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Deja vu: Contextual sparsity for efficient LLMs at inference time

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.299985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.912697Z digest=sha256:f3fa0a8432db8c6b19895cfa5856f6563b802fe22fde4f7d66a02b1d7a818170

Observation 96cf4937-1287-412e-a9e5-f09d44f4cf6a · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.915167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.915167Z digest=sha256:a2bf6452f30edeba10bc1d9d3c3d382bde03f5311562acce1dac2905f92c4167

Observation 6afc8624-bca3-443d-8f9b-b6d7c390a8ae · outbound

This paper cites LLM-Pruner : On the structural pruning of large language models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts LLM-Pruner : On the structural pruning of large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.287191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.917703Z digest=sha256:c710557db87bfa476a3b0b24b2cb798b73c443a54c1f09dee1800828942fbf90

Observation a74ca4bf-06c8-4a46-a61c-3d27bfb6f637 · outbound

This paper cites A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.279609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.920783Z digest=sha256:562e0ed8cd14159e28cd8b7f8ab9267c3729cbfe2bacf42666e5761d34adb7f0

Observation 50c8ad7f-3da4-4ee4-bd80-1e0aa9b5722f · outbound

This paper cites Pointer Sentinel Mixture Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Pointer Sentinel Mixture Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.923252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.923252Z digest=sha256:0974b6ede54122b857e0ed96ac49f7da18d89bf3c1ff69104fdb849bb88547f1

Observation 18549d90-38ae-4156-9f47-889ffbb24308 · outbound

This paper cites s1: Simple test-time scaling.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts s1: Simple test-time scaling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.925903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.925903Z digest=sha256:e1ec94e899b5f075f526fe2623b1ad073fdc4cccf58a3779de46e8c2254bf42f

Observation d236fffd-223c-4872-9ead-fbba587d87a9 · outbound

This paper cites an unresolved cited work.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.928816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.928816Z digest=sha256:a7e4d40678ff13ef3958c8f8d9dc3b9ebccf8957359cce498acea5e5a5529744

Observation 1f7b0fcc-ed40-4efd-9eb3-bd89f46f4956 · outbound

This paper cites Compressing large language models using low rank and low precision decomposition.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Compressing large language models using low rank and low precision decomposition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.267442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.931708Z digest=sha256:576424d2bafa31211e5b99702437a464f132ad71f7fa2ee5ed642a8c67a1bf80

Observation d1f11b3d-665c-4ddd-b9c9-f53f6941d291 · outbound

This paper cites Eigen Attention: Attention in Low-Rank Space for KV Cache Compression.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Eigen Attention: Attention in Low-Rank Space for KV Cache Compression

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.934567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.934567Z digest=sha256:1568e59cd53a4e97a0e42f031de59587469e67d940c463564f7b814f0f648a6b

Observation 02caee79-fff0-472b-bd1a-c8bbab540aca · outbound

This paper cites A., and Etzioni, O.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts A., and Etzioni, O

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.259185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.938191Z digest=sha256:08fc8f6e84db432077db5483e5919bf59238d809fda26a393838e41c945f0cd4

Observation 1aa9924e-3729-43b0-8e92-f803296899a9 · outbound

This paper cites Towards VQA models that can read.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Towards VQA models that can read

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.252038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.941617Z digest=sha256:a3c12554835a24e30f065a907d863dcba4444869b8375530779596c7b3497725

Observation e049edea-7f38-4cb7-a919-6349feb4e37a · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts A Simple and Effective Pruning Approach for Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.944075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.944075Z digest=sha256:2337c2da77b8de6c9623e3c464040d7a478aee08a541edc869194f432490d130

Observation e80ca720-5bda-4112-b45d-3cd1b3bd1363 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.947163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.947163Z digest=sha256:2850ad9beaeff6474ac72799d5ccb3d7b179b6d09fb3e77a168e4cb2d8ea06ae

Observation d6b19cdb-4c2a-495e-90ca-1dc772238a2a · outbound

This paper cites Neurons in Large Language Models: Dead, N-gram, Positional.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Neurons in Large Language Models: Dead, N-gram, Positional

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.950136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.950136Z digest=sha256:3bed7a5c92e1f4777189f8281f35f9096eacceddc5d876408d6d91aa1b425a00

Observation aee252e7-274c-4e48-941a-fc682c41b9f7 · outbound

This paper cites Q-VLM: Post-training Quantization for Large Vision-Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Q-VLM: Post-training Quantization for Large Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.953373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.953373Z digest=sha256:d87681c690d16e8499301493be667beca7632692547ca8a2ee4f51f25fbeb5f9

Observation 74356825-c5d8-4ae4-92fc-f145cd119c30 · outbound

This paper cites AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.956079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.956079Z digest=sha256:1423962a7190d6227745a8e05aae687058ac0f1953a3919cd7e08c5372b3ead7

Observation c80aa666-790e-4d4e-b564-876da6c1ae34 · outbound

This paper cites Emergent Abilities of Large Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Emergent Abilities of Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.959508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.959508Z digest=sha256:8eb8e3cbca17ca49caeb3aa5bbb87f51c06d826d1fb35323f1688554dd8ddf6c

Observation d6e59e3c-955e-45d6-9b20-4671e59d86f2 · outbound

This paper cites On the Impact of Calibration Data in Post-training Quantization and Pruning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts On the Impact of Calibration Data in Post-training Quantization and Pruning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.962578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.962578Z digest=sha256:3fdea45ddd2c460e6edf76851808aa8f197f6df8e1204bc1a454893d3a56f8ef

Observation 3a462e31-55e9-4610-b288-d2a854defeb3 · outbound

This paper cites Mixture of LoRA Experts.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Mixture of LoRA Experts

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.965720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.965720Z digest=sha256:6116b744c2e826bc93732d5597864ee8831b1edf4034f257ec00a7fc70cd1c4f

Observation 7fd4be8c-c017-4ec8-9b3e-8aea1444ce97 · outbound

This paper cites Automated fine-grained mixture-of-experts quantization.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Automated fine-grained mixture-of-experts quantization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.242855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.969148Z digest=sha256:80f1962a503c66a4897259759d484231d0f8002940213e8f73f7e09fcd77d6c6

Observation 8d23655c-c2c4-4b3b-a3f7-f75b967a4cfd · outbound

This paper cites and McAuley, J.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts and McAuley, J

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.235809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.972366Z digest=sha256:c4d752fd79e55102b0b61b8c019207fbc7c5daedd4ba059b0312eaf1d63c23e8

Observation 9a337ed3-03b6-45f9-8b61-c59b2437bec0 · outbound

This paper cites Dynamic DropConnect: Enhancing Neural Network Robustness through Adaptive Edge Dropping Strategies.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Dynamic DropConnect: Enhancing Neural Network Robustness through Adaptive Edge Dropping Strategies

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:34:14.043464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.975479Z digest=sha256:78f2f32a2190b57e0d19baadf3a98cc2ed4487825896a76324f56fae2747d23f

Observation fadc69d6-cdb5-4ef8-8e74-77bbbd2636ba · outbound

This paper cites B., Oh, G., and Gong, Y.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts B., Oh, G., and Gong, Y

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.228716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.978900Z digest=sha256:cb5574d1f0f564d9630944fe62e78eaf99347f946a65afb819beb0fcf8677a4e

Observation ea281f70-f4a6-4f13-a1c6-4c2b714b6f29 · outbound

This paper cites ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.981845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.981845Z digest=sha256:91f769928b5f8f7fc215477f8ce2b349f9a2b3f2b83f44a0c947550156d62e58

Observation 50a59906-288a-4c38-990d-e47abb93f1d2 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.984961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.984961Z digest=sha256:b32a1f4f069fe468d449064c621fe4244d74d4ca20f98cf8a098243c1c0ccaf6

Observation 6f5d9620-8a44-4c18-9b86-07c3e9a809a7 · outbound

This paper cites MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.988161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.988161Z digest=sha256:9ecd6fcc25c46832ddc38269fe1921a6fb6a19bf20f0d6c222292d825f00595f

Observation 976eab98-3f95-4cee-97ea-bc7c93868804 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts OPT: Open Pre-trained Transformer Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.991374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.991374Z digest=sha256:1dfce3069e50506d6de2973af71526621a26552d692214d1031dc5c87667a444

Observation 44f25647-acb7-479c-8089-72820110f2f4 · outbound

This paper cites A survey on model compression for large language models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts A survey on model compression for large language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.994884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.994884Z digest=sha256:563df857c62b8d2b31898ce2c3bd354945d45f64e8c2999eb0ca0620cfa79bd3

Pith citing papers

Observation ac90eb9b-e415-47cc-8e6d-dc15263edfaf · inbound

EinSort: Sorting is All We Need for Tensorizing LLM cites this paper.

EinSort: Sorting is All We Need for Tensorizing LLM $\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:07:26.479351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T18:31:01.804061Z digest=sha256:8b905c3ba55bcff193eb7be31f2bb2cee2a35cbd31c9f5a1679f6db4061e0984