Pith. sign in

Paper Citation Record · LEDGER

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts

As of 10 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 1 inbound Pith citation observation for arXiv:2505.18451.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18451 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:34:13.994884Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:31:01.804061Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:07:26.477842Z

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy22
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee876670-2330-4ec4-836b-459be42943bf · outbound

This paper cites write newline.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.787474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.787474Z digest=sha256:accee9f0ff31baf7acfe4df5c789950bea40125664b97798fbcdbf991239cdb6

Observation f949c38f-395b-4b26-97f6-3d4c5b2efa6b · outbound

This paper cites GPT-4 Technical Report.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.791615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.791615Z digest=sha256:742f4990889d34876b4e9f0176b67d357e439e287ff2464c4f3189e75f348314

Observation 46c4f29f-ac7b-42d1-9786-c20fe52bd580 · outbound

This paper cites and Frey, B.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts and Frey, B

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.444940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.794683Z digest=sha256:f982c9de2ed0dd777967f472f64259a5a1e1e16479c1506d173d8148521f57c0

Observation eef45565-0a43-42b4-9a08-abf148bccbae · outbound

This paper cites Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.798057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.798057Z digest=sha256:9fe187225de0d69a3a11df47cac0f2f5580d5c8560e516187103af05a89f9a67

Observation 9d3130fa-c900-4122-8c3b-b58019f6b5e4 · outbound

This paper cites SparseLLM: Towards Global Pruning for Pre-trained Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts SparseLLM: Towards Global Pruning for Pre-trained Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.801474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.801474Z digest=sha256:fd58746f53f19d89d7784637293025b38a54858360ada0eb4dc1e83cc09afa27

Observation 8b8ad919-6076-4717-8ca9-34265405e27f · outbound

This paper cites Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Rethinking the Role of Scale for In-Context Learning: An Interpretability-based Case Study at 66 Billion Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.805695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.805695Z digest=sha256:49d168f2190e8e713db0dede423d24f3b6c4c68046e43b6fc34bd784f45620de

Observation bdeb2113-6e95-425d-92ec-531e2e5e5e0f · outbound

This paper cites LoTR: Low Tensor Rank Weight Adaptation.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts LoTR: Low Tensor Rank Weight Adaptation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.808665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.808665Z digest=sha256:6672d7a8b055867ea23aa10d01f92054021c4fe542e036c249d0af98150dba17

Observation f7f71b19-f93c-47d1-9c1d-abc0c40fc676 · outbound

This paper cites J., Frankle, J., and Guttag, J.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts J., Frankle, J., and Guttag, J

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.812664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.812664Z digest=sha256:f03a0d3cbb04026279b924d424b87409f6c509aea555cc8138b9fbbaa431aa00

Observation 17da0630-eb63-4526-9f7e-bced5784154f · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.815632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.815632Z digest=sha256:1b233348e5051a19b63c38b3f5cab362742c6cea1c3b5aadf281355a338c3fa5

Observation 49fa680f-6d77-4547-a8d2-557d02095262 · outbound

This paper cites an unresolved cited work.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:34:14.430130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.818973Z digest=sha256:ac4e36ec54466704b20a969cafd066fc18540ea598eb5dfb207c956da1492ab4

Observation 5dd4b55e-e924-4d72-afe0-f14f479134ea · outbound

This paper cites Self-adaptive network pruning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Self-adaptive network pruning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.421042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.822074Z digest=sha256:d3d2e1110deea2ed5597cceb2c5b68d62c9e1c414f298bedc84fdc15b321e644

Observation 170aabea-943f-4017-b918-40a89dbc05e1 · outbound

This paper cites SuperLoRA : Parameter-efficient unified adaptation for large vision models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts SuperLoRA : Parameter-efficient unified adaptation for large vision models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.411224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.824784Z digest=sha256:70adf7dc156cf3c39f0c24d9614cbd60b23f695bfe81e1f0222eec4f6a967a35

Observation 83cd6933-ec06-4df0-b1eb-f65ae61f62c2 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.827834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.827834Z digest=sha256:4c434919ff2a35e3b527c00deac8b12907e900a91f33f16e53f3f10980b8765c

Observation 4bea1d77-3b4d-4ca1-9140-a9f5f1c369c1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.830800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.830800Z digest=sha256:9381b20e4dba6ed8dd50dc7c4282dfd435634f956fefc02651ce871b8a83aa9d

Observation 26e789da-3908-4f1e-bffc-69b277cff1cc · outbound

This paper cites Learning to prune deep neural networks via layer-wise optimal brain surgeon.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Learning to prune deep neural networks via layer-wise optimal brain surgeon

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.401276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.833811Z digest=sha256:e3fe47e13ce3107edef3a7515418433e0ecf06c067b2dcb075749abf2bbbcb72

Observation 1f21c389-ea7c-4695-b1af-3e62eca7df98 · outbound

This paper cites KronA: Parameter Efficient Tuning with Kronecker Adapter.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts KronA: Parameter Efficient Tuning with Kronecker Adapter

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.836356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.836356Z digest=sha256:59f0815f282bc5afa7b223541561a88b7d41617a605b538accda93d1161d4ae6

Observation ef101800-b5c3-4bef-a36b-b560552873b4 · outbound

This paper cites The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.839374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.839374Z digest=sha256:56f94dcb8797ea42acbe906a370483acc4afb0fd98b47f974cd18c87a4ca1c75

Observation 1a50bdf6-daaa-4108-b0d4-5b427fbc4a18 · outbound

This paper cites and Alistarh, D.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts and Alistarh, D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.392422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.842767Z digest=sha256:711d7716a0f25bd9b3339e96c9020b20077e07f5adb4ba1038fe530513c3741f

Observation 50985bf4-0104-480d-9ccf-95369343f996 · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.846159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.846159Z digest=sha256:d63a15bcf7c37075f5404323566227794b40bb0e9166c93fe3dbca2ea3acd50d

Observation 3c91d95b-5209-43c1-b8a8-b80365ffc148 · outbound

This paper cites Dynamic Channel Pruning: Feature Boosting and Suppression.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Dynamic Channel Pruning: Feature Boosting and Suppression

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.849123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.849123Z digest=sha256:c354e02f204efbd3ba7d036f2602048d26bd46f44e9c27d05fcf42d37ec5b458

Observation 0752d8cf-1b0a-44a9-8ccf-c2f6b0c09bd4 · outbound

This paper cites Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.852329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.852329Z digest=sha256:fbca03ab4b21d94ab41aa4215085124029cee5721fd5391136ab0a7ba93b16f4

Observation 5cb217d8-b2e4-47b1-8841-66d15dd0b820 · outbound

This paper cites Optimal brain surgeon: Extensions and performance comparisons.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Optimal brain surgeon: Extensions and performance comparisons

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.382469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.855376Z digest=sha256:fee0ef5abe7d738e90964e04eda148def247f5bea89a5181ea738b45f3d49a21

Observation 39f65531-9169-446a-8f27-0c4fd5e8ebda · outbound

This paper cites Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Distilling Step-by-Step! Outperforming Larger Language Models with Less Training Data and Smaller Model Sizes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.858580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.858580Z digest=sha256:e86d89ae846922c17d2f22f618ba4f795d27e96b05b25221de0f466da4947def

Observation 28fc2637-ca29-40d8-b830-58b9af7f2d71 · outbound

This paper cites J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., Chen, W., et al

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.373761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.861450Z digest=sha256:574c502a7533cab3b4d9237981fb9858a130b8e0550394790281fbc24d172cf5

Observation a6cdd1bb-35e6-48ee-b2f9-4bd6f3ca3117 · outbound

This paper cites M., Zhang, Z., and Suh, G.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts M., Zhang, Z., and Suh, G

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.364521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.864543Z digest=sha256:51e31ea966e766b87e2c7c22f1ef330f4b8d3b89dc91b07180f38205db5e110f

Observation d9013bcf-e0e6-4bfa-a425-d17696d0c489 · outbound

This paper cites PC-LoRA: Low-Rank Adaptation for Progressive Model Compression with Knowledge Distillation.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts PC-LoRA: Low-Rank Adaptation for Progressive Model Compression with Knowledge Distillation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.867971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.867971Z digest=sha256:b0b1bd52ab5978863163ca28ad1f43434f508b0520cd38c83ff9fdf18c550327

Observation eb6f0547-a548-44c9-b1de-e76b457560b2 · outbound

This paper cites Mixtral of Experts.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Mixtral of Experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.871657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.871657Z digest=sha256:2fe0a799f0d2a4416db70f732f3a8843911da29b6094e333799598932a5a71f9

Observation dd945c07-2587-4011-8445-152c43f3944b · outbound

This paper cites M., Bommarito, M.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts M., Bommarito, M

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.356067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.874682Z digest=sha256:ef760abda3ff11e61a77364e9ece523d004b64ec299f70fc4679f7475bb794f3

Observation 5adfa46e-d7cc-4f9c-8c71-6f722ba268b6 · outbound

This paper cites Quantum-PEFT: Ultra parameter-efficient fine-tuning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Quantum-PEFT: Ultra parameter-efficient fine-tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.877813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.877813Z digest=sha256:19eb31ec9413f5afc229ce17e396e2c0b3797427e5f8643da92f3f106875b39f

Observation 6c3b664d-2096-4764-b95b-9e2381eac861 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Scaling Laws for Fine-Grained Mixture of Experts

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.881134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.881134Z digest=sha256:3d4fba8d72f7ff0889967d92714a14dfa307248870d8542de59408040f474061

Observation 0cfdc33f-14ff-4ffd-a6b9-8c4e6bea281e · outbound

This paper cites Optimal brain damage.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Optimal brain damage

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.884160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.884160Z digest=sha256:c64ccda8d45ec361af275c2ddbabdef1d59cd0a3bbf9a1951549458e2a384c0b

Observation f33e0535-9d70-4269-8099-058e709249d7 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.887072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.887072Z digest=sha256:bec611d49ae46cab2b031e6362344c5f5024efb8d2fed68104e435e2256b080a

Observation 26512cfe-c1cc-4579-9fc1-f63942fc3ef0 · outbound

This paper cites Runtime neural pruning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Runtime neural pruning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.340848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.890370Z digest=sha256:3b38ae2863211aca9d7372f6e5a3afc785a4c184308f8270f4352ad28398cc6a

Observation 76e9a70a-242e-40ce-a59f-4595f3a4f6c0 · outbound

This paper cites AWQ : Activation-aware weight quantization for on-device LLM compression and acceleration.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts AWQ : Activation-aware weight quantization for on-device LLM compression and acceleration

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.333035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.893509Z digest=sha256:2e6ffbfd9192d8c4ba94ccece666cb5dc62d4ac449f31cad3fa93ee39d761c63

Observation 8927a97a-e79a-4fcb-a8d7-c3d2e3f3ece9 · outbound

This paper cites DeepSeek-V3 Technical Report.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts DeepSeek-V3 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.896670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.896670Z digest=sha256:93f43b2f6b40565b1386d66b481bdeee9e8246d8126b51b3f8ab168efaa19d75

Observation 03eda01b-3e32-41be-8467-27aba183f195 · outbound

This paper cites an unresolved cited work.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:34:14.324017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.899793Z digest=sha256:b67cdd4c33270c94672a17c2d5fda5c472ab922cb0e9492b11af0b49d5a26f94

Observation 3bab24ca-7da2-4940-82a7-61158f45f569 · outbound

This paper cites LoDA : Low-dimensional adaptation of large language models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts LoDA : Low-dimensional adaptation of large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.316279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.902989Z digest=sha256:b7ee2f0ede842311211e8554ef98334cf5848ccff5eb95162a7f6c1de8226dc4

Observation 61563575-7d27-413b-8563-453d8dc7da78 · outbound

This paper cites and Deng, J.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts and Deng, J

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.307513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.910492Z digest=sha256:53317b277cb9ecbe0de8da989acb2b7ec2c46119f72cc0a60487fe714a981a7e

Observation d32545e3-f1a2-4292-b7cf-8afdb7c15a47 · outbound

This paper cites Deja vu: Contextual sparsity for efficient LLMs at inference time.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Deja vu: Contextual sparsity for efficient LLMs at inference time

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.299985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.912697Z digest=sha256:fbb19db17c69a954575751cb0e7bea9c6083b988b6dfc4df668486dd0e35bfb7

Observation 96cf4937-1287-412e-a9e5-f09d44f4cf6a · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.915167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.915167Z digest=sha256:543abe376de85fabfc4fcb366853cf698eee3dd35db45023f2b240002951100e

Observation 6afc8624-bca3-443d-8f9b-b6d7c390a8ae · outbound

This paper cites LLM-Pruner : On the structural pruning of large language models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts LLM-Pruner : On the structural pruning of large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.287191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.917703Z digest=sha256:293700c2ff87e4c986179fc655abf7725ec93ca85ff3dc3b7daea35d736b39b0

Observation a74ca4bf-06c8-4a46-a61c-3d27bfb6f637 · outbound

This paper cites A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts A., MacIntyre, R., Bies, A., Ferguson, M., Katz, K., and Schasberger, B

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.279609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.920783Z digest=sha256:d474520cb69b06db68c8c61ea31df51a405143da74d5d93987c6ff499d353373

Observation 50c8ad7f-3da4-4ee4-bd80-1e0aa9b5722f · outbound

This paper cites Pointer Sentinel Mixture Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Pointer Sentinel Mixture Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.923252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.923252Z digest=sha256:68ef4deb310c288fb519f62820aa3370bd52f8123cd48a7e432a9c3ea8158e0d

Observation 18549d90-38ae-4156-9f47-889ffbb24308 · outbound

This paper cites s1: Simple test-time scaling.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts s1: Simple test-time scaling

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.925903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.925903Z digest=sha256:8ca191ddc6276aa5c8634f4b490784e546662680959ac6f313cfa42656e03d10

Observation d236fffd-223c-4872-9ead-fbba587d87a9 · outbound

This paper cites an unresolved cited work.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.928816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.928816Z digest=sha256:45235e69e2a901b18620a1e4a018e243423ed3b7878beb01f3b01b3916588f27

Observation 1f7b0fcc-ed40-4efd-9eb3-bd89f46f4956 · outbound

This paper cites Compressing large language models using low rank and low precision decomposition.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Compressing large language models using low rank and low precision decomposition

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.267442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.931708Z digest=sha256:396678bd8d8078cb027b4740c76a6cca56ec3ed13f53e15fc6ea3d6327ec7d1c

Observation d1f11b3d-665c-4ddd-b9c9-f53f6941d291 · outbound

This paper cites Eigen Attention: Attention in Low-Rank Space for KV Cache Compression.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Eigen Attention: Attention in Low-Rank Space for KV Cache Compression

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.934567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.934567Z digest=sha256:1d944233c489f35e69c0dd5d72636d9b175d38ef7e9fe3ce4ac2eb724634d911

Observation 02caee79-fff0-472b-bd1a-c8bbab540aca · outbound

This paper cites A., and Etzioni, O.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts A., and Etzioni, O

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.259185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.938191Z digest=sha256:9fbfa1dc2cddcc9dfa2d12404b2375a4fab3a3832e341b2c107c85c2b2b941d5

Observation 1aa9924e-3729-43b0-8e92-f803296899a9 · outbound

This paper cites Towards VQA models that can read.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Towards VQA models that can read

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.252038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.941617Z digest=sha256:1e666ddcc22059c9f0ff880fddf31973c0b39d843a12697d12ad68512227d6bb

Observation e049edea-7f38-4cb7-a919-6349feb4e37a · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts A Simple and Effective Pruning Approach for Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.944075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.944075Z digest=sha256:0595b7f560640869312d41ca0984f10924af63a33cd327d288def5583104e91d

Observation e80ca720-5bda-4112-b45d-3cd1b3bd1363 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.947163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.947163Z digest=sha256:4fb66eba7a91957828c904d3ac5a4572d935ff02c5b39c41c20031adaec33d78

Observation d6b19cdb-4c2a-495e-90ca-1dc772238a2a · outbound

This paper cites Neurons in Large Language Models: Dead, N-gram, Positional.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Neurons in Large Language Models: Dead, N-gram, Positional

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.950136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.950136Z digest=sha256:4af6ec1c3c70af105c83c9dfa65046dea4ecb45ebac736ed2ec85bd60f6e1c2e

Observation aee252e7-274c-4e48-941a-fc682c41b9f7 · outbound

This paper cites Q-VLM: Post-training Quantization for Large Vision-Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Q-VLM: Post-training Quantization for Large Vision-Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.953373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.953373Z digest=sha256:75b1486d081d52066732de18c255c4550b52a66b0597c3cd6b561163f94a7513

Observation 74356825-c5d8-4ae4-92fc-f145cd119c30 · outbound

This paper cites AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts AdaMix: Mixture-of-Adaptations for Parameter-efficient Model Tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.956079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.956079Z digest=sha256:91e908b67366e7197c1a776e1661619764dcc44b4de59762459a716ce59b8e52

Observation c80aa666-790e-4d4e-b564-876da6c1ae34 · outbound

This paper cites Emergent Abilities of Large Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Emergent Abilities of Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.959508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.959508Z digest=sha256:43802663f2f57e9bcd90a7b8a38a8880b28b5b6be7cda92aed0fb73a63b63b29

Observation d6e59e3c-955e-45d6-9b20-4671e59d86f2 · outbound

This paper cites On the Impact of Calibration Data in Post-training Quantization and Pruning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts On the Impact of Calibration Data in Post-training Quantization and Pruning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.962578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.962578Z digest=sha256:c77c0f8a77c011061cdd181e862877fed3808c32123d41e27e8602a0243103f8

Observation 3a462e31-55e9-4610-b288-d2a854defeb3 · outbound

This paper cites Mixture of LoRA Experts.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Mixture of LoRA Experts

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.965720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.965720Z digest=sha256:6e16f1cb692fb92088b60fd8e4cb0752c0c4486846758e472ef6cbdfd73afb76

Observation 7fd4be8c-c017-4ec8-9b3e-8aea1444ce97 · outbound

This paper cites Automated fine-grained mixture-of-experts quantization.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Automated fine-grained mixture-of-experts quantization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.242855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.969148Z digest=sha256:4696132edb862f234789898d8c796c6d76f8181ae626a77e1d56431a4a20feb8

Observation 8d23655c-c2c4-4b3b-a3f7-f75b967a4cfd · outbound

This paper cites and McAuley, J.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts and McAuley, J

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.235809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.972366Z digest=sha256:8bdf057801f56240bdeab9abe8e860fc934c2e2835043cf419a1a1fc3b869f94

Observation 9a337ed3-03b6-45f9-8b61-c59b2437bec0 · outbound

This paper cites Dynamic DropConnect: Enhancing Neural Network Robustness through Adaptive Edge Dropping Strategies.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts Dynamic DropConnect: Enhancing Neural Network Robustness through Adaptive Edge Dropping Strategies

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:34:14.043464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.975479Z digest=sha256:865eee6535de9a745d46880ceec8c42613c5aca6c88065625034da0b89db7f61

Observation fadc69d6-cdb5-4ef8-8e74-77bbbd2636ba · outbound

This paper cites B., Oh, G., and Gong, Y.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts B., Oh, G., and Gong, Y

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:34:14.228716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T14:34:13.978900Z digest=sha256:fefbf95880683f1862bb2bd6c1f2cf241b955868b9fafc996bbcc7a6c7ea81d5

Observation ea281f70-f4a6-4f13-a1c6-4c2b714b6f29 · outbound

This paper cites ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.981845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.981845Z digest=sha256:dc83f1ae5d2fc6564dba903c90dc4c3571d7204935ce789efa35b125b5569404

Observation 50a59906-288a-4c38-990d-e47abb93f1d2 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.984961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.984961Z digest=sha256:c6c0f320605a41087d789fc5365b764dab7817e8ce5c94cbfe18d06d55256cd0

Observation 6f5d9620-8a44-4c18-9b86-07c3e9a809a7 · outbound

This paper cites MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts MiLoRA: Efficient Mixture of Low-Rank Adaptation for Large Language Models Fine-tuning

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.988161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.988161Z digest=sha256:8df6766630142b9b3d30ac65fb9a2a048be1756d17f4b6969200c90bffa32cbf

Observation 976eab98-3f95-4cee-97ea-bc7c93868804 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts OPT: Open Pre-trained Transformer Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.991374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.991374Z digest=sha256:da1a7266490b58d78807d17e91227b1cdf9f9e2ce330115f09eaf3b1e4227329

Observation 44f25647-acb7-479c-8089-72820110f2f4 · outbound

This paper cites A survey on model compression for large language models.

$\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts A survey on model compression for large language models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:13.994884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:13.994884Z digest=sha256:05032538e4e32ba69702491d07ce57e1087f4bd6c053b8293a95f8adde86dc48

Pith citing papers

Observation ac90eb9b-e415-47cc-8e6d-dc15263edfaf · inbound

EinSort: Sorting is All We Need for Tensorizing LLM cites this paper.

EinSort: Sorting is All We Need for Tensorizing LLM $\mu$-MoE: Test-Time Pruning as Micro-Grained Mixture-of-Experts

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-02T23:07:26.479351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T18:31:01.804061Z digest=sha256:b043cf2e0d9e887f443c139668834aef15e176d6749511b11a5e7ce6ff8ae0ae