Pith. sign in

Paper Citation Record · LEDGER

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information

As of 23 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.03510.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03510 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:06:41.062180Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5556fa6b-2f0c-49df-abfd-fe95bc6ab418 · outbound

This paper cites Slicegpt: Compress large language models by deleting rows and columns.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Slicegpt: Compress large language models by deleting rows and columns

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:47.075297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:37.206536Z digest=sha256:bf0f8df1df33a8e4fbb3b6bfc863a45feb8d63738fd1d9ea86bc9a78cba47466

Observation aa1fee7e-3982-41fb-ad3d-30f40759b92d · outbound

This paper cites • ARC-Challenge and ARC-Easy[Clark et al., 2018] are datasets composed of grade-school level multiple-choice science questions.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information • ARC-Challenge and ARC-Easy[Clark et al., 2018] are datasets composed of grade-school level multiple-choice science questions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.653879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.829414Z digest=sha256:66587dea68f1c5d0a092e98fb9e8b4d0ecb5c92c1d6d051f2a02097c31a9a7d3

Observation e01af565-2563-4ba0-ac88-66611601c497 · outbound

This paper cites Pea-kd: Parameter-efficient and accurate knowledge distillation on bert.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Pea-kd: Parameter-efficient and accurate knowledge distillation on bert

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:46.934334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:37.414283Z digest=sha256:4147c3220bb87177d699d091bfec788d4fab1f73576905621857435d5b27ef66

Observation 38237da4-63d0-476f-9c26-bc362fbc1804 · outbound

This paper cites This demonstrates that five candidates are sufficient to find the sublayer with the least importance score.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information This demonstrates that five candidates are sufficient to find the sublayer with the least importance score

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.342946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.983994Z digest=sha256:d9ef6f73ff8cd6df8226ee262a9946f203bd8992f59bda31c735c7c4cd7f9cdf

Observation 662f402c-6d1d-4514-8715-854e9007b60e · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:45.856973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:37.893937Z digest=sha256:39b53347f38243fc30c935432a27d1c8f058f54249e70f87af8a69c816874762

Observation 7a9db158-1b3b-4329-bf3c-fedc0359d932 · outbound

This paper cites OPTQ: Accurate quantization for generative pre-trained transformers.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information OPTQ: Accurate quantization for generative pre-trained transformers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:45.688354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:37.954853Z digest=sha256:1852f2b1ca71fba902a5d06b74603d8678b695fe689dba5f3adb201345418532

Observation 0dd00334-b931-47a4-a8d4-f03df3027c78 · outbound

This paper cites Lazyllm: Dynamic token pruning for efficient long context llm inference.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Lazyllm: Dynamic token pruning for efficient long context llm inference

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:45.418760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:38.036607Z digest=sha256:f8a15e4f613d1ec92122fa7545a5087f2364f19ce06c0caf50c068c25b70c9f4

Observation 9c5207f5-51a4-4bd3-85b5-6cedc6725560 · outbound

This paper cites A frame- work for few-shot language model evaluation, 12.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information A frame- work for few-shot language model evaluation, 12

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:45.239781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:38.140621Z digest=sha256:c2a6b5dc5edaf49c6f80e2bee2c24f5e6e22112ef011cf97c92adf0d8a64a2f3

Observation 22e84dfc-d61e-4113-b2c8-148e3b09b1f1 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:45.068132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:38.218182Z digest=sha256:e846e810ac1d9fab11b06979ac07aef40649b8b87d63514eabda7e3c793075ab

Observation b75d018c-4b48-4a64-b762-152809afba74 · outbound

This paper cites Falcon: lightweight and accurate convo- lution based on depthwise separable convolution.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Falcon: lightweight and accurate convo- lution based on depthwise separable convolution

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.989062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:38.333525Z digest=sha256:dc7b1ec4d9134622c3a8df12538c17fa081a65b7d064dae8ee9310ddec9de472

Observation 90598526-0a8a-4c41-b45d-74f3b64d6507 · outbound

This paper cites an unresolved cited work.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:06:44.876957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:38.442187Z digest=sha256:0405ee6344d2209581aeafad91c661d2899cc3a05ac8e05b062e17ab86ae9c17

Observation 61086d0b-0683-4ab0-8507-943d620cd40c · outbound

This paper cites an unresolved cited work.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:06:44.783332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:38.550725Z digest=sha256:1077901e313b9ef7f0e56f8e0a88f2b1a7480eb98d3e39f72fc5dc6105b6205f

Observation dfd76e9a-163d-41eb-9838-4cd5703574b0 · outbound

This paper cites an unresolved cited work.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:06:44.580464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:38.743850Z digest=sha256:e2366608cfdea9eb219cd0f5bc2cfd08813c9bb0382a53fd8e2fb7fee82dfc5c

Observation 5235280a-ba91-4152-8169-b2c466d0e57f · outbound

This paper cites DDK: distilling domain knowledge for efficient large language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information DDK: distilling domain knowledge for efficient large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.347291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:38.951134Z digest=sha256:7a16ad4a0f0d47e5a601a2ac637ef28ce233f99a5edd390d1d1a2391752063e6

Observation 5634f1d0-04b7-40a3-bf6b-b7e174af23a0 · outbound

This paper cites Llm-pruner: On the structural pruning of large lan- guage models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Llm-pruner: On the structural pruning of large lan- guage models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.243254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:39.074936Z digest=sha256:3947144013676aa5d0f933412e21f60899e78738404e48269bd40827459b0f91

Observation bbaf37f5-70f7-46e7-9ad7-22660095dc21 · outbound

This paper cites Shortgpt: Layers in large language models are more redundant than you expect.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Shortgpt: Layers in large language models are more redundant than you expect

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.109191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:39.230392Z digest=sha256:7f631b4c1c3c52c9a7fa17b009c2ac8c6e405552e46c1193a800c00c4ad06d94

Observation 12d06413-1a80-4661-852a-7e0e5163877e · outbound

This paper cites Pointer sentinel mixture models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Pointer sentinel mixture models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.997597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:39.380345Z digest=sha256:6814eb056a6d1417aef6f10b5d74244edd2cfc789efd4e8d40c5f5669feb80fd

Observation 0f0aae67-aac6-4589-b922-b40e8174a583 · outbound

This paper cites Sensimix: Sensitivity-aware 8-bit index & 1-bit value mixed precision quantization for bert compression.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Sensimix: Sensitivity-aware 8-bit index & 1-bit value mixed precision quantization for bert compression

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.866307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:39.571943Z digest=sha256:f08e061c02306cc7cbf87be1bda62f635a44f35da339b68f4f610f452f5acbff

Observation 69e25463-b4f2-48c4-a910-89cb9613309a · outbound

This paper cites Mixture-of-depths: Dynamically al- locating compute in transformer-based language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Mixture-of-depths: Dynamically al- locating compute in transformer-based language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.775359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:39.666897Z digest=sha256:3e2557a00ca37c4e1054282a9c13b7781fdb0d227113277880b189d58079421e

Observation 2e88c0cc-9fc4-499e-8b3a-59aee2d02567 · outbound

This paper cites Exaone 3.0 7.8 b instruction tuned language model.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Exaone 3.0 7.8 b instruction tuned language model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.642116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:39.844672Z digest=sha256:f381d7d29563efa24aca8287e2c87a8a416998461272861dded9b2dcfbaccc9e

Observation 8fa26fd2-3181-48d5-b2bb-ec9cf99d3a37 · outbound

This paper cites Confident adaptive language modeling.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Confident adaptive language modeling

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.540375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.001779Z digest=sha256:c822531b11ff7da8129311b29b9dae9e1cc00b84b36798ea95515ef1f605c6a3

Observation 010dc779-83e3-4e6f-acb5-d19675455bbf · outbound

This paper cites Omniquant: Omnidi- rectionally calibrated quantization for large language mod- els.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Omniquant: Omnidi- rectionally calibrated quantization for large language mod- els

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.373199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.117420Z digest=sha256:6fdb4541fc2fb78829be97b8ef2e27499ec6e754297f758f5aed29f739115c41

Observation 4b1d685d-2559-4158-89a4-b03b97987496 · outbound

This paper cites Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.227336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.178601Z digest=sha256:30c28df341a97c63c98b1be2eba9917fd905fe5c6d45694c74a22839824a02f3

Observation 2140cf69-46e2-4cc5-b1dc-37371f098a7d · outbound

This paper cites A simple and effective pruning approach for large language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information A simple and effective pruning approach for large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.038919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.242222Z digest=sha256:6603dd2204ee8634916dffca30fcb28b2e8a776e1c100cc8521fde7d3fa7bde8

Observation 97329a70-27f7-43d8-870c-50f274ae52ed · outbound

This paper cites Gemini: a family of highly capable multimodal models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Gemini: a family of highly capable multimodal models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.890677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.313820Z digest=sha256:b530361a38ce90a087bfd0fad3ddc51248ac5493cb0cdd87adc348b94c6359c4

Observation 472f7a99-b653-4b6b-bb98-fe8f39d00ee4 · outbound

This paper cites Accelerating llama infer- ence by enabling intermediate layer decoding via instruc- tion tuning with lite.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Accelerating llama infer- ence by enabling intermediate layer decoding via instruc- tion tuning with lite

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.736879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.358457Z digest=sha256:1474f059321060e8e5be2bf3a3be8811f339faa07fd7cc96ff4f0013b2e0e559

Observation b6ec1720-b2c9-45df-8ca1-06f156af4020 · outbound

This paper cites Attention is all you need.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Attention is all you need

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.606699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.415973Z digest=sha256:ea8eca662df43494b7f86dac6f8dc9e62ceb6d5b4e5de4c42ee19762fad9d6c8

Observation 09eba7dd-b119-43f7-a49d-34a9e96886ab · outbound

This paper cites Outlier weighed layerwise sparsity (OWL): A missing secret sauce for pruning llms to high sparsity.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Outlier weighed layerwise sparsity (OWL): A missing secret sauce for pruning llms to high sparsity

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.367414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.533464Z digest=sha256:74b688991ad2c834de4248ca83d5d812ee189dd30096ff51991cea80db1f689e

Observation 10c44411-5b20-437a-9bd8-69aab4373dad · outbound

This paper cites Knowledge extraction with no observable data.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Knowledge extraction with no observable data

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.232594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.597396Z digest=sha256:4b14694503c93df11e534697ca7bc53ab05629f3721ed24ff9f7b526828b2a36

Observation b9686ceb-7242-4416-a864-7bc9aead71ef · outbound

This paper cites Hellaswag: Can a ma- chine really finish your sentence? arXiv,.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Hellaswag: Can a ma- chine really finish your sentence? arXiv,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.083604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.672976Z digest=sha256:7a184bed7fa65c8d4b9554a4168e7fdffee8f2443d0b98583b26de2fdb388d58

Observation e09bf8af-4cc2-4796-a241-0ad8ba18e944 · outbound

This paper cites Opt: Open pre-trained transformer language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Opt: Open pre-trained transformer language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.951101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.737335Z digest=sha256:a6aa8c31c830e24feb18cb8b99981f1b69913b8a170912b70f6617766621b973

Observation e680e9dd-125f-4e8c-99b4-24799b1a1e9d · outbound

This paper cites Blockpruner: Fine- grained pruning for large language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Blockpruner: Fine- grained pruning for large language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.819965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.777241Z digest=sha256:9baf47e7abf60c8c1c8f41344a88c408f05cd1eacf2050228e0ad744af052103

Observation a6c6d7c0-a273-46d9-9568-efd467e9c860 · outbound

This paper cites LLM-Pruner cannot prune Llama-2 70B and Llama-3 8B, 70B since it does not support group query attention.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information LLM-Pruner cannot prune Llama-2 70B and Llama-3 8B, 70B since it does not support group query attention

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.205996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:41.062180Z digest=sha256:c3990a55a95553c80730764d07d5c636166e72d3a41970b57c8f990cafa4d436

Observation 34f0ca9a-7111-4823-8ce8-3ce34f438b41 · outbound

This paper cites We use a single GPU to evaluate Llama-2 7B, 13B and Llama-3 8B.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information We use a single GPU to evaluate Llama-2 7B, 13B and Llama-3 8B

Reference 1024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.502576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.892994Z digest=sha256:b737ced27f5ccee08193deeb3d6c8357477c8315b32352fd19010cc236ce0538

Observation a8a4870c-8c8c-4888-b5eb-97a159a6e20b · outbound

This paper cites Qa-lora: Quantization-aware low-rank adaptation of large language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Qa-lora: Quantization-aware low-rank adaptation of large language models

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.486150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:40.478109Z digest=sha256:4d539639c734061ba7d9e675806f8d38a70463d14caa941d74e8aa5cae43684a

Observation 6208b8cd-86ff-4936-a3f7-bf9d8d13b290 · outbound

This paper cites Boolq: Exploring the surprising diffi- culty of natural yes/no questions.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Boolq: Exploring the surprising diffi- culty of natural yes/no questions

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:46.351963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:37.701919Z digest=sha256:da2f9a0816067edec41e2bef717db352d52d0df2e6526714284aa0c5ef75659a

Observation c3660369-e8ee-4ff8-ba96-58a416a8c0a0 · outbound

This paper cites The llama 3 herd of models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information The llama 3 herd of models

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:46.034194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:37.810918Z digest=sha256:a11f167712ec09437ee5eb493333bf58276144788933e719db41e4ce96b4fb8e

Observation 262e6f9e-725a-4f8c-b8cd-7eca616d768b · outbound

This paper cites Language models are few-shot learners.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Language models are few-shot learners

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:47.056001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:37.324634Z digest=sha256:b1156587473f41804ee8b6aeee109244864b789b4b26619384d17371587692ae

Observation 76a1307c-49d8-4f7d-bccd-a6c255ac0ab9 · outbound

This paper cites Flexround: Learnable rounding based on element-wise division for post-training quantiza- tion.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Flexround: Learnable rounding based on element-wise division for post-training quantiza- tion

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.456308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:38.853165Z digest=sha256:8665a662e87a0039b222aab8097520adaee7fcd02d6de33cea4d53542949ea4d

Observation 75c45ef0-2e10-403a-b356-6c4eade2493a · outbound

This paper cites Palm: Scaling language modeling with pathways.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Palm: Scaling language modeling with pathways

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:46.717315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:37.510415Z digest=sha256:5e72e62f9c6f2e8f699928b9858a3d123163246702c5ee3167c3b3d2e12be907

Observation ee1a10d3-c294-4b29-834f-e99a59aef860 · outbound

This paper cites Think you have solved question answer- ing? try arc, the ai2 reasoning challenge.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Think you have solved question answer- ing? try arc, the ai2 reasoning challenge

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:46.487145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:37.633240Z digest=sha256:637ce173319dbaaed891d2c51704cc272467cdca1a1d455454b2811802ec8834

Observation 08a4932e-0ad1-4204-ae89-03e97ba2dec3 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Piqa: Reasoning about physical commonsense in natural language

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:47.066226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:37.233404Z digest=sha256:aefc4f57602c6b7962c733e86fbd9f09bb0891673b63edca02243d44a6250ab2

Observation 83424917-ff2e-41e3-bc0e-ba0caa9b00ec · outbound

This paper cites Distillm: Towards streamlined distil- lation for large language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Distillm: Towards streamlined distil- lation for large language models

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.670504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:06:38.628367Z digest=sha256:c1e11fcabc08265be32d9882ccdb37638c09406e016f189ec674dce6d835936d

Pith citing papers

No inbound Pith citation observations are available.