Pith. sign in

Paper Citation Record · LEDGER

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information

As of 23 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.03510.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03510 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:06:41.062180Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5556fa6b-2f0c-49df-abfd-fe95bc6ab418 · outbound

This paper cites Slicegpt: Compress large language models by deleting rows and columns.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Slicegpt: Compress large language models by deleting rows and columns

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:47.075297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:37.206536Z digest=sha256:edff22c0b0276a2ef8eb8660682225ec4a2f9c37f3f57c5926e6cd9d84bb811e

Observation aa1fee7e-3982-41fb-ad3d-30f40759b92d · outbound

This paper cites • ARC-Challenge and ARC-Easy[Clark et al., 2018] are datasets composed of grade-school level multiple-choice science questions.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information • ARC-Challenge and ARC-Easy[Clark et al., 2018] are datasets composed of grade-school level multiple-choice science questions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.653879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.829414Z digest=sha256:b3b63a93560e4cc094291bb70bb1b5fc2d56c12971b806c29a054ca09be82f20

Observation e01af565-2563-4ba0-ac88-66611601c497 · outbound

This paper cites Pea-kd: Parameter-efficient and accurate knowledge distillation on bert.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Pea-kd: Parameter-efficient and accurate knowledge distillation on bert

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:46.934334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:37.414283Z digest=sha256:f03c487c20ce227b35aa997086cf7b5b202ff228df73e20aa715d1590fd162ec

Observation 38237da4-63d0-476f-9c26-bc362fbc1804 · outbound

This paper cites This demonstrates that five candidates are sufficient to find the sublayer with the least importance score.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information This demonstrates that five candidates are sufficient to find the sublayer with the least importance score

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.342946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.983994Z digest=sha256:2707e77d58e0fdbe50ef32e6bace60310cd2dde31460e7a2e7a8e47e5c484b66

Observation 662f402c-6d1d-4514-8715-854e9007b60e · outbound

This paper cites Sparsegpt: Massive language models can be accurately pruned in one-shot.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Sparsegpt: Massive language models can be accurately pruned in one-shot

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:45.856973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:37.893937Z digest=sha256:8b2939cc89e95fa3840a87e571489d64298c845125ec780714a9729a5be6ec40

Observation 7a9db158-1b3b-4329-bf3c-fedc0359d932 · outbound

This paper cites OPTQ: Accurate quantization for generative pre-trained transformers.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information OPTQ: Accurate quantization for generative pre-trained transformers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:45.688354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:37.954853Z digest=sha256:5a3ea5f6f44dc9c0c986b505ad1a906eddcb3633038dd18fe85b61ddac6ea3b3

Observation 0dd00334-b931-47a4-a8d4-f03df3027c78 · outbound

This paper cites Lazyllm: Dynamic token pruning for efficient long context llm inference.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Lazyllm: Dynamic token pruning for efficient long context llm inference

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:45.418760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:38.036607Z digest=sha256:0e0cb74ec662d135a4f7652ce8e7597d5f58f65d17e07b809e470ddb4e4686af

Observation 9c5207f5-51a4-4bd3-85b5-6cedc6725560 · outbound

This paper cites A frame- work for few-shot language model evaluation, 12.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information A frame- work for few-shot language model evaluation, 12

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:45.239781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:38.140621Z digest=sha256:490bdefc67c2a6715bb5b32e85f66126d8c8288982181e70396468f7e8c8d137

Observation 22e84dfc-d61e-4113-b2c8-148e3b09b1f1 · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:45.068132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:38.218182Z digest=sha256:990eefc8ce24224b9750c4b33883f0cfdeb95f980701eaad7f39524b509c6eac

Observation b75d018c-4b48-4a64-b762-152809afba74 · outbound

This paper cites Falcon: lightweight and accurate convo- lution based on depthwise separable convolution.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Falcon: lightweight and accurate convo- lution based on depthwise separable convolution

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.989062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:38.333525Z digest=sha256:8311663991393ff042736959f898a5fdfde9fe96b8d1bb902c8aa5333c5604f5

Observation 90598526-0a8a-4c41-b45d-74f3b64d6507 · outbound

This paper cites an unresolved cited work.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:06:44.876957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:38.442187Z digest=sha256:ef10f7817f2d1b34d270a4af973e64fd02f2e2ff4c62a99d753f77717e0cf9f8

Observation 61086d0b-0683-4ab0-8507-943d620cd40c · outbound

This paper cites an unresolved cited work.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:06:44.783332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:38.550725Z digest=sha256:96233627ad571b0a5a424e75f70e2262a984e9feba8c2f2f6722dfc4047d96b0

Observation dfd76e9a-163d-41eb-9838-4cd5703574b0 · outbound

This paper cites an unresolved cited work.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:06:44.580464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:38.743850Z digest=sha256:f5a75a809844f82a9913281e87df67f1c2ac944dc18b3ebcadef65fd97741806

Observation 5235280a-ba91-4152-8169-b2c466d0e57f · outbound

This paper cites DDK: distilling domain knowledge for efficient large language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information DDK: distilling domain knowledge for efficient large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.347291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:38.951134Z digest=sha256:8938e9fcdbc4b73adf6fd69cc69f47d291361b410b9f04d6980d7b8b474d0f59

Observation 5634f1d0-04b7-40a3-bf6b-b7e174af23a0 · outbound

This paper cites Llm-pruner: On the structural pruning of large lan- guage models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Llm-pruner: On the structural pruning of large lan- guage models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.243254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:39.074936Z digest=sha256:22575ea197e2538981c18b5b87b6f06e8666bdd7a0ad9d88999b7986346dd5b6

Observation bbaf37f5-70f7-46e7-9ad7-22660095dc21 · outbound

This paper cites Shortgpt: Layers in large language models are more redundant than you expect.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Shortgpt: Layers in large language models are more redundant than you expect

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.109191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:39.230392Z digest=sha256:b5e097ec04066abb2c0f8e9b5eec5f57564a88f7c2078edb4885d9faaf02bcbe

Observation 12d06413-1a80-4661-852a-7e0e5163877e · outbound

This paper cites Pointer sentinel mixture models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Pointer sentinel mixture models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.997597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:39.380345Z digest=sha256:da36a0ef565e767955c5bcc7e06eedc949ed9f95b4f3a6e1e86219ca6790aca7

Observation 0f0aae67-aac6-4589-b922-b40e8174a583 · outbound

This paper cites Sensimix: Sensitivity-aware 8-bit index & 1-bit value mixed precision quantization for bert compression.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Sensimix: Sensitivity-aware 8-bit index & 1-bit value mixed precision quantization for bert compression

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.866307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:39.571943Z digest=sha256:7646a3078495c5db5350628a81de90212a4484b9f8381c80c65f972fc5a785e0

Observation 69e25463-b4f2-48c4-a910-89cb9613309a · outbound

This paper cites Mixture-of-depths: Dynamically al- locating compute in transformer-based language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Mixture-of-depths: Dynamically al- locating compute in transformer-based language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.775359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:39.666897Z digest=sha256:51014dd5722bfb57137dfe9a82a21e95475b4e8367a8353dab8d440b3ba45578

Observation 2e88c0cc-9fc4-499e-8b3a-59aee2d02567 · outbound

This paper cites Exaone 3.0 7.8 b instruction tuned language model.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Exaone 3.0 7.8 b instruction tuned language model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.642116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:39.844672Z digest=sha256:46ccb5cada6d0219fc2a327699cb83b6c68e97cdbc96c807e7206ca7761a8e51

Observation 8fa26fd2-3181-48d5-b2bb-ec9cf99d3a37 · outbound

This paper cites Confident adaptive language modeling.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Confident adaptive language modeling

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.540375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.001779Z digest=sha256:685cfcbba7980629351cf356eb9cf003aa8a50eb508fc2149bd6472a62797c95

Observation 010dc779-83e3-4e6f-acb5-d19675455bbf · outbound

This paper cites Omniquant: Omnidi- rectionally calibrated quantization for large language mod- els.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Omniquant: Omnidi- rectionally calibrated quantization for large language mod- els

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.373199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.117420Z digest=sha256:7843c7f27826a5ec5297e32ec671449f66e170037996ead9411ec1abd7996756

Observation 4b1d685d-2559-4158-89a4-b03b97987496 · outbound

This paper cites Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Sleb: Streamlining llms through redundancy verification and elimination of transformer blocks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.227336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.178601Z digest=sha256:bb502695347b321fc7ba6bd1eae280150df868bc19f96d44182b3aa67b05f58b

Observation 2140cf69-46e2-4cc5-b1dc-37371f098a7d · outbound

This paper cites A simple and effective pruning approach for large language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information A simple and effective pruning approach for large language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:43.038919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.242222Z digest=sha256:461961d20c193c79fb074377d5825c28d10a0ac44a8ca8b0b7b0e4113fa9dec7

Observation 97329a70-27f7-43d8-870c-50f274ae52ed · outbound

This paper cites Gemini: a family of highly capable multimodal models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Gemini: a family of highly capable multimodal models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.890677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.313820Z digest=sha256:dcee9f5f23c932e3eed33cb51fdc517c1dd410584776052f493a0e7e33154142

Observation 472f7a99-b653-4b6b-bb98-fe8f39d00ee4 · outbound

This paper cites Accelerating llama infer- ence by enabling intermediate layer decoding via instruc- tion tuning with lite.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Accelerating llama infer- ence by enabling intermediate layer decoding via instruc- tion tuning with lite

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.736879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.358457Z digest=sha256:54312fd660c65baa53d6dfb538428b58536aa68e29bf7459a5a90b1bbb8b9990

Observation b6ec1720-b2c9-45df-8ca1-06f156af4020 · outbound

This paper cites Attention is all you need.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Attention is all you need

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.606699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.415973Z digest=sha256:22cabfd06019278616967b17298c3c9f02157a65b0dab3d53cbefc8d531e9041

Observation 09eba7dd-b119-43f7-a49d-34a9e96886ab · outbound

This paper cites Outlier weighed layerwise sparsity (OWL): A missing secret sauce for pruning llms to high sparsity.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Outlier weighed layerwise sparsity (OWL): A missing secret sauce for pruning llms to high sparsity

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.367414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.533464Z digest=sha256:fc7498222a5f0cb324d29aa20bca7afba44d7344525d840acb8edda3994d3be0

Observation 10c44411-5b20-437a-9bd8-69aab4373dad · outbound

This paper cites Knowledge extraction with no observable data.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Knowledge extraction with no observable data

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.232594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.597396Z digest=sha256:80ccbed09238674020f7d7bedd116ede4b38ccf12ce22d5e064dfd250a61f3a4

Observation b9686ceb-7242-4416-a864-7bc9aead71ef · outbound

This paper cites Hellaswag: Can a ma- chine really finish your sentence? arXiv,.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Hellaswag: Can a ma- chine really finish your sentence? arXiv,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.083604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.672976Z digest=sha256:cdefb8c1e30a6f27ee2679d89c28f67e909f3d0a1021b3ccf838453fccd627de

Observation e09bf8af-4cc2-4796-a241-0ad8ba18e944 · outbound

This paper cites Opt: Open pre-trained transformer language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Opt: Open pre-trained transformer language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.951101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.737335Z digest=sha256:a2edeed3007f92cc20af1e1b01a6c895093bd965d188f765ebc23f1efc8cc2da

Observation e680e9dd-125f-4e8c-99b4-24799b1a1e9d · outbound

This paper cites Blockpruner: Fine- grained pruning for large language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Blockpruner: Fine- grained pruning for large language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.819965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.777241Z digest=sha256:42e6302d6a57eae70bfbd501ead42eb530c92f7b1f77d13c2a4687a862ff9253

Observation a6c6d7c0-a273-46d9-9568-efd467e9c860 · outbound

This paper cites LLM-Pruner cannot prune Llama-2 70B and Llama-3 8B, 70B since it does not support group query attention.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information LLM-Pruner cannot prune Llama-2 70B and Llama-3 8B, 70B since it does not support group query attention

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.205996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:41.062180Z digest=sha256:c3d0f862465dfa3e30f6a9fa81efa6e90133e97051da9d3618ae90d62b24147e

Observation 34f0ca9a-7111-4823-8ce8-3ce34f438b41 · outbound

This paper cites We use a single GPU to evaluate Llama-2 7B, 13B and Llama-3 8B.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information We use a single GPU to evaluate Llama-2 7B, 13B and Llama-3 8B

Reference 1024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:41.502576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.892994Z digest=sha256:be012b09f2b0a52b31d3a4ac09aaabeee36da3d12e2d472e96754ddf5c64ead9

Observation a8a4870c-8c8c-4888-b5eb-97a159a6e20b · outbound

This paper cites Qa-lora: Quantization-aware low-rank adaptation of large language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Qa-lora: Quantization-aware low-rank adaptation of large language models

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:42.486150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:40.478109Z digest=sha256:c180e723ea2674ea809ffb7f9977a0b4cb94dc2aafbddcca47c20cc22896b6a1

Observation 6208b8cd-86ff-4936-a3f7-bf9d8d13b290 · outbound

This paper cites Boolq: Exploring the surprising diffi- culty of natural yes/no questions.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Boolq: Exploring the surprising diffi- culty of natural yes/no questions

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:46.351963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:37.701919Z digest=sha256:2e004d25ad4822cba7dddac9df7d9dff6fcf1571826935baa4d5175e1eedf242

Observation c3660369-e8ee-4ff8-ba96-58a416a8c0a0 · outbound

This paper cites The llama 3 herd of models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information The llama 3 herd of models

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:46.034194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:37.810918Z digest=sha256:377817649eb40986894f33e49c93906aa430c1a3fa5cd1bcd21c57d84389a4b9

Observation 262e6f9e-725a-4f8c-b8cd-7eca616d768b · outbound

This paper cites Language models are few-shot learners.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Language models are few-shot learners

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:47.056001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:37.324634Z digest=sha256:8a159ee0e3d4fc33eae526bcb30c5da26d998a40909c40e4258bb12ce5b2b36e

Observation 76a1307c-49d8-4f7d-bccd-a6c255ac0ab9 · outbound

This paper cites Flexround: Learnable rounding based on element-wise division for post-training quantiza- tion.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Flexround: Learnable rounding based on element-wise division for post-training quantiza- tion

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.456308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:38.853165Z digest=sha256:f61e542f063b6210082d1f9dd55e959aa0669f8b784b3ce6d7a267ea6580d940

Observation 75c45ef0-2e10-403a-b356-6c4eade2493a · outbound

This paper cites Palm: Scaling language modeling with pathways.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Palm: Scaling language modeling with pathways

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:46.717315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:37.510415Z digest=sha256:03605454e44131f58be24d3d8f70c17268f1e983845e2f6158325c43db4f4d32

Observation ee1a10d3-c294-4b29-834f-e99a59aef860 · outbound

This paper cites Think you have solved question answer- ing? try arc, the ai2 reasoning challenge.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Think you have solved question answer- ing? try arc, the ai2 reasoning challenge

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:46.487145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:37.633240Z digest=sha256:b2cf0c9f9c635e204cfe83b78cfbaebf4034f41e824518c1b1ee8e1d7dde9c9b

Observation 08a4932e-0ad1-4204-ae89-03e97ba2dec3 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Piqa: Reasoning about physical commonsense in natural language

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:47.066226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:37.233404Z digest=sha256:d4dd24b310487cda8ae7a3e7768aae00411ead381fff85a468208fd50e049df1

Observation 83424917-ff2e-41e3-bc0e-ba0caa9b00ec · outbound

This paper cites Distillm: Towards streamlined distil- lation for large language models.

Accurate Sublayer Pruning for Large Language Models by Exploiting Latency and Tunability Information Distillm: Towards streamlined distil- lation for large language models

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:06:44.670504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T11:06:38.628367Z digest=sha256:51a0ba5dd68dd1cf7a189d8699dac6ef32cc6f98ffa14127b446308991959d35

Pith citing papers

No inbound Pith citation observations are available.