Pith. sign in

Paper Citation Record · LEDGER

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices

As of 16 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 0 inbound Pith citation observations for arXiv:2411.17720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17720 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:20:03.210254Z

measured 80 of 80 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact2
  • verified fuzzy39
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08ddaea7-4bee-4635-8789-45e8964ad17f · outbound

This paper cites https://docs.nvidia.com/deeplearning/tensorrt/archives/tensorrt-803/best-practices/index.html.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices https://docs.nvidia.com/deeplearning/tensorrt/archives/tensorrt-803/best-practices/index.html

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.471320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.783745Z digest=sha256:b5bb467ce2f3efabfaa3595a640687006799f13e4cfae6a4455dd20b121f75f3

Observation a4970715-898a-4066-a788-ea774ac0a9a8 · outbound

This paper cites https://developer.apple.com/documentation/accelerate/bnns, a.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices https://developer.apple.com/documentation/accelerate/bnns, a

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.448648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.789786Z digest=sha256:1a1570686cdc8b5503c93403b23fb9e9d3234d4ed212c0dad1f0ce447d9512f7

Observation 3243a98a-67db-467f-a5ee-58fedbce9656 · outbound

This paper cites https://apple.github.io/coremltools/docs-guides/source/opt-palettization-overview.html, b.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices https://apple.github.io/coremltools/docs-guides/source/opt-palettization-overview.html, b

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.427703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.795001Z digest=sha256:e4999c8d15fc83a0d867eb4abb7146d6f3ec78932331668360804d7b60341435

Observation bdab1616-2a0f-4fa2-82e6-baf3a1b60459 · outbound

This paper cites https://developer.apple.com/documentation/metalperformanceshadersgraph, c.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices https://developer.apple.com/documentation/metalperformanceshadersgraph, c

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.409736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.800450Z digest=sha256:adff3abc72c02cd973fe9295f8b39b478124755e37532861f763acd348f8d170

Observation 15003672-7012-4a93-acc2-1847b4cba90a · outbound

This paper cites an unresolved cited work.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:15.391770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.805879Z digest=sha256:49ba25f17a6efc671e38a75f6dc142c6ee44e1afb6620b0d3f92f86b13376cc3

Observation d5088739-c155-403a-b9ec-bc3906eb6fd3 · outbound

This paper cites https://www.tensorflow.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices https://www.tensorflow

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.373257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.811069Z digest=sha256:9cc15d482bd63b64afcd8b44a60b746528420138264ab441d2dfc6a23a9f98c8

Observation cb93ec47-6c47-45ec-ba94-95a0c8ff885d · outbound

This paper cites Y., Rajbhandari, S., Awan, A.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Y., Rajbhandari, S., Awan, A

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.817329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.817329Z digest=sha256:d56f85494c9fe0a1614d1eb71541d92e070f545eaed70131f25d249ac66f0e01

Observation 57db820d-51e3-4f8d-8de9-c98fe27b3422 · outbound

This paper cites B., Del Sozzo, E., Akkas, A., Zhang, Y., Suriana, P., Kamil, S., and Amarasinghe, S.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices B., Del Sozzo, E., Akkas, A., Zhang, Y., Suriana, P., Kamil, S., and Amarasinghe, S

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.336420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.822792Z digest=sha256:e2b34839ea1e3a22221d9c0486dc2acb688ece540afc8060837593857bf966a8

Observation dc2bee9f-c6ff-4cb5-9eed-eed8944db39f · outbound

This paper cites \ TVM \ : An automated \ End-to-End \ optimizing compiler for deep learning.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices \ TVM \ : An automated \ End-to-End \ optimizing compiler for deep learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.309182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.830076Z digest=sha256:46bef22721db0d65c8e4dd6b909980169275c8f9986a52b88847dd5ac24ee331

Observation 43f3e696-0448-442b-8e1a-b2c1b6a68453 · outbound

This paper cites Learning to optimize tensor programs.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Learning to optimize tensor programs

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.280430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.836041Z digest=sha256:0572efd6a0386c30edfeeb885ee091e3a5d6394ec6f0566a56fa270519a20219

Observation 55476913-9260-4e8d-8641-16e3905322c1 · outbound

This paper cites Eyeriss v2: A flexible accelerator for emerging deep neural networks on mobile devices.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Eyeriss v2: A flexible accelerator for emerging deep neural networks on mobile devices

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.259117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.842077Z digest=sha256:ca96795554f8d761473945697ed2444cd9e5d1841db20fea0cafe5bacd899d31

Observation 211ed60a-b480-419d-942b-dcae66d7d2f5 · outbound

This paper cites DKM: Differentiable K-Means Clustering Layer for Neural Network Compression.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices DKM: Differentiable K-Means Clustering Layer for Neural Network Compression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.848229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.848229Z digest=sha256:849e30deb7aeb4fa5fa578b2c1838614be6489cbcfb1383a5c5f28f9bc88500a

Observation b1cd0939-2d2c-4056-9593-6f0adb04c06c · outbound

This paper cites KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices KV-Runahead: Scalable Causal LLM Inference by Parallel Key-Value Cache Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:20:14.369076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.854975Z digest=sha256:9498287a0606a6ef113a99b112d1d67da391cb120bf7ff50a220bd062e56653f

Observation f7170fa5-29ef-4020-b071-a53080d6c696 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.862330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.862330Z digest=sha256:7f6ac5d12eb81d896741e3b129bedca606b5e76843751f6b5efd08649aeabf8f

Observation 657bbd5a-c650-4f11-a93a-19b054cab6f9 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.867992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.867992Z digest=sha256:6cf3e1fda9f26269719a755868ec36aebe98b4cdede6a18097f23b1acb4dca22

Observation d7eac942-454e-4b2b-ab0f-2ea4d77b8de9 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.873180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.873180Z digest=sha256:942de48e2586699650805083f96f0f56bb9234ea60564cb09b02384bed098116

Observation c218a399-5fdd-41d8-ba87-b2981d6d77f0 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.878218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.878218Z digest=sha256:7170f282d019526faf4c2fea84fbf43c2ec6560b0b90942d2dc6f7d38b948dcf

Observation 1085b971-9779-4f53-872a-6ea70244359e · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Scaling rectified flow transformers for high-resolution image synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.883210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.883210Z digest=sha256:d96581b97fead2fbb7da0270557ad6274543500c11cd0fde92ef476732333e00

Observation 4cb0e3f2-cc72-49a5-8545-a3c623e882ff · outbound

This paper cites Videoagent: A memory-augmented multimodal agent for video understanding.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Videoagent: A memory-augmented multimodal agent for video understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.209274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.889268Z digest=sha256:60ad203620ce91450133dcdbf3e8604d6dbfc523bd6a3fdf9c09df888501c59c

Observation 422ae86d-9f4f-4d19-86d5-a3a043454b43 · outbound

This paper cites A., Yang, Y., Sajjad, H., Nakov, P., Chen, D., and Winslett, M.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices A., Yang, Y., Sajjad, H., Nakov, P., Chen, D., and Winslett, M

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.181293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.895539Z digest=sha256:c5c82ae8b2ec3bec3aa357cf350c0cdda58d487ca2e39278f50cec0212e9da6f

Observation f1c2b330-65ae-48bb-9f1f-8e8f0f4838ab · outbound

This paper cites Collective loop fusion for array contraction.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Collective loop fusion for array contraction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.155247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.900740Z digest=sha256:2175ca4185839a774807375d4cf683e602e4236e2b6f4f5e35c34803f62a9939

Observation 74701642-78c7-4236-8b3d-53d277a454e8 · outbound

This paper cites Improving alignment of dialogue agents via targeted human judgements.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Improving alignment of dialogue agents via targeted human judgements

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.906203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.906203Z digest=sha256:fe35191708989edee26034f33d1cc9d74790d0ec194d5d64f0aba23cf4e67bc8

Observation 2e33714f-b51c-4fb2-993d-70bc49c617ef · outbound

This paper cites Progressive Knowledge Distillation Of Stable Diffusion XL Using Layer Level Loss.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Progressive Knowledge Distillation Of Stable Diffusion XL Using Layer Level Loss

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.911109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.911109Z digest=sha256:e6fe3126951e6b5744d04098cdb22e8f9679911ca7126afbfe3065746021466c

Observation 615c5762-3583-416c-8876-20a484ea1b5b · outbound

This paper cites FlashDecoding++: Faster Large Language Model Inference on GPUs.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.915968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.915968Z digest=sha256:26db05df4336028fa0592d07b0a1061e222318526365370eae74bccbc11ea7cf

Observation 25e62ec3-a073-4629-a8fc-b6b2faebed7c · outbound

This paper cites Knowledge diffusion for distillation.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Knowledge diffusion for distillation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.129432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.920773Z digest=sha256:aba3fac9a2e5cea1b2fd684c4f6b769b6516aada56e8c72d743a8edcbec3c236

Observation e59812d8-4cad-4ae5-b3e9-b6405a920dc5 · outbound

This paper cites Data movement is all you need: A case study on optimizing transformers.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Data movement is all you need: A case study on optimizing transformers

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.110169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.925474Z digest=sha256:5a4f05750b681f116a8034172cb7b251b0bd826f7405543e85a0d6cc53f8933c

Observation 6a5d504d-efc5-4178-800a-4a1300b2c002 · outbound

This paper cites P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices P., Young, C., Patil, N., Patterson, D., Agrawal, G., Bajwa, R., Bates, S., Bhatia, S., Boden, N., Borchers, A., et al

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.089089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.932705Z digest=sha256:cd12c8cd61169590fdd0b91cdc89506aedf533f0d35ff2600314cc4461fa257c

Observation 3ad6a27f-c2a2-4f61-b25c-f5ba091fd81c · outbound

This paper cites P., Yoon, D.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices P., Yoon, D

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.937725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.937725Z digest=sha256:cb352ed40804ee8803a2ba2068bac832bb28e82beb7501e430a88b61dd04e6ed

Observation aa871e7b-7345-4c69-93df-ad5768587f41 · outbound

This paper cites Flat: An optimized dataflow for mitigating attention bottlenecks.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Flat: An optimized dataflow for mitigating attention bottlenecks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.049068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.942641Z digest=sha256:f6664831df6c0d4173f0329edc745714ee936f48fc4d9349e48c383b22af4b08

Observation 070363a4-5661-4352-a63a-91533cb51988 · outbound

This paper cites Scaling Laws for Neural Language Models.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Scaling Laws for Neural Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.947873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.947873Z digest=sha256:252c2e55a4358be2b0ccb51948f34a0edc69d048f32f781effa8004c3fcba9ab

Observation 97e04be2-41ff-4b7a-a040-a4e2640bdcc7 · outbound

This paper cites an unresolved cited work.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:20:15.028721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.956659Z digest=sha256:7e11c60386ea303032a0a839e834dbf0095512ae410d424fe96489edf6433272

Observation 569d0a15-f1a2-4c9f-ba9c-3b0b9a387fee · outbound

This paper cites Reformer: The Efficient Transformer.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Reformer: The Efficient Transformer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.962077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.962077Z digest=sha256:2c9db69c441f923ae93e24be866e6e378dc83cca5377581f62c988586c1b9acb

Observation 13f8cc53-bd8d-4be6-9fcb-8354ca092ca8 · outbound

This paper cites The tensor algebra compiler.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices The tensor algebra compiler

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:15.000367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.967349Z digest=sha256:d84476d854f8cb75b1e22f0cb7e15e8122ae7579882f84a65c6e072c85f3cd0f

Observation aaa3a126-5fc9-44ae-92f2-b558261aea1e · outbound

This paper cites Maeri: Enabling flexible dataflow mapping over dnn accelerators via reconfigurable interconnects.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Maeri: Enabling flexible dataflow mapping over dnn accelerators via reconfigurable interconnects

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.983792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.972663Z digest=sha256:593caf6e7d80e9f8b3cfba36efb20be2d07f5fc408fed787888c621b3547c4b8

Observation bd67bee7-737f-4788-8072-a4de1a0eb552 · outbound

This paper cites H., Gonzalez, J., Zhang, H., and Stoica, I.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices H., Gonzalez, J., Zhang, H., and Stoica, I

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.978310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.978310Z digest=sha256:898ec4ca3265774b25618bac99b96668ddc06d84b1d67fbd068d9b9a0d921b85

Observation 39fc4848-91d3-4769-b67d-ebc210374f78 · outbound

This paper cites Cross-lingual Language Model Pretraining.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Cross-lingual Language Model Pretraining

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:02.983558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:02.983558Z digest=sha256:39680fa2509d41c4b7ebef7f076343b87d50cba29907b6b9a6c7dc420abbadbd

Observation 4a51a011-c030-4bb1-aa98-75b0c8640d66 · outbound

This paper cites Multimodal foundation models: From specialists to general-purpose assistants.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Multimodal foundation models: From specialists to general-purpose assistants

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.952238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.988978Z digest=sha256:cf3e81204e9703128b780e952cbfec9a09085f22eddd6dd4d8b2a42909bf78b5

Observation 8ca1f9a9-53a8-476b-97d4-9d718451e61b · outbound

This paper cites onednn graph compiler: A hybrid approach for high-performance deep learning compilation.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices onednn graph compiler: A hybrid approach for high-performance deep learning compilation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.934454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:02.994280Z digest=sha256:70115836bc57b8b51d59fa7282c9a4a747d2f1896dc5ee8b1b4a07f482511570

Observation 6f007858-1bc0-4d1f-85be-eda741ddb22f · outbound

This paper cites Q-vit: Accurate and fully quantized low-bit vision transformer.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Q-vit: Accurate and fully quantized low-bit vision transformer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.916406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.000080Z digest=sha256:dff4b53cd1375c16c0d1e1ed2276268d0fc0be8df92d9fe24158a0654c9b16ca

Observation 8b3bfa95-725f-4604-af59-7f418b843ada · outbound

This paper cites and Gu, Q.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices and Gu, Q

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.896384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.005303Z digest=sha256:ddea199bc6cf201c03ebfc66b717e7455351341a4b6bfcf42ce3159ca26083f7

Observation 387af77b-2df8-4b2d-accf-3a4946d51a77 · outbound

This paper cites Davinci: A scalable architecture for neural network computing.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Davinci: A scalable architecture for neural network computing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.876522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.010957Z digest=sha256:c85cc7b0be53b60964e4b5f81fc3b2d94bc328ce5e71ef8b60c054dacfedc878

Observation 625db241-1661-40fc-a0c5-98b2bd369dfa · outbound

This paper cites FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices FQ-ViT: Post-Training Quantization for Fully Quantized Vision Transformer

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.015840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.015840Z digest=sha256:287d70ecb9ae5238fbae0aabffaa810c87276424030a886df2adca51c7229cf5

Observation aa7dbaca-1be1-4882-aeae-022dee9fcd61 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.021686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.021686Z digest=sha256:1a20f15aa30fc6bd6b75767d49a4565b9534157bf48e75a65687361d5225e56e

Observation 88ad4273-175f-4742-a4ff-5fee35e01732 · outbound

This paper cites Post-training quantization for vision transformer.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Post-training quantization for vision transformer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.858186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.027499Z digest=sha256:cc861dcca68fc4a4ecdbb80a93dc40f6ce7cb001a2615b49e9022d33f24d341e

Observation 0ed8ecc8-2895-4b75-8eb6-691ee52ec6bd · outbound

This paper cites Tprune: Efficient transformer pruning for mobile devices.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Tprune: Efficient transformer pruning for mobile devices

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.838320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.032104Z digest=sha256:c046dbd72ff4ab347c45a09a53c38f20fa24aa43c6eeff0b58117f025a98d649

Observation f8ee8da5-b38d-4077-8b0b-383e1b4b42de · outbound

This paper cites OpenELM: An Efficient Language Model Family with Open Training and Inference Framework.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices OpenELM: An Efficient Language Model Family with Open Training and Inference Framework

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.037584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.037584Z digest=sha256:afe4b46f4f838d51e1818c3040ad46200bb00d6366f579adce610494e3e5ef98

Observation 4776660d-adb1-4ad3-975a-da4b14a28e88 · outbound

This paper cites Defines: Enabling fast exploration of the depth-first scheduling space for dnn accelerators through analytical modeling.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Defines: Enabling fast exploration of the depth-first scheduling space for dnn accelerators through analytical modeling

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.818518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.042845Z digest=sha256:f88c1e722f64a9e4a982e6e8218d38ed7234ed5c09912b1d55d61d181cf8640b

Observation 6f9af71f-4511-4291-833b-d5d342481dd3 · outbound

This paper cites Efficient large-scale language model training on gpu clusters using megatron-lm.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Efficient large-scale language model training on gpu clusters using megatron-lm

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.047677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.047677Z digest=sha256:54b1d393bd19d8ab3c1e56964a211fedb4aab0c668a994cf24aeea0e256ea409

Observation 9e5f5e38-9f5e-4c17-91e7-deeb4abb448b · outbound

This paper cites O., Pellauer, M., Emer, J.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices O., Pellauer, M., Emer, J

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.053659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.053659Z digest=sha256:f58494d6c98ac798b666efa2a125d35808bc9fc691e1bedd7e86c93c444b9024

Observation 80725d92-7ce5-45de-855b-9582a78e3831 · outbound

This paper cites Dnnfusion: accelerating deep neural networks execution with advanced operator fusion.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Dnnfusion: accelerating deep neural networks execution with advanced operator fusion

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.782306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.058778Z digest=sha256:4c8fb5098590360a9042ac87b2bdadcf9cdb1e845f19336cdd4f8444f8d9dae5

Observation 5bfb9de1-0142-4b59-b842-07954cde34b8 · outbound

This paper cites Training language models to follow instructions with human feedback.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Training language models to follow instructions with human feedback

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.064414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.064414Z digest=sha256:5fda5959abb5c64591b795e524254907c2974965faa1d4f43b3c24fe489f83b3

Observation 18194aa1-6223-4ac5-b4ad-e169f3f075fd · outbound

This paper cites S., Chen, Y.-H., Ying, V.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices S., Chen, Y.-H., Ying, V

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.738941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.069821Z digest=sha256:9d635175e9fcd213ed87a480dbe8961752a608ca692fdea2becc29de69bfad15

Observation 87b29825-d987-4b8c-bc11-51cff97aba81 · outbound

This paper cites Splitwise: Efficient generative llm inference using phase splitting.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Splitwise: Efficient generative llm inference using phase splitting

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.713331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.075016Z digest=sha256:5547a4db395410e2de8058cb25a7576cc15fec23907d79ab62fbda5ccd8eca54

Observation add3a1c1-9f93-419c-b611-856cb36133aa · outbound

This paper cites and Xie, S.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices and Xie, S

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.080119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.080119Z digest=sha256:65f821fe8df2fdd46079c33109b067bc8d7cd6ca072dbbaaff76022846cdceaa

Observation 6c14337a-f0f9-4d91-968c-5edbfc34738b · outbound

This paper cites Accelerating transformer-based deep learning models on fpgas using column balanced block pruning.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Accelerating transformer-based deep learning models on fpgas using column balanced block pruning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.679219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.085645Z digest=sha256:9bcdaaa415d526195b0158364dfb0257881f92d127e1a467f219d577e8abd9ed

Observation abd6822c-fb4f-494f-ba47-c296261c1175 · outbound

This paper cites Sensimix: Sensitivity-aware 8-bit index & 1-bit value mixed precision quantization for bert compression.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Sensimix: Sensitivity-aware 8-bit index & 1-bit value mixed precision quantization for bert compression

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.659045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.090817Z digest=sha256:1e4e93aedffebd97084abbb007199c67b37700c8d308901acf958f0bf7b8f799

Observation 2872df2a-25f8-4999-9c0d-a0d25d4d3bc3 · outbound

This paper cites DreamFusion: Text-to-3D using 2D Diffusion.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices DreamFusion: Text-to-3D using 2D Diffusion

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.095869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.095869Z digest=sha256:447bbd3ecda71db1a3bf907ec046c705c0c09126199196d76939ed9e8226fb80

Observation c136a145-3aa9-40d2-895f-ea5a9980a11a · outbound

This paper cites Improving language understanding by generative pre-training.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Improving language understanding by generative pre-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.101666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.101666Z digest=sha256:3caa138d923b0abd849dad000879101b03299d496f7c07d00681c8f0f15a1f28

Observation 5f0c9c46-f664-4789-b55c-77084f7ee513 · outbound

This paper cites an unresolved cited work.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.106594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.106594Z digest=sha256:bd6c28e9fc16281a4be794cbfa4addccd8c447dd5140951f15110ce3e2296f7f

Observation 81203344-da31-447d-a4e6-1bf7e9f1b1e9 · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.111284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.111284Z digest=sha256:4f1d8c25ae075db269847741fd50eb4f453880c909284efe961a33193a92c6ee

Observation 914bfc6f-707e-4d7c-9b13-2b7c03fa503a · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.115971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.115971Z digest=sha256:6a4bb528579430980811112058d5e7c3efee9438495bdae6a4df99c16173c65c

Observation e132c2de-bb0a-4801-8867-756b16a441cf · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.120751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.120751Z digest=sha256:d4dc3ea5f7f6f6c2024e0eea6442bde8523fefdabc5ef146d0db3c8fa98fe939

Observation 193681a3-7ed0-4996-90b7-2541e155a979 · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Patient Knowledge Distillation for BERT Model Compression

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.125648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.125648Z digest=sha256:42b965eff90f18b728ee80643b484f7b37eaf4e6748c084ca7995deeb871b558

Observation 664c87d5-e70c-47b7-a61e-c85db9a2556a · outbound

This paper cites Improving the efficiency of transformers for resource-constrained devices.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Improving the efficiency of transformers for resource-constrained devices

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.593717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.130572Z digest=sha256:1d42d2f7631132782824ab943d0a7394ed20131f61852463f589e5f292333822

Observation e2cb3bf3-da56-4378-aebe-8a4bd7979bea · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices LLaMA: Open and Efficient Foundation Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.135423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.135423Z digest=sha256:af2d7c14f44210ced9d0fc7511b3f1bf43b775d17b2f28bd1d4023548833df2f

Observation e10fc1db-b711-458e-92cd-fb706b1e38a6 · outbound

This paper cites N., Kaiser, ., and Polosukhin, I.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices N., Kaiser, ., and Polosukhin, I

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.140419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.140419Z digest=sha256:c31a1baac4acf5f1b3d1647ea7c29cdc33bf77a17d8443d4e29f3a7d42522c20

Observation 306620eb-8751-4452-8a76-eadb6a7025a0 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.145702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.145702Z digest=sha256:ed86f2cea12c19ba9ef7dc9e09fdbbe250fc873b1f25237ec1a15e4ba590e84f

Observation 4417ca96-4b92-4ad2-8bed-f169a120af87 · outbound

This paper cites C., Venkataramani, S., Sen, S., Chen, C.-Y., El Maghraoui, K., Srinivasan, V.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices C., Venkataramani, S., Sen, S., Chen, C.-Y., El Maghraoui, K., Srinivasan, V

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.555143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.150817Z digest=sha256:3c20649d511df7c6e252864ecdfee6b5012cbb69926f95d1f23483dfca3858f4

Observation a0a64d52-cdb4-4bfd-a69e-d4999e7ed5dd · outbound

This paper cites Cluster-Former: Clustering-based Sparse Transformer for Long-Range Dependency Encoding.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Cluster-Former: Clustering-based Sparse Transformer for Long-Range Dependency Encoding

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:20:03.324238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.156474Z digest=sha256:123dff41466e4eeb84e3f6d08dc2cee9ec000561bcdd97b7865ed0771729a676

Observation c5423b7d-d1a1-44ea-aa1f-3ca1828e13ba · outbound

This paper cites MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices MiniLMv2: Multi-Head Self-Attention Relation Distillation for Compressing Pretrained Transformers

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.161573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.161573Z digest=sha256:fe896689228eec50364a025d9074f9f54dac5a536e5d9b7a397c00e701ca6b2a

Observation f84c3fab-eeba-4ecc-ab12-502be72f0fb8 · outbound

This paper cites Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Minilm: Deep self-attention distillation for task-agnostic compression of pre-trained transformers

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.535907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.166375Z digest=sha256:e4bcdcef3fbc74f0f864d5682eb26643fd403b0d4597522044df2e01bc267ad4

Observation 08a15981-5a58-4b34-82d3-c8c6f209b796 · outbound

This paper cites N., Emer, J.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices N., Emer, J

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.515715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.171117Z digest=sha256:f00659da279fa45933c6e84db82b4089a2c36094d4a6213f3f366d22514d52bc

Observation d62d813a-88ba-48c8-b259-283f78199d63 · outbound

This paper cites Zeroquant: Efficient and affordable post-training quantization for large-scale transformers.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Zeroquant: Efficient and affordable post-training quantization for large-scale transformers

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.175603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.175603Z digest=sha256:41826a114675be21699fe4b42b0218c956ecd37326b445adf44bb9063cea8716

Observation 7925b971-d851-4c7e-9663-1c097be4b72d · outbound

This paper cites Boost vision transformer with gpu-friendly sparsity and quantization.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Boost vision transformer with gpu-friendly sparsity and quantization

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.485611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.180775Z digest=sha256:b6239ce615314a47e2c9a9ea1ddf8fdfe5195051d4acca467b983955030a30d0

Observation 88d768d8-3359-4ad1-b447-c23fdc91e314 · outbound

This paper cites Width & depth pruning for vision transformers.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Width & depth pruning for vision transformers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.464747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.185884Z digest=sha256:75d662ad67d47ddfdf64de99b9794840661486b81c4e7836759ed425ada77a8d

Observation 3f674856-9f03-4622-88b2-34a8b7de35e4 · outbound

This paper cites Unified Visual Transformer Compression.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Unified Visual Transformer Compression

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.190444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.190444Z digest=sha256:6e8065f2039ac85f9a2da852b561d780359daa363d1b64e1144defb8d795ff4f

Observation 3653e378-8e32-4bd2-9629-a6609408d2bc · outbound

This paper cites AppAgent: Multimodal Agents as Smartphone Users.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices AppAgent: Multimodal Agents as Smartphone Users

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.195310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.195310Z digest=sha256:b13b2951d7a167b4b686abdbf92aa120dfa0b92dde8cc66134f4b085322f3f60

Observation 1128f696-60b7-4e44-a028-c2562ea3a7bf · outbound

This paper cites Tileflow: A framework for modeling fusion dataflow via tree-based analysis.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices Tileflow: A framework for modeling fusion dataflow via tree-based analysis

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.446934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.200194Z digest=sha256:8fbcb482004d79a5e914bb3206bcf1213eb7b6564203a540a8813f919e2e8d06

Observation 4f5c5103-a284-4dfd-8fe5-48f6e81bd9de · outbound

This paper cites and Yang, K.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices and Yang, K

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:20:14.426817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-12T16:20:03.204980Z digest=sha256:69f6f28049d296fd7e3a5ae975e10eae4041f0214b9106935f9a7c59bfc2b9a4

Observation 55b38dc0-c758-4f83-a6cd-dfa2cea1fcca · outbound

This paper cites write newline.

MAS-Attention: Memory-Aware Stream Processing for Attention Acceleration on Resource-Constrained Edge Devices write newline

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T16:20:03.210254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T16:20:03.210254Z digest=sha256:9e98d04daa429dcd7cc57026f02157849ac3dde4c64f9f7fe7340de69632ec69

Pith citing papers

No inbound Pith citation observations are available.