Pith. sign in

Paper Citation Record · LEDGER

MixFormer: Linear Transformer with Mixture of Memory Experts

As of 19 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2608.09468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09468 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:02:06.937137Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy30
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7bd4a140-c4f3-4aca-8f13-fe8cf5344647 · outbound

This paper cites Attention is all you need.Proceedings of the 2017 Advances in Neural Information Processing Systems, NeuraIPS, pages 5998–6008, 2017.

MixFormer: Linear Transformer with Mixture of Memory Experts Attention is all you need.Proceedings of the 2017 Advances in Neural Information Processing Systems, NeuraIPS, pages 5998–6008, 2017

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.657671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.733397Z digest=sha256:e8036bc8b1e120397c20ad67ed2ad89100e217857b1859b5582b4f1bf2d1cc37

Observation 2a19aaee-0952-4df1-9c0a-856f4aa08113 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MixFormer: Linear Transformer with Mixture of Memory Experts LLaMA: Open and Efficient Foundation Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.739145Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.739145Z digest=sha256:e10746e961b5e5c60c590de154552dec9d434fb26b7a92049946f288c791262f

Observation f1eba3b0-f82b-450e-84bb-51e0e2505169 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

MixFormer: Linear Transformer with Mixture of Memory Experts Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.744470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.744470Z digest=sha256:6a024c92ba20e0f54c28c626a33b9934f8ef346187f8893010653d6eb530c1e7

Observation 0fb668da-41b8-48ee-9f55-4115ea066474 · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407, 2024.

MixFormer: Linear Transformer with Mixture of Memory Experts The llama 3 herd of models.arXiv e-prints, pages arXiv–2407, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.749457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.749457Z digest=sha256:3d89987072710c68b88a4bae0078bfc1587f09a2c96fb89f8ad86c39ff17e1b8

Observation 77a00a8d-2fa9-42c1-81bf-23e2aa35a402 · outbound

This paper cites Molmo and pixmo: open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024.

MixFormer: Linear Transformer with Mixture of Memory Experts Molmo and pixmo: open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.630860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.755293Z digest=sha256:460b22c661221772990af1ac4a75f08a6d18e619039e532424d9370684940081

Observation 6a6a3153-d04c-4a01-aef2-f61433ec6a6b · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

MixFormer: Linear Transformer with Mixture of Memory Experts NVLM: Open Frontier-Class Multimodal LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.760415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.760415Z digest=sha256:86bb49f18524960366369faaa7754dd0a4bbf631ff4e9a050d3ed720480a2fd9

Observation c4a816df-d03b-4ba8-8028-3b5009247283 · outbound

This paper cites Efficient transformers: a survey.ACM Computing Survey, 55(6), 2022.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficient transformers: a survey.ACM Computing Survey, 55(6), 2022

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.615287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.767004Z digest=sha256:835187fed10b85528630de1c0348841b8a6c2148884fa71929d1d3efc1608ffb

Observation 2f607eb3-f53f-4eee-b132-8c578e47b8de · outbound

This paper cites A survey on efficient training of transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts A survey on efficient training of transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.598109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.772520Z digest=sha256:4232107978429e34d4db40c109a40cca73b17f4d877404e817577c06fdac4c8a

Observation e1c29fcc-6bc8-4563-a0c0-8ef0102c0970 · outbound

This paper cites an unresolved cited work.

MixFormer: Linear Transformer with Mixture of Memory Experts Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:02:07.581779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.777736Z digest=sha256:7c6b492e5679dab9ac1a607e21542e104a7329359d205b2293e5f187320efc96

Observation 1eb31702-f9b0-4ebf-b0f6-2acd135bf81d · outbound

This paper cites A survey on vision transformer.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):87–110, 2022.

MixFormer: Linear Transformer with Mixture of Memory Experts A survey on vision transformer.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(1):87–110, 2022

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.566504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.783024Z digest=sha256:38b9d54c1f28444f91610835a5187c92d0197488f8978c276423f54a4fa785b2

Observation 1df62038-dde7-43b7-b1fb-94f733cdd8a7 · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficiently modeling long sequences with structured state spaces

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.551246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.789114Z digest=sha256:ff654056c1bfca5aaac8a513e80e6e9708912c6f488ffc13ceb704f64d87cf3a

Observation dfa9049a-e736-42ae-b409-4e28161aa6c2 · outbound

This paper cites Diagonal state spaces are as effective as structured state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Diagonal state spaces are as effective as structured state spaces

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.534221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.794479Z digest=sha256:251b298f5e89300e74a1785487e932c4853a075607fb3bc1d7fa6357af86903c

Observation 5a49165a-d784-4664-b347-f09b3875177e · outbound

This paper cites an unresolved cited work.

MixFormer: Linear Transformer with Mixture of Memory Experts Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:02:07.519106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.799632Z digest=sha256:dc363cb94d5021ff18dbeef11476f953b3947c972ebd10e16986144c212b09f1

Observation fedf7564-1b39-4c14-a9b6-39fad529bd22 · outbound

This paper cites Liquid structural state-space models.

MixFormer: Linear Transformer with Mixture of Memory Experts Liquid structural state-space models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.503407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.804447Z digest=sha256:19719b4e5c255ba10e0d50559978946fc983211d1d1db0098f1751cf39b2f962

Observation 7484d9bd-91c1-4cbc-97c9-821d3219d5db · outbound

This paper cites Simplified state space layers for sequence modeling.

MixFormer: Linear Transformer with Mixture of Memory Experts Simplified state space layers for sequence modeling

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.488124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.809142Z digest=sha256:cb30e7b2a012d4068c591cfe3441e80af351f0d9e316f345c573fe6535694a1f

Observation affa02d5-02da-4bd1-9626-0ad06068fbf0 · outbound

This paper cites Retentive Network: A Successor to Transformer for Large Language Models.

MixFormer: Linear Transformer with Mixture of Memory Experts Retentive Network: A Successor to Transformer for Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.813888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.813888Z digest=sha256:f8aee2f6432bd4ff51951040dbced2ad1d4a63855f1f079fd2f75f00548bf714

Observation a4d0d1ab-6067-46f6-b07d-9b14f63f7418 · outbound

This paper cites Gated linear attention transformers with hardware-efficient training.

MixFormer: Linear Transformer with Mixture of Memory Experts Gated linear attention transformers with hardware-efficient training

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.472604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.818921Z digest=sha256:8bfe1aaaf76ade97115c22bcd71e52bc59e9547477972f47aca5f0d9cfc56264

Observation d97789cb-2202-4d0a-abf9-2067798803ac · outbound

This paper cites Transformers are rnns: fast autoregressive transformers with linear attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Transformers are rnns: fast autoregressive transformers with linear attention

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.456246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.823323Z digest=sha256:b268867833f29bc027f36dcea30331f5276118125e467e616fdb15435f81ebfb

Observation 3efb5ef3-2a30-4674-972b-c354604964b1 · outbound

This paper cites Efficient attention: attention with linear complexities.

MixFormer: Linear Transformer with Mixture of Memory Experts Efficient attention: attention with linear complexities

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.439782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.827677Z digest=sha256:a5852dd9557da2e332a77b7b09245a64d8e299ae3dfe29bd49899a057de8050f

Observation 7af4b9ed-8273-4831-8def-e00b15592c32 · outbound

This paper cites Linear transformers are secretly fast weight programmers.

MixFormer: Linear Transformer with Mixture of Memory Experts Linear transformers are secretly fast weight programmers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.423943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.832401Z digest=sha256:469a92415a75e77e6e6cc6f0c47056406985664c4651621a32ec46c632a0848f

Observation 349fb060-3e53-4895-95f2-46bc8956a66f · outbound

This paper cites Flatten transformer: vision transformer using focused linear attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Flatten transformer: vision transformer using focused linear attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.408559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.837260Z digest=sha256:c700bfb108b29fe6054e595805078a73926346d9cdfeb2525a82fc7611e6231b

Observation 902f38e1-f551-4875-817b-b88076078383 · outbound

This paper cites Polaformer: polarity-aware linear attention for vision transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts Polaformer: polarity-aware linear attention for vision transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.389436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.841842Z digest=sha256:8c2d522587238b86b622849f591cf55eda0501149dfbd767b19b4ac9bb325759

Observation e04374ce-f3de-4284-ac12-4545da574cf5 · outbound

This paper cites Linear Attention Mechanism: An Efficient Attention for Semantic Segmentation.

MixFormer: Linear Transformer with Mixture of Memory Experts Linear Attention Mechanism: An Efficient Attention for Semantic Segmentation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.846478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.846478Z digest=sha256:156609096347638ecada0d726e76e63eb2c9e2832447b8b4012e108c55cde1bb

Observation ec9b2a3f-6171-4bbc-836e-9a14c1f6e354 · outbound

This paper cites Random feature attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Random feature attention

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.370028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.851757Z digest=sha256:b957bb02001cba3ee2446aedc62a0d9c5c013e652eaee07a14cb5fba1505dc60

Observation df083ca3-9ea6-4d5e-8adb-151c05bc334b · outbound

This paper cites Rethinking attention with performers.

MixFormer: Linear Transformer with Mixture of Memory Experts Rethinking attention with performers

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.354293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.856505Z digest=sha256:0b7951f7a2eeeb47826c4cc2a70b599e4d5c13dcfdfce5744cab408d289e34a7

Observation 207c6656-18d6-4fc9-ab07-ec9c10cf617b · outbound

This paper cites Mamba: linear-time sequence modeling with selective state spaces.

MixFormer: Linear Transformer with Mixture of Memory Experts Mamba: linear-time sequence modeling with selective state spaces

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.337411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.861100Z digest=sha256:576bd65868cd4c146075a4e1056adce60c54f79f170ddc970640b3213a8e71b8

Observation 6bcdf29d-0268-42a2-936e-30e01f8fa144 · outbound

This paper cites Transformers are ssms: generalized models and efficient algorithms through structured state space duality.

MixFormer: Linear Transformer with Mixture of Memory Experts Transformers are ssms: generalized models and efficient algorithms through structured state space duality

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.321854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.865999Z digest=sha256:d2c5778792d434fe8755a5451177f7ffea7b59a6212c3806cf1eff5e7320804e

Observation b3b9a948-da53-456b-918c-639b4a1da352 · outbound

This paper cites An Attention Free Transformer.

MixFormer: Linear Transformer with Mixture of Memory Experts An Attention Free Transformer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.870460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.870460Z digest=sha256:ea02a044e707a4c7dfa711bddad506efda6c55657a1897f4851d03521383c42a

Observation 366b724a-9899-4afc-89e5-c6030e9a6aa5 · outbound

This paper cites Rwkv: Reinventing rnns for the transformer era.

MixFormer: Linear Transformer with Mixture of Memory Experts Rwkv: Reinventing rnns for the transformer era

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.305242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.875703Z digest=sha256:fed3f2c2d448b9d7364da57b5a2536b4fde0fda18b07c86f24ebb6ec9494eba0

Observation e551ac0a-a7a5-42e3-954b-488100443664 · outbound

This paper cites Group normalization.

MixFormer: Linear Transformer with Mixture of Memory Experts Group normalization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.880665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.880665Z digest=sha256:c202d0eb6e552a47d85e1909818040535d8ea707604e43afbbb0b25bc472d43f

Observation fded8756-0e96-4f2e-b5e6-fe98c74c77b2 · outbound

This paper cites EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction.

MixFormer: Linear Transformer with Mixture of Memory Experts EfficientViT: Multi-Scale Linear Attention for High-Resolution Dense Prediction

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.885235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.885235Z digest=sha256:4ace5c5bebedda77c5a671c1f999a4724586f842d4eb707e6181f56507653c2a

Observation dd78c79f-08f4-42e2-ae0f-67a5679d4a35 · outbound

This paper cites Soft: softmax-free transformer with linear complexity.Proceedings of the 2021 Advances in Neural Information Processing Systems, NeuraIPS, 34:21297–21309, 2021.

MixFormer: Linear Transformer with Mixture of Memory Experts Soft: softmax-free transformer with linear complexity.Proceedings of the 2021 Advances in Neural Information Processing Systems, NeuraIPS, 34:21297–21309, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.274337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.890292Z digest=sha256:040f665f8a4647704adf6eb4c4c15d353eefd14f00a6d403aed2eaa8d838b033

Observation f91b6789-a3d1-4e26-9a39-a2b67165695a · outbound

This paper cites Nystr¨omformer: a nystr ¨om-based algorithm for approximating self-attention.

MixFormer: Linear Transformer with Mixture of Memory Experts Nystr¨omformer: a nystr ¨om-based algorithm for approximating self-attention

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.254217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.895575Z digest=sha256:493f462a54d777b2c53d1d29d442972169ddb95afe6eb1c44439d4a6b42838b4

Observation 594cde54-6c43-4f90-abf2-a79a4cdf328c · outbound

This paper cites Outrageously large neural networks: the sparsely-gated mixture-of-experts layer.

MixFormer: Linear Transformer with Mixture of Memory Experts Outrageously large neural networks: the sparsely-gated mixture-of-experts layer

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.236513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.900238Z digest=sha256:fd95cab8d8929443c07e21098f3b020b37742b86b75570a4e2b1e9e9605a6cbf

Observation 11156257-0005-4f0a-aff0-29a3e7c52132 · outbound

This paper cites Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization.

MixFormer: Linear Transformer with Mixture of Memory Experts Don’t give me the details, just the summary! topic-aware convolutional neural networks for extreme summarization

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.217429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.904806Z digest=sha256:c32cc21df798da6b2c7177e44cc3fa94056640012e123b1e75c9085ee2aaff54

Observation 12c2d22a-8cee-4abb-9529-43cf6d45b9cb · outbound

This paper cites Long range arena: a benchmark for efficient transformers.

MixFormer: Linear Transformer with Mixture of Memory Experts Long range arena: a benchmark for efficient transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.198005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.909593Z digest=sha256:92a7466ea63f15c22f5a757a991872622553a4800b16ec0c341dd2d1ce434fa2

Observation 985d0c57-9069-4181-bcf6-3803bf82d05e · outbound

This paper cites Listops: a diagnostic dataset for latent tree learning.

MixFormer: Linear Transformer with Mixture of Memory Experts Listops: a diagnostic dataset for latent tree learning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.180622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.914154Z digest=sha256:891aee7aa256c5d2a0dcad4649fd7bcb1fc3da33ba7b7271e0dfa929c0aebdd8

Observation b2114500-42d9-4d0a-8f99-32bfeb3f4060 · outbound

This paper cites Learning word vectors for sentiment analysis.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning word vectors for sentiment analysis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.162071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.918732Z digest=sha256:10416333d2e94ef1cce5d8f480f00aae36fd73100faf9e46fd3a9e0cabe47a9d

Observation e8449253-7c69-42bd-b3e2-7f360bc81a69 · outbound

This paper cites The acl anthology network corpus.Language Resources and Evaluation, 47(4):919–944, 2013.

MixFormer: Linear Transformer with Mixture of Memory Experts The acl anthology network corpus.Language Resources and Evaluation, 47(4):919–944, 2013

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.146128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.923391Z digest=sha256:cbf51d4ff421c680d88d60e2de65f1732e76d60636179d89931f371a8109ef5a

Observation a66fb356-0427-48a0-ba50-37fff294600b · outbound

This paper cites Learning long-range spatial dependencies with horizontal gated recurrent units.Proceedings of the 2018 Advances in Neural Information Processing Systems, NeuraIPS, 31, 2018.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning long-range spatial dependencies with horizontal gated recurrent units.Proceedings of the 2018 Advances in Neural Information Processing Systems, NeuraIPS, 31, 2018

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.128493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.927739Z digest=sha256:e0fba759a6b9548be3737c9f65b30f23b658146a1f7ede15af3a9b5ad8e6c4e3

Observation 867a0f1f-591f-4288-9ec6-cfd59ae19a5d · outbound

This paper cites Learning multiple layers of features from tiny images.

MixFormer: Linear Transformer with Mixture of Memory Experts Learning multiple layers of features from tiny images

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:02:06.932524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:02:06.932524Z digest=sha256:b30952b2eaf54a81c2b2abe945039bd6e3c9675577d5ac62cd3c0a15c73d7979

Observation 12adb45a-901c-4675-b703-276e1affadcb · outbound

This paper cites Emnist: Extending mnist to handwritten letters.

MixFormer: Linear Transformer with Mixture of Memory Experts Emnist: Extending mnist to handwritten letters

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:02:07.101239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T17:02:06.937137Z digest=sha256:1c7ae4ffd5d948a370268d55a15a711c424485976163390759ebc0bdfae93271

Pith citing papers

No inbound Pith citation observations are available.