Pith. sign in

Paper Citation Record · LEDGER

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning

As of 9 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 2 inbound Pith citation observations for arXiv:2508.18756.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.18756 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:17:42.154186Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:13:54.031036Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:06:43.816739Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 06d29732-220e-4398-b214-e33e734496f2 · outbound

This paper cites Program Synthesis with Large Language Models.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Program Synthesis with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.862060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.862060Z digest=sha256:bca259f256d9b9fdf7d16803726cb0c08eb8f3af514a28444988dcc1e06fa1fd

Observation bf4eed67-8daf-4c43-8846-cac6c833c7be · outbound

This paper cites Memory Layers at Scale.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Memory Layers at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.871650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.871650Z digest=sha256:4375acb3081aba36465406377a611769e72c643756b3ebd1f166341f9cc5a375

Observation 8ec2a355-90da-4f99-b3b6-d2b9780417c3 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Piqa: Reasoning about physical commonsense in natural language

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:43.095234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:17:41.882241Z digest=sha256:a0d2455b83f4809f66230f8a0b674d82acced4acfb8f6c08a244ad7ccb655dbc

Observation 7176ecf7-38b0-4c4f-8966-32b3acc9f0f5 · outbound

This paper cites an unresolved cited work.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.891161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.891161Z digest=sha256:700346b0e10e26a40d920b2d09696e200ab161dd33f5468e4c117c1daead061a

Observation 4ec88c91-b50e-4e60-9e65-e55879060f8b · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.898090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.898090Z digest=sha256:38bc455401295e2a70dd8c05d8ed8b7aba74f53579c312795667f9a257a08af9

Observation 6e86179a-0fa9-41f1-8fdf-703e4db0a865 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.903827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.903827Z digest=sha256:dbe452f952dc9719a70c7e49e0a0d3588c9649c9b3edc12a15595a5abdfc7708

Observation f1c6fb34-e731-41cc-aa3f-ff880a8298ca · outbound

This paper cites Approximating Two-Layer Feedforward Networks for Efficient Transformers.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Approximating Two-Layer Feedforward Networks for Efficient Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.909750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.909750Z digest=sha256:a44c9d83a92cf71450365f10702765e205ce7790cebecbc18fa5d5d5cf2350c2

Observation bd542213-d772-41c0-9435-d9a454db8c4f · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.916736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.916736Z digest=sha256:cd10efd4474cf575b81a90398ab605f6b6e7ccaa5b0bd694b2dbf3412f8944fb

Observation 21d0a2ff-763f-45c3-83b6-db892710129a · outbound

This paper cites DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:43.057371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:17:41.924101Z digest=sha256:8d884236dfcd0e067016c74bd43ee260e492fea6ebd5e9a43e655c2b85f97bd8

Observation b28f6439-323b-4962-b986-44410a875be0 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research, 23(120):1–39, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.932009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.932009Z digest=sha256:a0774d9f90b37c6e23d5d7b4452a9568de9af064a88f3e2b4faa6355016fed66

Observation 90764af7-1ac7-4b06-8082-f40da840398e · outbound

This paper cites FastMoE: A Fast Mixture-of-Expert Training System.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning FastMoE: A Fast Mixture-of-Expert Training System

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.941220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.941220Z digest=sha256:99fc90d7b8569c6cdd7b7a890bca5038e7c1574128c27cfa0b5b44b31fb6678a

Observation 7b5d0f0f-6f70-4f1f-9b93-74b345fadb35 · outbound

This paper cites Mixture of A Million Experts.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Mixture of A Million Experts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.950238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.950238Z digest=sha256:92f084d56849cd0c6f430350474d436910cc58d42be967e235c9f84195282c87

Observation e8412630-7a60-4108-8d70-e5f05c4506fb · outbound

This paper cites Aligning ai with shared human values.Proceedings of the International Conference on Learning Representations (ICLR), 2021.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Aligning ai with shared human values.Proceedings of the International Conference on Learning Representations (ICLR), 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.956630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.956630Z digest=sha256:312ec8ed82151adcd669698fec31fa7f1d188eff3ecbb5e43e62edf6077f752f

Observation 74431a28-f424-4255-9527-53e7dcb99fd1 · outbound

This paper cites Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Measuring massive multitask language understanding.Proceedings of the International Conference on Learning Representations (ICLR), 2021

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.963854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.963854Z digest=sha256:92da9be37d76e4b998853d114483e75c4f137214aa38fe34c948372f799e55aa

Observation 552c83df-6bf9-4d15-91a2-b1eebbd700d8 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.NeurIPS, 2021.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Measuring mathematical problem solving with the math dataset.NeurIPS, 2021

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.973737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.973737Z digest=sha256:a6ce5689b42bf800ba3fe09a638559968121f90596ae6f4fa9cef9ce879998dc

Observation 2dd5a2a9-6fac-496a-b3a1-432231d05997 · outbound

This paper cites Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Over-Tokenized Transformer: Vocabulary is Generally Worth Scaling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.981651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.981651Z digest=sha256:08b7a9cae17178a6a18ff36e56b8b8a9354e70cef0359d2ca397dcde159f60d3

Observation d9fc7fbc-9f52-419b-9d0c-6cbae2d58457 · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:42.964852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:17:41.989216Z digest=sha256:9e924d2f6a30642f3e57a8ee5a212693140c644b4f973a75f739ef8b745b0ea2

Observation e8ecbb92-aa3c-432f-9a7c-a9f8c04ef171 · outbound

This paper cites Ultra-Sparse Memory Network.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Ultra-Sparse Memory Network

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:41.998354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:41.998354Z digest=sha256:b870df5697e055c53b24ed693ca444fce42b3e7087b243ee2d43de8d84c45f5c

Observation f339dd7c-1b5f-4c73-b747-c9875d60c6f5 · outbound

This paper cites dots.llm1 Technical Report.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning dots.llm1 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.004470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.004470Z digest=sha256:6e803e7c9fff791430b0bf5110fb3afbf9673dcb01b9552235e36353310a0887

Observation aa57d30e-03f4-47d6-8ff9-e67b51d831d7 · outbound

This paper cites Product quantization for nearest neighbor search.IEEE transactions on pattern analysis and machine intelligence, 33(1):117–128, 2010.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Product quantization for nearest neighbor search.IEEE transactions on pattern analysis and machine intelligence, 33(1):117–128, 2010

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.010279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.010279Z digest=sha256:b165546e9651bda1a0e7a49a025642c19d2908328531ee0b6481f99a5bf31d89

Observation ac4e10c0-960f-4123-be6b-4c0c545c5a69 · outbound

This paper cites Mixtral of Experts.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Mixtral of Experts

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.015268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.015268Z digest=sha256:7289bef3d66e4aa79d360af06034b1571bf6ebc818636c9a98e417a35d188b15

Observation 5448c1fb-c3ec-4903-9c3c-64346f545032 · outbound

This paper cites an unresolved cited work.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.020585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.020585Z digest=sha256:4ba0e52c9feab465d5f25c06e3399646b361b2729981e34ba0ef1711d24ac581

Observation c4b2915f-07ec-48e6-9918-706388389735 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.025904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.025904Z digest=sha256:6228fedc5a13b846e3e6b39c6ce8073c6967af4b2f180bbdf30ae8469643197a

Observation e005e7a5-2a00-4426-85f7-543c19ac5dec · outbound

This paper cites Large Product Key Memory for Pretrained Language Models.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Large Product Key Memory for Pretrained Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.032313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.032313Z digest=sha256:36971deaa2f7f4d16b5266831d75dae5a29fa303b594db51b8b3c4b3f99cd713

Observation e766e8a8-41d6-4b09-a7bd-4f0b18b05e96 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Scaling Laws for Fine-Grained Mixture of Experts

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.037489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.037489Z digest=sha256:292ce31618b84111b3ff1ba0c8e5e79b7925320f79a527368a01a7e55fae6a8c

Observation 33ab54b5-1975-4c31-8a67-f73c40d16c14 · outbound

This paper cites Large memory layers with product keys.Advances in Neural Information Processing Systems, 32, 2019.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Large memory layers with product keys.Advances in Neural Information Processing Systems, 32, 2019

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:42.918347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:17:42.043484Z digest=sha256:97bded212721a80b8dabdab884baa5ce38de01a708bca63aae7b18c30dee1d01

Observation f8dc512d-f6d8-4d4e-8e75-efa4d47529ac · outbound

This paper cites DeepSeek-V3 Technical Report.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning DeepSeek-V3 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.050746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.050746Z digest=sha256:431c3a636a9e7fe22aad20f8b8a22086ee2513b2ea8df0218a0e199ce9320fbf

Observation 21d9ec13-c829-41cc-9880-81543435c15d · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.058600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.058600Z digest=sha256:a12da86ce559735b81d25ce18fd4cc0bce28a02f21deff1d90321123a9578269

Observation d9140e8c-5cb1-454c-9934-2386a73cfa2f · outbound

This paper cites OLMoE: Open Mixture-of-Experts Language Models.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning OLMoE: Open Mixture-of-Experts Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.066072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.066072Z digest=sha256:c61b2b322b7e76b4e5e987dbeccaf3997ffd98f6e9e29dda67d56d6f86990080

Observation 0b698c83-5376-4f15-832e-f29ac974cc19 · outbound

This paper cites Transformers without Tears: Improving the Normalization of Self-Attention.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Transformers without Tears: Improving the Normalization of Self-Attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.074129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.074129Z digest=sha256:4f04e4a448444b3a69e6ec50c191c89260996997e9ac026733e7a98d35246e59

Observation 673c66de-f8de-4e25-af8b-aa7ff99e8d7e · outbound

This paper cites Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.081799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.081799Z digest=sha256:ab293e2d70c97fdef1a39d54fb806756c8ce33badc698ff9679e94d5255ca540

Observation 7e3617df-af46-4e77-bc6f-63f84a873655 · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.088345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.088345Z digest=sha256:9ed1e28ffe3615db235a195fbf131dc1b10f9c11fb8fd4c112a883b936cb7c08

Observation 04d84cf5-9f83-46ab-8cbd-00cb00d46f2f · outbound

This paper cites GLU Variants Improve Transformer.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning GLU Variants Improve Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.097567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.097567Z digest=sha256:2ae176f798cf6b5d5921745dbdd1165fad94de5ee2a4d457d3ed762e15b0202e

Observation d62372b7-4833-4869-904d-1f40b18da95c · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.104052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.104052Z digest=sha256:0d14f9c558ee932ab4908fd68ef919755cd09ce72fef5cc777a2c0017c123f10

Observation ce438d7d-5e54-4a3c-b4a0-78845ca3dc63 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.109834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.109834Z digest=sha256:388b408f9944fbfdd0bcea36a54735da2c6afe536c55995e294ee3e180133466

Observation e9893b79-0301-4daa-88ca-f6b7c2aa2e11 · outbound

This paper cites CommonsenseQA: A question answering challenge targeting commonsense knowledge.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning CommonsenseQA: A question answering challenge targeting commonsense knowledge

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.115682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.115682Z digest=sha256:5ee33bd0a05bc9632867bd6ca3e7e2f7b82c8fad0493a5a9f1d9b81594a1450a

Observation 7a0b0e3b-efbc-4799-98ed-7533ba8244f1 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:42.858993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:17:42.120614Z digest=sha256:6832aebf09ca63142defd3b5bcd3c3e3c31dee576f526c3d5292f21522f6b061

Observation 55063835-0b19-4ea0-8d12-83f01057f58c · outbound

This paper cites Moec: Mixture of expert clusters.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Moec: Mixture of expert clusters

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:42.834215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:17:42.125390Z digest=sha256:526b58f3d75447d141f332bd82ac7e59f16133b6a3e4e184d23e24c1ce000ac7

Observation 5dcddd93-72b7-4b4a-bf0d-487940f0af07 · outbound

This paper cites Qwen3 Technical Report.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Qwen3 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.130675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.130675Z digest=sha256:8a1bad79fc5540184f5b6107aa738dca1ff61b095623623134feea451e3636ca

Observation 277491ed-0343-46cb-b108-e7c227a182f7 · outbound

This paper cites Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Expert Race: A Flexible Routing Strategy for Scaling Diffusion Transformer with Mixture of Experts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.136085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.136085Z digest=sha256:32adc0d5bf1236aff9cee5243f251f87f229338430f067fd7263d71370bfdd22

Observation 0fd5e50b-593b-4303-934a-0ef98b620337 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Hellaswag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.142224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.142224Z digest=sha256:859850de3e3c88cff7fef9673306493000cd1fb5a3fef34c8a07dcc50fa886d2

Observation f386348f-bd3d-400b-8408-925a8d03cd26 · outbound

This paper cites Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T16:17:42.148465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:17:42.148465Z digest=sha256:22fa24fd8a5321140d06a522a5f4493139554400fc77f75eb799c918dedde758

Observation d97cc08f-2759-43f4-9634-36c8a1599b67 · outbound

This paper cites Pre-values.

UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning Pre-values

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:17:42.787037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T16:17:42.154186Z digest=sha256:6e3a5c4b715bb08d9fa1e049de2e3f8109c5f1d3a5fee4e9f38b196efba4b15f

Pith citing papers

Observation 9c5b9c36-33bd-4b84-8514-b1935b865e6d · inbound

MIDUS: Memory-Infused Depth Up-Scaling cites this paper.

MIDUS: Memory-Infused Depth Up-Scaling UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:03:36.159265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T22:02:42.297041Z digest=sha256:d6061584c2ceb3450e2380b5ec271ab40d8b6fd2a7482a3c21bff5bd0ba36799

Observation 18678fdf-ed1f-4050-b3ff-cb9b215df4f0 · inbound

SinkRec: Mitigating Semantic State Sink in Long Sequence Recommendation with Memory-Conditioned Gated Delta Networks cites this paper.

SinkRec: Mitigating Semantic State Sink in Long Sequence Recommendation with Memory-Conditioned Gated Delta Networks UltraMemV2: Memory Networks Scaling to 120B Parameters with Superior Long-Context Learning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:43.818439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T07:13:54.031036Z digest=sha256:9ce3f5386cb9fef89467f55093cc3d5bd81ae6804872cd8799e6529b27aaeac0