REVIEW 3 major objections 5 minor 59 references
GRACE: Generative Recommender Acceleration Engine for Real-Time Ads Retrieval
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Generative Target Matching moves advertiser eligibility checks into the decoding loop, lifting the exact targeting pass rate of generated ads from 23.55% to 40.42% while specialized kernels cut decoder latency 11.1x.
desk verdict GTM is a genuine extension of constrained decoding to personalized eligibility, and the wide-beam decoder kernels are a real engineering win; the pass-rate gain, however, is measured through a Bloom filter that may include false positives, so read 23.55%→40.42% with caution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Generative Target Matching (GTM): a decode-time eligibility layer that augments a catalog-valid constrained-decoding trie with per-node matchers — packed 64-bit-integer bitmasks for low-cardinality targeting attributes and $w=256$-bit Bloom filters for high-cardinality location attributes — so a candidate token advances only if its SID prefix still contains at least one matching ad; the matcher for a node is the OR-union of its subtree's ads, which is conservative but false-positive-prone, and $k$-way partitioning is offered as a mitigation. The compute story is carried by four decoder optimizations for the wide-beam, short-sequence regime: a beam-as-query cross-attention reshape that lets a
What would settle it
Re-run the Section 6.2 pass-rate experiment with a real production location distribution (or a heavily concentrated one, e.g., a few locations per user) instead of the synthetic 64 random locations, keeping the same 30M-SID index and matcher settings; if the CD+GTM final ad-level pass rate moves back toward the 23.55% CD-only baseline, the eligibility claim does not transfer to production-shaped data. A complementary check: instrument the OR-union matchers' false-positive rate per trie level and confirm the bucket-level pass-rate breakdown in Table 3 reproduces.
Extended reading notes
Core claim
The paper's central claim is that ads generative retrieval can be served in real time only if advertiser eligibility is enforced inside the decoding loop, and that this is feasible with the right matchers and kernels. GRACE's Generative Target Matching (GTM) augments catalog-valid constrained decoding: at each SID-token position, a candidate token survives only if the prefix it extends still contains at least one ad whose targeting rules match the request, using packed bitmask matchers for low-cardinality attributes (country, age, gender) and 256-bit Bloom filter matchers for high-cardinality location attributes, stored on the nodes of the SID trie. Because each matcher is an OR-union over t
Load-bearing premise
The eligibility result is measured on synthetic user-location data — 64 random locations per user — against one fixed 30M-SID index; if production location data is far more concentrated or the targeting rule mix differs, the measured jump from 23.55% to 40.42% could change materially.
Editorial extensions
If this is right
- Eligibility becomes a decode-time property: with GTM, only SID prefixes that still contain at least one request-eligible ad advance in beam search, so a much larger fraction of generated ads survive exact ad-level targeting (23.55% to 40.42%).
- The distribution of generated ads shifts: the share of requests landing in the ineligible-heavy <5k generated-ads bucket drops from 50.78% to 31.84% while the 10k+ bucket grows from 12.50% to 29.10%, so downstream filtering and ranking receive more usable candidates.
- The latency path fits production budgets: with dynamic beam sizes and all GTM matchers enabled, full-model P99 latency is 53.6 ms, inside the ~70 ms compute window after batch accumulation and under a P99 <100 ms end-to-end budget.
- $k$-way partitioning is not worth its cost: despite lowering matcher fill rates, it does not improve final pass rate (40.27%/40.45% vs. 40.42% for unpartitioned CD+GTM) and adds storage and kernel overhead, so the preferred operating point is the unpartitioned layout.
- Dynamic per-step beam sizes reduce model FLOPs per request by 16.2% (533 to 446 GFLOPs) with no reported quality loss, and GTM overhead stays small enough that the full bitmask+Bloom configuration remains within budget.
Reading between the lines
- The pass-rate numbers rest on synthetic user locations — 64 random locations per user — and the paper never tests how OR-union false positives behave on production location distributions, which are likely far more concentrated; with fewer distinct locations per request, the Bloom ANY-of-many predicate could pass more often, moving the 40.42% figure either way.
- The beam-as-query cross-attention reshape and coalesced short-sequence self-attention are layout and kernel transformations with correctness arguments independent of ads; they would transfer to any encoder-decoder beam-search decode with short outputs, such as speech or translation lattices, a generalization the paper does not claim.
- The paper evaluates eligibility only as pass rate, not downstream ranking quality; whether the extra ~17 percentage points of eligible ads translate into better ad selection or business outcomes is an open question that a production A/B test would need to settle.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. GRACE addresses two serving problems for productionizing generative ads retrieval with encoder-decoder Transformers: eligibility (generated Semantic IDs must satisfy advertiser targeting rules) and compute (wide-beam decoding with strict latency). For eligibility, the paper proposes Generative Target Matching (GTM), which augments catalog-valid constrained decoding with per-node bitmask matchers for low-cardinality attributes and Bloom filter matchers for high-cardinality location attributes, optionally with k-way partition matchers. Evaluation on a 30M-SID index reports that SID-level GTM raises the final ad-level pass rate from 23.55% to 40.42% over constrained decoding alone (Table 3). For compute, GRACE contributes a beam-as-query cross-attention layout, coalesced and fused short-sequence self-attention kernels, a paged KV cache for beam rearrangement, dynamic per-step beam sizes, and a pipelined four-stage serving path. On GH200, the optimized decoder reduces P50 latency from 196.7 ms to 17.8 ms (Table 6), and the full pipeline with GTM remains within a 70 ms compute window (D3 full P99 = 53.6 ms). The paper also reports a negative result: k-way partitioning does not materially improve final pass rate over unpartitioned CD+GTM.
Significance. If the results hold, the paper makes a useful contribution to generative retrieval serving. It clearly identifies the gap between catalog-valid constrained decoding and personalized audience targeting, and it proposes a practical token-level filtering mechanism. The decoder optimizations are well-motivated for the wide-beam, short-sequence regime and are supported by detailed microbenchmarks and NCU counters; the arithmetic in the latency tables is internally consistent. The cross-attention layout correctness argument in Appendix B is a nice touch. The main eligibility result, however, is currently measured through Bloom-based matching on synthetic location data, so its external validity is not yet established. The paper would be significantly stronger if it quantified the gap between Bloom-semantics pass rate and true eligibility. The compute contribution is more solid and would likely stand on its own.
major comments (3)
- [§3.1, §4.3, Table 3] The headline pass-rate metric is called 'exact ad-level target matching pass rate,' but Section 3.1 says the CPU-side 'ad-level exact target matcher' enforces 'the same bitmask and Bloom filter targeting semantics as GTM.' Section 4.3 explicitly acknowledges that Bloom matching has false positives. Consequently, the final pass rate in Table 3 counts Bloom hash collisions as passes. The reported improvement (23.55% to 40.42%) is concentrated in the Bloom column (40.74% to 65.70%), and Table 2 shows Bloom root fill 0.979 at the root and 0.475 at Pos1. With 64 user locations and ANY-of-many semantics, the fraction of Bloom passes that are true string-set intersections may be substantially lower. Please quantify Bloom false positives against exact string-set intersection and report the pass rate under true eligibility, or explicitly reframe the central claim as 'Bloom-semantics-compatible pa
- [§6.1, §6.2] The pass-rate evaluation uses synthetic user-location data with 64 random locations per user, but the paper gives no description of the generation procedure, the location universe, or how this compares with production request-location distribution. GTM's Bloom predicate is ANY-of-many over the user's locations, so the pass rate is strongly sensitive to Nloc and to how locations are distributed. Without sensitivity analysis (e.g., varying Nloc, using a real location sample, or reporting location-population statistics), the central 23.55%->40.42% improvement cannot be taken as representative of production targeting behavior.
- [§6.3, Tables 4–6] Latency results are reported as point estimates (P50/P99) without repetitions, confidence intervals, or number of trials. This is especially important for the P99 claims (e.g., D3 full P99 = 53.6 ms vs. the 70 ms compute window), where run-to-run variance can determine feasibility. Please report multiple runs or otherwise characterize the variance of the reported latency numbers.
minor comments (5)
- [§3.1, Figure 1] The label 'exact ad-level target filtering' in Figure 1 and Section 3.1 is misleading given that the matcher uses Bloom filters; consider renaming to 'ad-level target filtering' or 'ad-granularity target filtering.'
- [§6.3] The sentence 'About 30 ms is used to accumulate batches of 16 users' is presented as a fact but is not measured or referenced. Please provide support or soften the budget statement.
- [§3.2] The 533 GFLOPs per request figure should state explicitly that it is for the decoder (or for the full model) and whether it includes the encoder and GTM kernels; the dynamic-beam 446 GFLOPs figure then needs the same clarification.
- [References] The author name 'Yavuz Y etim' is missing letters; it should be 'Yavuz Yetim.' Also, the title in the full text is rendered without spaces: 'GENERATIVERECOMMENDERACCELERATIONENGINE.'
- [§6.2] Table 3 would benefit from a sentence explicitly noting that the 'Bloom' column is the pass rate conditional on the bitmask pass, since the final column is the product of the two preceding columns; this is currently implied but not stated.
Circularity Check
No significant circularity: GRACE's pass-rate and latency results are empirical measurements against fixed external baselines, with no fitted parameter renamed as a prediction and no load-bearing self-citation chain.
full rationale
GRACE's central claims are measured outcomes, not derived predictions that recycle their own inputs. The ad-level pass-rate gain (Section 6.2, Table 3) compares CD-only, CD+GTM, and k-way variants through the same CPU-side ad-level matcher and the same 30M-SID index; GTM is a deliberately coarser SID-prefix OR-union filter (Sections 4.2-4.3), so the improvement is not the evaluation criterion by construction. The latency results (Tables 4-6) are benchmarked against FlashAttention-2/3 and disclosed model hyperparameters and beam schedules, i.e., external baselines rather than self-referential fits. The paper's self-citations to Meta infrastructure such as STATIC for trie layout, InterFormer for the encoder, and PLUM/OneRec for generative-retrieval background are building blocks or related work, not uniqueness theorems or ansatz-forcing citations. Section 6.1's synthetic user-location data and Section 4.3's Bloom containment semantics create a production-transfer validity caveat, but they do not make the comparison circular because all methods are evaluated with the same matcher and baseline. No load-bearing step in the paper reduces to its own input; accordingly, no circularity steps are identified.
Assumptions & free parameters
free parameters (5)
- Dynamic beam schedule (M1,M2,M3,M4) =
(1,512,1024,1024)
- Static k-way partition sizes =
k=(64,64,32,8) across decode positions
- Dynamic k-way k_max and fill-rate target =
k_max=32, target fill rate 0.25
- Matcher widths =
bitmask 7x64-bit, Bloom 4x64-bit (w=256)
- Synthetic user-location data size =
64 locations per user
assumptions (5)
- domain assumption The GTM SID trie and matcher tables are built asynchronously and are assumed to reflect current advertiser targeting rules during decode.
- domain assumption The Bloom containment test with OR-ed subtree filters is a conservative proxy for exact ad-level location targeting.
- standard math The cross-attention beam-as-query reshape is a layout-only transformation.
- domain assumption The encoder-decoder model has fixed SID length 4, vocabulary 512, 3 layers, 16 heads, head dim 128.
- domain assumption GPU microbenchmark measurements on GH200 with batch size 16 generalize to production serving conditions.
Cite this review
Pith. "Pith review of GRACE: Generative Recommender Acceleration Engine for Real-Time Ads Retrieval." pith.science (2026). https://pith.science/paper/BQ4RKJQW
@misc{pith2026260800938,
author = {Pith},
title = {Pith review of: GRACE: Generative Recommender Acceleration Engine for Real-Time Ads Retrieval},
year = {2026},
howpublished = {\url{https://pith.science/paper/BQ4RKJQW}},
note = {Machine review of arXiv:2608.00938}
}
read the original abstract
Productionizing generative recommenders for high-volume, real-time ads retrieval creates two serving challenges: eligibility, ensuring that each generated ad is eligible for the request under the advertiser's audience targeting rules, and compute, which requires meeting strict latency and GPU cost requirements while remaining capable of generating thousands of ads per request with wide-beam decoding. This paper presents GRACE, a serving system for ads generative retrieval that addresses both challenges. For eligibility, GRACE introduces Generative Target Matching (GTM), which extends catalog-valid constrained decoding with personalized filtering over Semantic ID (SID) prefixes using bitmask and Bloom filter matchers derived from targeting rules. SID-level GTM improves final ad-level target matching pass rate from 23.55% to 40.42% over constrained decoding alone. For compute-cost and latency, GRACE targets encoder-decoder Transformers, which are more lightweight than LLMs. It redesigns the decoder around the wide-beam, short-sequence regime, covering attention kernels, KV cache, and beam search optimizations. On NVIDIA GH200, compared with the faster of FlashAttention-2 and FlashAttention-3 baselines, GRACE improves cross-attention latency by 68.0 times and self-attention latency by 23.4-25.8 times across decode steps. Together, these changes reduce decoder latency by 11.1 times, keeping ads generative retrieval within latency and compute requirements.
Figures
Reference graph
Works this paper leans on
-
[1]
and Ermon, Stefano and Rudra, Atri and R
Dao, Tri and Fu, Daniel Y. and Ermon, Stefano and Rudra, Atri and R. Proceedings of the 36th International Conference on Neural Information Processing Systems , pages =
-
[2]
Dao, Tri , booktitle =
-
[3]
Shah, Jay and Bikshandi, Ganesh and Zhang, Ying and Thakkar, Viral and Ramani, Pradeep and Dao, Tri , booktitle =
-
[4]
Zadouri, Ted and Hoehnerbach, Markus and Shah, Jay and Liu, Timmy and Thakkar, Vijay and Dao, Tri , booktitle =
-
[5]
Ye, Zihao and Chen, Lequn and Lai, Ruihang and Lin, Wuwei and Zhang, Yineng and Wang, Stephanie and Chen, Tianqi and Kasikci, Baris and Grover, Vinod and Krishnamurthy, Arvind and Ceze, Luis , booktitle =
-
[6]
and Zhang, Hao and Stoica, Ion , booktitle =
Kwon, Woosuk and Li, Zhuohan and Zhuang, Siyuan and Sheng, Ying and Zheng, Lianmin and Yu, Cody Hao and Gonzalez, Joseph E. and Zhang, Hao and Stoica, Ion , booktitle =
-
[7]
and Han, Ningren , booktitle =
Su, Zhengyang and Katsman, Isay and Wang, Yueqi and He, Ruining and Heldt, Lukasz and Keshavan, Raghunandan and Wang, Shao-Chuan and Yi, Xinyang and Gao, Mingyan and Dalal, Onkar and Hong, Lichan and Chi, Ed H. and Han, Ningren , booktitle =
-
[8]
and Samost, Jonathan and Kula, Maciej and Chi, Ed H
Rajput, Shashank and Mehta, Nikhil and Singh, Anima and Hulikal Keshavan, Raghunandan and Vu, Trung and Heldt, Lukas and Hong, Liang and Tay, Yi and Tran, Vinh Q. and Samost, Jonathan and Kula, Maciej and Chi, Ed H. and Sathiamoorthy, Maheswaran , booktitle =
Show all 59 references
-
[9]
He, Ruining and Heldt, Lukasz and Hong, Lichan and Keshavan, Raghunandan and Mao, Shifan and Mehta, Nikhil and Su, Zhengyang and Tsai, Alicia and Wang, Yueqi and Wang, Shao-Chuan and Yi, Xinyang and Baugher, Lexi and Cakici, Baykal and Chi, Ed and Goodrow, Cristos and Han, Nin...
-
[10]
Sun, Chuan and Yu, Nan and Lu, Hao and Wang, Luming and Pu, Yun and Liu, Guang and Bhatia, Nikhil and Musumeci, GP , year =
-
[11]
Li, Huayu and Liu, Xiaoyi and Nie, Jade and Wen, Ellie and Yang, Chunzhi and Yang, Jiyan and Yu, Nancy and Beg, Habiya and Arditi, Gil and Bhatia, Neeraj , year =
-
[12]
2016 , pages =
Covington, Paul and Adams, Jay and Sargin, Emre , booktitle =. 2016 , pages =
2016
-
[13]
Shane , booktitle =
Gallagher, Luke and Chen, Ruey-Cheng and Blanco, Roi and Culpepper, J. Shane , booktitle =
-
[14]
Wang, Xuewei and Jin, Qiang and Huang, Shengyu and Zhang, Min and Liu, Xi and Zhao, Zhengli and Chen, Yukun and Zhang, Zhengyu and Yang, Jiyan and Wen, Ellie and Chordia, Sagar and Chen, Wenlin and Huang, Qin , booktitle =
-
[15]
and Tong, Hanghang and Yang, Jiyan , booktitle =
Zeng, Zhichen and Liu, Xiaolong and Hang, Mengyue and Liu, Xiaoyi and Zhou, Qinghai and Yang, Chaofei and Liu, Yiqun and Ruan, Yichen and Chen, Laming and Chen, Yuxin and Hao, Yujia and Xu, Jiaqi and Nie, Jade and Liu, Xi and Zhang, Buyun and Wen, Wei and Yuan, Siyang and Zhan...
-
[16]
Zhang, Buyun and Luo, Liang and Liu, Xi and Li, Jay and Chen, Zeliang and Zhang, Weilin and Wei, Xiaohan and Hao, Yuchen and Tsang, Michael and Wang, Wenjun and Liu, Yang and Li, Huayu and Badr, Yasmine and Park, Jongsoo and Yang, Jiyan and Mudigere, Dheevatsa and Wen, Ellie ,...
-
[17]
and Kaiser, Lukasz and Polosukhin, Illia , booktitle =
Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N. and Kaiser, Lukasz and Polosukhin, Illia , booktitle =
-
[18]
2025 , archivePrefix =
Deng, Jiaxin and Wang, Shiyao and Cai, Kuo and Ren, Lejian and Hu, Qigen and Ding, Weifeng and Luo, Qiang and Zhou, Guorui , journal =. 2025 , archivePrefix =
2025
-
[19]
Zhai, Jiaqi and Liao, Lucy and Liu, Xing and Wang, Yueming and Li, Rui and Cao, Xuan and Gao, Leon and Gong, Zhaojie and Gu, Fangda and He, Jiayuan and Lu, Yinghai and Shi, Yu , booktitle =
-
[20]
Naumov, Maxim and Mudigere, Dheevatsa and Shi, Hao-Jun Michael and Huang, Jianyu and Sundaraman, Narayanan and Park, Jongsoo and Wang, Xiaodong and Gupta, Udit and Wu, Carole-Jean and Azzolini, Alisson G. and Dzhulgakov, Dmytro and Mallevich, Andrey and Cherniavskii, Ilia and ...
2019
-
[21]
and Jain, Sagar and Lin, Dong and Hong, Lichan and Chi, Ed H
Wang, Ruoxi and Shivanna, Rakesh and Cheng, Derek Z. and Jain, Sagar and Lin, Dong and Hong, Lichan and Chi, Ed H. , booktitle =
-
[22]
De Cao, Nicola and Izacard, Gautier and Riedel, Sebastian and Petroni, Fabio , booktitle =
-
[23]
and Dehghani, Mostafa and Ni, Jianmo and Bahri, Dara and Mehta, Harsh and Qin, Zhen and Hui, Kai and Zhao, Zhe and Gupta, Jai and Schuster, Tal and Cohen, William W
Tay, Yi and Tran, Vinh Q. and Dehghani, Mostafa and Ni, Jianmo and Bahri, Dara and Mehta, Harsh and Qin, Zhen and Hui, Kai and Zhao, Zhe and Gupta, Jai and Schuster, Tal and Cohen, William W. and Metzler, Donald , booktitle =
-
[24]
Zheng, Bowen and Hou, Yupeng and Lu, Hongyu and Chen, Yu and Zhao, Wayne Xin and Chen, Ming and Wen, Ji-Rong , booktitle =
-
[25]
Lu, Ximing and Welleck, Sean and Hessel, Jack and Jiang, Liwei and Qin, Lianhui and West, Peter and Ammanabrolu, Prithviraj and Choi, Yejin , booktitle =
-
[26]
Poesia, Gabriel and Polozov, Oleksandr and Le, Vu and Tiwari, Ashish and Soares, Gustavo and Meek, Christopher and Gulwani, Sumit , booktitle =
-
[27]
Ye, Haotian and Jain, Himanshu and You, Chong and Suresh, Ananda Theertha and Lin, Haowei and Zou, James and Yu, Felix , booktitle =
-
[28]
2026 , publisher =
Gang Liao and Hongsen Qin and Ying Wang and Alicia Golden and Michael Kuchnik and Yavuz Yetim and Ruichao Xiao and Jia Jiunn Ang and Chunli Fu and Yihan He and Samuel Hsia and Zewei Jiang and Roman Levenstein and Dianshi Li and Liyuan Li and Ajit Mathews and Varna Puvvada and ...
2026
-
[29]
and Ren, Manman and Wang, Lei and Nay, Shane and Kanuparthy, Partha and Pan, Zaifeng and Hu, Zhengding and Ding, Yufei , journal =
Guan, Yue and Yu, Hongtao and Chen, Peng and Shi, Daohang and Manivannan, Karthik and Riasanovsky, Nicholas J. and Ren, Manman and Wang, Lei and Nay, Shane and Kanuparthy, Partha and Pan, Zaifeng and Hu, Zhengding and Ding, Yufei , journal =. 2026 , archivePrefix =
2026
-
[30]
Deep Neural Networks for YouTube Recommendations
Covington, P., Adams, J., and Sargin, E. Deep Neural Networks for YouTube Recommendations . In Proceedings of the 10th ACM Conference on Recommender Systems, pp.\ 191--198. ACM, September 2016
2016
-
[31]
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Dao, T. FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning . In International Conference on Learning Representations, May 2024
2024
-
[32]
Y., Ermon, S., Rudra, A., and R \'e , C
Dao, T., Fu, D. Y., Ermon, S., Rudra, A., and R \'e , C. FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness . In Proceedings of the 36th International Conference on Neural Information Processing Systems, pp.\ 16344--16359, 2022
2022
-
[33]
Autoregressive Entity Retrieval
De Cao, N., Izacard, G., Riedel, S., and Petroni, F. Autoregressive Entity Retrieval . In International Conference on Learning Representations, 2021
2021
-
[34]
OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment
Deng, J., Wang, S., Cai, K., Ren, L., Hu, Q., Ding, W., Luo, Q., and Zhou, G. OneRec: Unifying Retrieve and Rank with Generative Recommender and Iterative Preference Alignment . arXiv preprint arXiv:2502.18965, 2025
2025 arXiv
-
[35]
Gallagher, L., Chen, R.-C., Blanco, R., and Culpepper, J. S. Joint Optimization of Cascade Ranking Models . In Proceedings of the Twelfth ACM International Conference on Web Search and Data Mining, pp.\ 15--23. ACM, January 2019
2019
-
[36]
J., Ren, M., Wang, L., Nay, S., Kanuparthy, P., Pan, Z., Hu, Z., and Ding, Y
Guan, Y., Yu, H., Chen, P., Shi, D., Manivannan, K., Riasanovsky, N. J., Ren, M., Wang, L., Nay, S., Kanuparthy, P., Pan, Z., Hu, Z., and Ding, Y. TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-Scale Production Environments . arXiv preprint arXiv:2605.10905, 2026
2026 arXiv
-
[37]
PLUM: Adapting Pre-trained Language Models for Industrial-scale Generative Recommendations
He, R., Heldt, L., Hong, L., Keshavan, R., Mao, S., Mehta, N., Su, Z., Tsai, A., Wang, Y., Wang, S.-C., Yi, X., Baugher, L., Cakici, B., Chi, E., Goodrow, C., Han, N., Ma, H., Rosales, R., Van Soest, A., Tandon, D., Wu, S.-L., Yang, W., and Zheng, Y. PLUM: Adapting Pre-trained...
2026
-
[38]
H., Gonzalez, J
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., and Stoica, I. Efficient Memory Management for Large Language Model Serving with PagedAttention . In Proceedings of the 29th Symposium on Operating Systems Principles, pp.\ 611--626. ACM...
2023
-
[39]
Meta's Generative Ads Model (GEM): The Central Brain Accelerating Ads Recommendation AI Innovation
Li, H., Liu, X., Nie, J., Wen, E., Yang, C., Yang, J., Yu, N., Beg, H., Arditi, G., and Bhatia, N. Meta's Generative Ads Model (GEM): The Central Brain Accelerating Ads Recommendation AI Innovation . Meta Engineering Blog https://engineering.fb.com/2025/11/10/ml-applications/m...
2025
-
[40]
J., Fu, C., He, Y., Hsia, S., Jiang, Z., Levenstein, R., Li, D., Li, L., Mathews, A., Puvvada, V., Shi, F., Yan, N., Yu, X., Pashkevich, U., Steiner, M., Wu, C.-J., and Liu, G
Liao, G., Qin, H., Wang, Y., Golden, A., Kuchnik, M., Yetim, Y., Xiao, R., Ang, J. J., Fu, C., He, Y., Hsia, S., Jiang, Z., Levenstein, R., Li, D., Li, L., Mathews, A., Puvvada, V., Shi, F., Yan, N., Yu, X., Pashkevich, U., Steiner, M., Wu, C.-J., and Liu, G. KernelEvolve: Sca...
2026
-
[41]
NeuroLogic Decoding: (Un)supervised Neural Text Generation with Predicate Logic Constraints
Lu, X., Welleck, S., Hessel, J., Jiang, L., Qin, L., West, P., Ammanabrolu, P., and Choi, Y. NeuroLogic Decoding: (Un)supervised Neural Text Generation with Predicate Logic Constraints . In Proceedings of the 2021 Conference of the North American Chapter of the Association for...
2021
-
[42]
M., Huang, J., Sundaraman, N., Park, J., Wang, X., Gupta, U., Wu, C.-J., Azzolini, A
Naumov, M., Mudigere, D., Shi, H.-J. M., Huang, J., Sundaraman, N., Park, J., Wang, X., Gupta, U., Wu, C.-J., Azzolini, A. G., Dzhulgakov, D., Mallevich, A., Cherniavskii, I., Lu, Y., Krishnamoorthi, R., Yu, A., Kondratenko, V., Pereira, S., Chen, X., Chen, W., Rao, V., Jia, B...
1906 arXiv
-
[43]
NVIDIA GH200 Grace Hopper Superchip
NVIDIA . NVIDIA GH200 Grace Hopper Superchip . https://www.nvidia.com/en-us/data-center/grace-hopper-superchip/, 2026
2026
-
[44]
Synchromesh: Reliable Code Generation from Pre-trained Language Models
Poesia, G., Polozov, O., Le, V., Tiwari, A., Soares, G., Meek, C., and Gulwani, S. Synchromesh: Reliable Code Generation from Pre-trained Language Models . In International Conference on Learning Representations, 2022
2022
-
[45]
Q., Samost, J., Kula, M., Chi, E
Rajput, S., Mehta, N., Singh, A., Hulikal Keshavan, R., Vu, T., Heldt, L., Hong, L., Tay, Y., Tran, V. Q., Samost, J., Kula, M., Chi, E. H., and Sathiamoorthy, M. Recommender Systems with Generative Retrieval . In Advances in Neural Information Processing Systems, volume 36, p...
2023
-
[46]
FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-Precision
Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., and Dao, T. FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-Precision . In Proceedings of the 38th International Conference on Neural Information Processing Systems, pp.\ 68658--68685, 2024
2024
-
[47]
H., and Han, N
Su, Z., Katsman, I., Wang, Y., He, R., Heldt, L., Keshavan, R., Wang, S.-C., Yi, X., Gao, M., Dalal, O., Hong, L., Chi, E. H., and Han, N. Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators . In Proceedings of the 32nd ACM S...
2026
-
[48]
Meta Andromeda: Supercharging Advantage+ Automation with the Next-Gen Personalized Ads Retrieval Engine
Sun, C., Yu, N., Lu, H., Wang, L., Pu, Y., Liu, G., Bhatia, N., and Musumeci, G. Meta Andromeda: Supercharging Advantage+ Automation with the Next-Gen Personalized Ads Retrieval Engine . Meta Engineering Blog https://engineering.fb.com/2024/12/02/production-engineering/meta-an...
2024
-
[49]
Q., Dehghani, M., Ni, J., Bahri, D., Mehta, H., Qin, Z., Hui, K., Zhao, Z., Gupta, J., Schuster, T., Cohen, W
Tay, Y., Tran, V. Q., Dehghani, M., Ni, J., Bahri, D., Mehta, H., Qin, Z., Hui, K., Zhao, Z., Gupta, J., Schuster, T., Cohen, W. W., and Metzler, D. Transformer Memory as a Differentiable Search Index . In Advances in Neural Information Processing Systems, volume 35, pp.\ 2183...
2022
-
[50]
N., Kaiser, L., and Polosukhin, I
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. Attention is All You Need . In Advances in Neural Information Processing Systems, volume 30, pp.\ 5998--6008, 2017
2017
-
[51]
Z., Jain, S., Lin, D., Hong, L., and Chi, E
Wang, R., Shivanna, R., Cheng, D. Z., Jain, S., Lin, D., Hong, L., and Chi, E. H. DCN V2: Improved Deep & Cross Network and Practical Lessons for Web-scale Learning to Rank Systems . In Proceedings of the Web Conference 2021, pp.\ 1785--1797. ACM, April 2021
2021
-
[52]
Towards the Better Ranking Consistency: A Multi-task Learning Framework for Early Stage Ads Ranking
Wang, X., Jin, Q., Huang, S., Zhang, M., Liu, X., Zhao, Z., Chen, Y., Zhang, Z., Yang, J., Wen, E., Chordia, S., Chen, W., and Huang, Q. Towards the Better Ranking Consistency: A Multi-task Learning Framework for Early Stage Ads Ranking . In Proceedings of the AdKDD Workshop, 2023
2023
-
[53]
T., Lin, H., Zou, J., and Yu, F
Ye, H., Jain, H., You, C., Suresh, A. T., Lin, H., Zou, J., and Yu, F. Efficient and Asymptotically Unbiased Constrained Decoding for Large Language Models . In International Conference on Artificial Intelligence and Statistics, volume 258, pp.\ 4483--4491. PMLR, April 2025 a
2025
-
[54]
FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving
Ye, Z., Chen, L., Lai, R., Lin, W., Zhang, Y., Wang, S., Chen, T., Kasikci, B., Grover, V., Krishnamurthy, A., and Ceze, L. FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving . In Proceedings of the 8th Annual Conference on Machine Learning and S...
2025
-
[55]
FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling
Zadouri, T., Hoehnerbach, M., Shah, J., Liu, T., Thakkar, V., and Dao, T. FlashAttention-4: Algorithm and Kernel Pipelining Co-Design for Asymmetric Hardware Scaling . In Proceedings of the 9th Annual Conference on Machine Learning and Systems, May 2026
2026
-
[56]
S., Tong, H., and Yang, J
Zeng, Z., Liu, X., Hang, M., Liu, X., Zhou, Q., Yang, C., Liu, Y., Ruan, Y., Chen, L., Chen, Y., Hao, Y., Xu, J., Nie, J., Liu, X., Zhang, B., Wen, W., Yuan, S., Zhang, X., Wang, K., Chen, W.-Y., Han, Y., Li, H., Yang, C., Long, B., Yu, P. S., Tong, H., and Yang, J. InterForme...
2025
-
[57]
Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations
Zhai, J., Liao, L., Liu, X., Wang, Y., Li, R., Cao, X., Gao, L., Gong, Z., Gu, F., He, J., Lu, Y., and Shi, Y. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations . In Proceedings of the 41st International Conference on Mac...
2024
-
[58]
DHEN: A Deep and Hierarchical Ensemble Network for Large-Scale Click-Through Rate Prediction
Zhang, B., Luo, L., Liu, X., Li, J., Chen, Z., Zhang, W., Wei, X., Hao, Y., Tsang, M., Wang, W., Liu, Y., Li, H., Badr, Y., Park, J., Yang, J., Mudigere, D., and Wen, E. DHEN: A Deep and Hierarchical Ensemble Network for Large-Scale Click-Through Rate Prediction . In Workshop ...
2022
-
[59]
X., Chen, M., and Wen, J.-R
Zheng, B., Hou, Y., Lu, H., Chen, Y., Zhao, W. X., Chen, M., and Wen, J.-R. Adapting Large Language Models by Integrating Collaborative Semantics for Recommendation . In 2024 IEEE 40th International Conference on Data Engineering (ICDE), pp.\ 1435--1448. IEEE, May 2024
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.