{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2025:SDI5XCOC4L5OAFROSGD55SLAAF","short_pith_number":"pith:SDI5XCOC","schema_version":"1.0","canonical_sha256":"90d1db89c2e2fae0162e9187dec9600147e8b3b44ed334d67ec76c97004d3729","source":{"kind":"arxiv","id":"2506.17298","version":1},"attestation_state":"computed","paper":{"title":"Mercury: Ultra-Fast Language Models Based on Diffusion","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"Diffusion LLMs generate code at over 1100 tokens per second while matching frontier quality.","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Aditya Grover, Akash Palrecha, Eric Wang, Harshit Varma, Inception Labs, Samar Khanna, Sawyer Birnbaum, Shufan Li, Siddhant Kharbanda, Stefano Ermon, Volodymyr Kuleshov, Yanis Miraoui, Ziyang Luo","submitted_at":"2025-06-17T17:06:18Z","abstract_excerpt":"We present Mercury, a new generation of commercial-scale large language models (LLMs) based on diffusion. These models are parameterized via the Transformer architecture and trained to predict multiple tokens in parallel. In this report, we detail Mercury Coder, our first set of diffusion LLMs designed for coding applications. Currently, Mercury Coder comes in two sizes: Mini and Small. These models set a new state-of-the-art on the speed-quality frontier. Based on independent evaluations conducted by Artificial Analysis, Mercury Coder Mini and Mercury Coder Small achieve state-of-the-art thro"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"2506.17298","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","primary_cat":"cs.CL","submitted_at":"2025-06-17T17:06:18Z","cross_cats_sorted":["cs.AI","cs.LG"],"title_canon_sha256":"0b90f499788fdc9bde41d138574e56eb58d75f2a1b9562e557221998c6185b6f","abstract_canon_sha256":"42d05162ad40c8417f9651c838467a46967600527416aa5176c6ec3f44016011"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-05-17T23:38:15.453770Z","signature_b64":"Ob5bmBinHdHkEeHxx8gBte02uQXg5YqJDK0b4rwlTwTZ3xB0O3CMWnRgm4FDU5VRDbYpqG8y8ONgZEwBE6njDw==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"90d1db89c2e2fae0162e9187dec9600147e8b3b44ed334d67ec76c97004d3729","last_reissued_at":"2026-05-17T23:38:15.453003Z","signature_status":"signed_v1","first_computed_at":"2026-05-17T23:38:15.453003Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Mercury: Ultra-Fast Language Models Based on Diffusion","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"Diffusion LLMs generate code at over 1100 tokens per second while matching frontier quality.","cross_cats":["cs.AI","cs.LG"],"primary_cat":"cs.CL","authors_text":"Aditya Grover, Akash Palrecha, Eric Wang, Harshit Varma, Inception Labs, Samar Khanna, Sawyer Birnbaum, Shufan Li, Siddhant Kharbanda, Stefano Ermon, Volodymyr Kuleshov, Yanis Miraoui, Ziyang Luo","submitted_at":"2025-06-17T17:06:18Z","abstract_excerpt":"We present Mercury, a new generation of commercial-scale large language models (LLMs) based on diffusion. These models are parameterized via the Transformer architecture and trained to predict multiple tokens in parallel. In this report, we detail Mercury Coder, our first set of diffusion LLMs designed for coding applications. Currently, Mercury Coder comes in two sizes: Mini and Small. These models set a new state-of-the-art on the speed-quality frontier. Based on independent evaluations conducted by Artificial Analysis, Mercury Coder Mini and Mercury Coder Small achieve state-of-the-art thro"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Mercury Coder Mini and Small achieve state-of-the-art throughputs of 1109 tokens/sec and 737 tokens/sec on NVIDIA H100 GPUs and outperform speed-optimized frontier models by up to 10x on average while maintaining comparable quality.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"The assumption that independent evaluations by Artificial Analysis and Copilot Arena rankings accurately measure both speed and quality in a way that generalizes beyond the tested benchmarks and real-world developer use.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Mercury Coder diffusion LLMs achieve throughputs of 1109 and 737 tokens per second on H100 GPUs, up to 10x faster than frontier models with comparable quality.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Diffusion LLMs generate code at over 1100 tokens per second while matching frontier quality.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"aaf35451ec928680a9851599a036a3a4176bdc921a77895a29919a8ab0e143ce"},"source":{"id":"2506.17298","kind":"arxiv","version":1},"verdict":{"id":"8e07e1f6-e29f-4422-a9ed-815abe6362d5","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-17T02:00:43.664501Z","strongest_claim":"Mercury Coder Mini and Small achieve state-of-the-art throughputs of 1109 tokens/sec and 737 tokens/sec on NVIDIA H100 GPUs and outperform speed-optimized frontier models by up to 10x on average while maintaining comparable quality.","one_line_summary":"Mercury Coder diffusion LLMs achieve throughputs of 1109 and 737 tokens per second on H100 GPUs, up to 10x faster than frontier models with comparable quality.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"The assumption that independent evaluations by Artificial Analysis and Copilot Arena rankings accurately measure both speed and quality in a way that generalizes beyond the tested benchmarks and real-world developer use.","pith_extraction_headline":"Diffusion LLMs generate code at over 1100 tokens per second while matching frontier quality."},"references":{"count":41,"sample":[{"doi":"","year":null,"title":"URLhttps://api.semanticscholar","work_id":"5cd6d826-7f9a-43eb-b3c8-c8e8a02f6495","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2024,"title":"Top latest ai code generator statistics and trends in 2024, 2024","work_id":"c769cabb-ea69-4f3b-9161-de2305770d0f","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2023,"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","ref_index":3,"cited_arxiv_id":"2303.08774","is_internal_anchor":true},{"doi":"","year":2021,"title":"Structured denoising diffusion models in discrete state-spaces.Advances in Neural Infor- mation Processing Systems, 34:17981–17993, 2021","work_id":"7537045d-579b-4a5a-a11a-57bab1056159","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2021,"title":"Program Synthesis with Large Language Models","work_id":"fd241a05-03b9-4de2-9588-9d77ce176125","ref_index":5,"cited_arxiv_id":"2108.07732","is_internal_anchor":true}],"resolved_work":41,"snapshot_sha256":"8507d88fa94c3c0e220d80b426a54bfd761d67eb3500e7f2656a190c1a34c1d0","internal_anchors":15},"formal_canon":{"evidence_count":3,"snapshot_sha256":"7452bfbcac4233280fa4d20edc11419b66e40010f913dbb7f27aa86c7a9a8d25"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2506.17298","created_at":"2026-05-17T23:38:15.453137+00:00"},{"alias_kind":"arxiv_version","alias_value":"2506.17298v1","created_at":"2026-05-17T23:38:15.453137+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2506.17298","created_at":"2026-05-17T23:38:15.453137+00:00"},{"alias_kind":"pith_short_12","alias_value":"SDI5XCOC4L5O","created_at":"2026-05-18T12:33:37.589309+00:00"},{"alias_kind":"pith_short_16","alias_value":"SDI5XCOC4L5OAFRO","created_at":"2026-05-18T12:33:37.589309+00:00"},{"alias_kind":"pith_short_8","alias_value":"SDI5XCOC","created_at":"2026-05-18T12:33:37.589309+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":22,"internal_anchor_count":22,"sample":[{"citing_arxiv_id":"2508.19982","citing_title":"Diffusion Language Models Know the Answer Before Decoding","ref_index":12,"is_internal_anchor":true},{"citing_arxiv_id":"2509.20624","citing_title":"FS-DFM: Fast and Accurate Long Text Generation with Few-Step Diffusion Language Models","ref_index":4,"is_internal_anchor":true},{"citing_arxiv_id":"2510.18165","citing_title":"Saber: An Efficient Sampling with Adaptive Acceleration and Backtracking Enhanced Remasking for Diffusion Language Model","ref_index":14,"is_internal_anchor":true},{"citing_arxiv_id":"2509.08827","citing_title":"A Survey of Reinforcement Learning for Large Reasoning Models","ref_index":259,"is_internal_anchor":true},{"citing_arxiv_id":"2512.14067","citing_title":"Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed","ref_index":31,"is_internal_anchor":true},{"citing_arxiv_id":"2601.18681","citing_title":"ART for Diffusion Sampling: A Reinforcement Learning Approach to Timestep Schedule","ref_index":7,"is_internal_anchor":true},{"citing_arxiv_id":"2602.16813","citing_title":"Flow Map Language Models: One-step Language Modeling via Continuous Denoising","ref_index":4,"is_internal_anchor":true},{"citing_arxiv_id":"2604.08557","citing_title":"Re-Mask and Redirect: Exploiting Denoising Irreversibility in Diffusion Language Models","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2604.08564","citing_title":"Attention-Based Sampler for Diffusion Language Models","ref_index":6,"is_internal_anchor":true},{"citing_arxiv_id":"2605.11726","citing_title":"Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models","ref_index":26,"is_internal_anchor":true},{"citing_arxiv_id":"2605.11726","citing_title":"Block-R1: Rethinking the Role of Block Size in Multi-domain Reinforcement Learning for Diffusion Large Language Models","ref_index":26,"is_internal_anchor":true},{"citing_arxiv_id":"2605.11854","citing_title":"Self-Distilled Trajectory-Aware Boltzmann Modeling: Bridging the Training-Inference Discrepancy in Diffusion Language Models","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"2605.10518","citing_title":"Infinite Mask Diffusion for Few-Step Distillation","ref_index":7,"is_internal_anchor":true},{"citing_arxiv_id":"2604.26985","citing_title":"Simple Self-Conditioning Adaptation for Masked Diffusion Models","ref_index":7,"is_internal_anchor":true},{"citing_arxiv_id":"2604.19856","citing_title":"ChipCraftBrain: Validation-First RTL Generation via Multi-Agent Orchestration","ref_index":20,"is_internal_anchor":true},{"citing_arxiv_id":"2502.09992","citing_title":"Large Language Diffusion Models","ref_index":74,"is_internal_anchor":true},{"citing_arxiv_id":"2604.13413","citing_title":"Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models","ref_index":14,"is_internal_anchor":true},{"citing_arxiv_id":"2604.14001","citing_title":"Diffusion Language Models for Speech Recognition","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"2604.18738","citing_title":"Remask, Don't Replace: Token-to-Mask Refinement in Diffusion Large Language Models","ref_index":10,"is_internal_anchor":true},{"citing_arxiv_id":"2604.15750","citing_title":"DepCap: Adaptive Block-Wise Parallel Decoding for Efficient Diffusion LM Inference","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"2604.18471","citing_title":"NI Sampling: Accelerating Discrete Diffusion Sampling by Token Order Optimization","ref_index":8,"is_internal_anchor":true},{"citing_arxiv_id":"2604.20079","citing_title":"On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks","ref_index":18,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":3,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/SDI5XCOC4L5OAFROSGD55SLAAF","json":"https://pith.science/pith/SDI5XCOC4L5OAFROSGD55SLAAF.json","graph_json":"https://pith.science/api/pith-number/SDI5XCOC4L5OAFROSGD55SLAAF/graph.json","events_json":"https://pith.science/api/pith-number/SDI5XCOC4L5OAFROSGD55SLAAF/events.json","paper":"https://pith.science/paper/SDI5XCOC"},"agent_actions":{"view_html":"https://pith.science/pith/SDI5XCOC4L5OAFROSGD55SLAAF","download_json":"https://pith.science/pith/SDI5XCOC4L5OAFROSGD55SLAAF.json","view_paper":"https://pith.science/paper/SDI5XCOC","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2506.17298&json=true","fetch_graph":"https://pith.science/api/pith-number/SDI5XCOC4L5OAFROSGD55SLAAF/graph.json","fetch_events":"https://pith.science/api/pith-number/SDI5XCOC4L5OAFROSGD55SLAAF/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/SDI5XCOC4L5OAFROSGD55SLAAF/action/timestamp_anchor","attest_storage":"https://pith.science/pith/SDI5XCOC4L5OAFROSGD55SLAAF/action/storage_attestation","attest_author":"https://pith.science/pith/SDI5XCOC4L5OAFROSGD55SLAAF/action/author_attestation","sign_citation":"https://pith.science/pith/SDI5XCOC4L5OAFROSGD55SLAAF/action/citation_signature","submit_replication":"https://pith.science/pith/SDI5XCOC4L5OAFROSGD55SLAAF/action/replication_record"}},"created_at":"2026-05-17T23:38:15.453137+00:00","updated_at":"2026-05-17T23:38:15.453137+00:00"}