REVIEW 3 major objections 4 minor 57 references
Tracezip: Efficient Distributed Tracing via Trace Compression
T0 review · 3 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash
Pith's one-line read Tracezip compresses trace spans on the fly, cutting tracing overhead by 10–45% on top of gzip.
desk verdict A clever online trace-compression system whose end-to-end savings are not yet proven, because the evaluation omits SRT sync bytes from the wire count. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The Span Retrieval Tree (SRT) is a prefix tree whose paths spell out the key-value pairs common to all spans of a given span name. Each non-leaf node holds a shared key-value pair, while the leaf node stores the names of fields whose values vary too much to be shared, and a time base at the root lets timestamps be sent as small offsets. The SRT is restructured by reordering keys by their number of distinct values, compressed by replacing repeated keys and values with short identifiers, and synchronized with the backend through a differential update so only changed paths are transmitted. This structure is what converts a stream of verbose spans into compact identifiers plus local values, and back again.
What would settle it
Run Tracezip against a corpus of traces from a production system whose instrumentation changes over time or whose spans include optional attributes, and see whether every span is reconstructed byte-for-byte. If any span whose name already exists in the SRT arrives with a new key or a nested structure that differs from the stored path, reconstruction will fail or silently drop the new field, which would falsify the structural-locality premise.
Extended reading notes
Core claim
Tracezip's central claim is that the redundancy distributed across trace spans is large enough to be exploited at generation time, without waiting for backend-side batch compression. The authors introduce the Span Retrieval Tree (SRT), a prefix-tree data structure that records, for each span name, the key-value pairs shared by many spans as a single path. A span is transmitted as the identifier of its path plus the values of the few 'local' fields that differ from span to span, such as identifiers and timestamps, and the backend reconstructs the original span by looking up the path and reattaching the local values. Because the SRT and its companion dictionary are synchronized incrementally, the compression structures themselves cost only a few megabytes. The paper reports compression ratios of roughly 4 to 6 for Tracezip alone, and 10%–45% additional size reduction when combined with general-purpose compressors, with full fidelity of the reconstructed traces.
Load-bearing premise
The scheme assumes that every span with the same name has the same set of fields, so a single shared path can stand for all of them; if real spans vary their attributes from one request to the next, path lookup fails or reconstruction is wrong.
Editorial extensions
If this is right
- Tracing can cover all requests rather than only a sampled subset, because the added transmission cost of full tracing is sharply reduced.
- Tracezip composes with existing sampling and log-compression systems, so operators can layer it on top of current tracing stacks without changing their APIs.
- On the Train Ticket benchmark, collection throughput rises from about 14 MB/s to nearly 110 MB/s, implying that the bottleneck shifts from network transfer to span generation.
- The memory footprint stays at the single-digit-megabyte scale even with a high threshold for what counts as a shared field, so the compression structures fit in typical service containers.
- When combined with lzma on 26 GB of production traces, the data shrink from 3.07 GB to 1.91 GB, an improvement of about 38%.
- When combined with lzma on 26 GB of production traces, the data shrink from 3.07 GB to 1.91 GB, an improvement of about 38%.
Reading between the lines
- Because Tracezip's gain comes from global redundancy rather than the local windows used by gzip, its advantage should grow with trace volume and diversity; a natural test is to measure compression gain as a function of the number of traces within a fixed time window.
- The structural-locality assumption suggests an extension: instead of one SRT per service instance, grouping spans by instrumentation point or operation template could make the scheme robust to schema drift, since each group's paths would remain stable even when the overall span schema evolves.
- If adoption spreads, the SRT itself becomes a compact digest of a service's 'normal' behavior, so Tracezip could double as a cheap detector of new or changed span structures, flagging spans that fail to match any existing path.
- The throughput result implies that serialization becomes the next bottleneck; integrating span compression directly into serialization formats such as Protobuf could yield further gains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Tracezip, an online compression layer for distributed tracing. Tracezip maintains a Span Retrieval Tree (SRT) and a dictionary at the service side, encodes each span as a path identifier plus a small set of local-field values, and synchronizes the SRT and dictionary to the backend through differential updates. The authors implement the system inside the OpenTelemetry Collector and evaluate it on Train Ticket, six application components (gRPC, Kafka, Servlet, MySQL, Redis, MongoDB), and Alibaba production traces, reporting compression-ratio improvements over gzip, bzip2, and lzma as well as a roughly eight-fold throughput gain.
Significance. If the claimed gains remain valid when all bytes actually sent to the backend are counted, Tracezip is a useful contribution: it targets the transmission and ingestion cost of full-fidelity tracing, is orthogonal to sampling, and ships a working OpenTelemetry implementation with public code and data. The evaluation honestly compares against external general-purpose compressors, and the threshold parameter ψ is treated as a swept design knob rather than fitted to the data, so there is no circularity in the main compression measurements. However, the load-bearing performance claim depends on a wire-accounting metric that the current evaluation does not fully implement, and the central structural-locality assumption is not validated on the production dataset; both need to be addressed before the headline results can be accepted.
major comments (3)
- [Section 5.1.2, Table 2, Figure 6, Table 3] The CR metric defined in Section 5.1.2 as CR = Original File Size / Compressed File Size is computed on the compressed span stream only and omits the SRT and dictionary synchronization bytes that Section 4.2 explicitly says are transmitted to the backend. Since the paper's central claim is reduced trace transmission overhead (Section 3.1), the evaluation should count total bytes on the wire, i.e., compressed spans plus all differential update payloads. For a concrete scale, the Train Ticket row reports raw 21.0 MB and Tracezip 5.19 MB; including even a few MB of cumulative sync traffic would reduce the CR from 4.05 to well below that, and the stated 10–45% improvements would shrink correspondingly. Please report the cumulative synchronization-byte volume per dataset and recompute all CR values and the throughput numbers in Table 3 with this overhead included.
- [Section 3.2, Algorithm 1] The structural-locality assumption — that all spans sharing a span Name have the identical set of keys, differing only in values — is load-bearing for path lookup and lossless reconstruction, but it is not validated on the Alibaba production dataset or on the open-source systems. Production spans can exhibit optional attributes, nested structures that vary by occurrence, or schema drift after instrumentation or framework upgrades; under such conditions Algorithm 1's traversal (lines 9–12) has no defined behavior when a key expected at a given depth is missing from the span. Please measure, per span Name, the fraction of spans whose key set deviates from the registered SRT path, and specify the fallback behavior on mismatch (e.g., creating a new branch, transmitting the span raw, or reporting an error).
- [Section 5.2, Figure 7] The headline results in Table 2 and Figure 6 do not state the configuration used for the threshold ψ, the SRT depth limit, the time_base reset period, or the SRT memory cap, even though Section 5.3.1 shows that ψ materially changes compression effectiveness. Please state the exact parameter settings for each experiment and include a sensitivity analysis for the Table 2 systems; without this, the claimed 'around 10%–45%' gains are not reproducible and cannot be attributed to a fixed Tracezip configuration.
minor comments (4)
- [Table 2, Kafka row] The Kafka/Tracezip(lzma) entry reports a compressed size of 0.012 MB, which would imply a compression ratio of roughly 205.8 rather than the listed 20.58 and is inconsistent with the stated 9.6% improvement over lzma; please correct this apparent typo (the value should likely be around 0.116–0.120 MB).
- [Throughout] There are several typos and small presentation issues: Section 2.2 has 'capabiltiy', Section 3.3 has 'readibility', Figure 2's axis label says 'Radio' instead of 'Ratio', Figure 5's caption says 'SFT: time_base' instead of 'SRT', and Section 5.3.2 says 'the results are present' and 'the time token' instead of 'the results are presented' and 'the time taken'.
- [Section 5.1–5.3] No repeated runs, error bars, or variance information are reported for the CR or throughput measurements; since the paper's conclusions are comparative percentage improvements, a statement about run-to-run variability would strengthen the evaluation.
- [Section 2.2 vs. Section 5.1] The redundancy study reports collecting more than 40 GB of trace data, but the open-source evaluation datasets in Table 2 are on the order of tens of MB; please clarify the relationship between the 40 GB figure and the datasets used for the compression experiments.
Circularity Check
No significant circularity: Tracezip's compression gains are empirical measurements against external baselines, and its design choices are stated assumptions or swept parameters, not fitted to the results they are used to explain.
full rationale
Tracezip's central claim is that its SRT-based compression reduces trace transmission size and improves throughput. The derivation chain is an engineering construction (Algorithm 1 plus optimizations), not a mathematical derivation from fitted parameters. The compression ratio is defined in Section 5.1.2 as CR = Original File Size / Compressed File Size and then independently measured on Train Ticket, six application components, and Alibaba production traces against external baselines gzip, bzip2, and lzma. The psi threshold is an explicit design knob swept across values 1, 10, 100, and 1,000 in Section 5.3.1, not a parameter fitted to make a target result come out. The structural-locality assumption in Section 3.2 is stated as an assumption and is empirically falsifiable (schema drift would break path lookup); it is not imported from a self-citation. Self-citations in the reference list, such as TraceMesh [4] and incident-management work [5], appear as related-work context and are not load-bearing for the SRT compression mechanism. The skeptic's concern that the CR metric omits SRT/dictionary synchronization bytes is an evaluation-validity issue about whether the headline 10-45% gain holds end-to-end, not circularity: the metric does not equal its input by construction, and the paper does not rename a fitted quantity as a prediction. No step in the paper's derivation reduces to its own inputs by definition or by self-citation, so the appropriate score is 0.
Assumptions & free parameters
free parameters (4)
- threshold ψ (psi) =
unspecified, swept from 1 to 10,000
- SRT depth limit =
2 (default)
- time_base reset period =
1 second
- SRT memory cap =
5 MB
assumptions (3)
- domain assumption Spans with the same span Name have identical key structure
- domain assumption The backend always receives SRT updates before the compressed spans that depend on them
- domain assumption A single production trace dataset (Alibaba) is representative of diverse cloud workloads
Cite this review
Pith. "Pith review of Tracezip: Efficient Distributed Tracing via Trace Compression." pith.science (2026). https://pith.science/paper/DNN5FJH6
@misc{pith2026250206318,
author = {Pith},
title = {Pith review of: Tracezip: Efficient Distributed Tracing via Trace Compression},
year = {2026},
howpublished = {\url{https://pith.science/paper/DNN5FJH6}},
note = {Machine review of arXiv:2502.06318}
}
read the original abstract
Distributed tracing serves as a fundamental building block in the monitoring and testing of cloud service systems. To reduce computational and storage overheads, the de facto practice is to capture fewer traces via sampling. However, existing work faces a trade-off between the completeness of tracing and system overhead. On one hand, head-based sampling indiscriminately selects requests to trace when they enter the system, which may miss critical events. On the other hand, tail-based sampling first captures all requests and then selectively persists the edge-case traces, which entails the overheads related to trace collection and ingestion. Taking a different path, we propose Tracezip in this paper to enhance the efficiency of distributed tracing via trace compression. Our key insight is that there exists significant redundancy among traces, which results in repetitive transmission of identical data between services and the backend. We design a new data structure named Span Retrieval Tree (SRT) that continuously encapsulates such redundancy at the service side and transforms trace spans into a lightweight form. At the backend, the complete traces can be seamlessly reconstructed by retrieving the common data that are already delivered by previous spans. Tracezip includes a series of strategies to optimize the structure of SRT and a differential update mechanism to efficiently synchronize SRT between services and the backend. Our evaluation on microservices benchmarks, popular cloud service systems, and production trace data demonstrates that Tracezip can achieve substantial performance gains in trace collection with negligible overhead. We have implemented Tracezip inside the OpenTelemetry Collector, making it compatible with existing tracing APIs.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
Paul Barham, Rebecca Isaacs, Richard Mortier, and Dushyanth Narayanan. 2003. Magpie: Online Modelling and Performance-aware Systems. In Proceedings of HotOS’03: 9th Workshop on Hot Topics in Operating Systems, May 18-21, 2003, Lihue (Kauai), Hawaii, USA , Michael B. Jones (Ed.). USENIX, 85–90. https://www.usenix.org/conference/hotos- ix/magpie-online-mode...
work page 2003
-
[2]
Protocol Buffers. 2024. Protocol Buffers: language-neutral, platform-neutral extensible mechanisms for serializing structured data. Retrieved August, 2024 from https://protobuf.dev/
work page 2024
-
[3]
Anupam Chanda, Alan L. Cox, and Willy Zwaenepoel. 2007. Whodunit: transactional profiling for multi-tier applications. In Proceedings of the 2007 EuroSys Conference, Lisbon, Portugal, March 21-23, 2007 , Paulo Ferreira, Thomas R. Gross, and Luís Veiga (Eds.). ACM, 17–30. doi:10.1145/1272996.1273001
arXiv 2007
-
[4]
Zhuangbin Chen, Zhihan Jiang, Yuxin Su, Michael R. Lyu, and Zibin Zheng. 2024. Tracemesh: Scalable and Streaming Sampling for Distributed Traces. In 17th IEEE International Conference on Cloud Computing, CLOUD 2024, Shenzhen, China, July 7-13, 2024 , Rong N. Chang, Carl K. Chang, Jingwei Yang, Nimanthi L. Atukorala, Zhi Jin, Michael Sheng, Jing Fan, Kenne...
arXiv 2024
-
[5]
Zhuangbin Chen, Yu Kang, Liqun Li, Xu Zhang, Hongyu Zhang, Hui Xu, Yangfan Zhou, Li Yang, Jeffrey Sun, Zhangwei Xu, Yingnong Dang, Feng Gao, Pu Zhao, Bo Qiao, Qingwei Lin, Dongmei Zhang, and Michael R. Lyu. 2020. Towards intelligent incident management: why we need it and how we make it. InESEC/FSE ’20: 28th ACM Joint European Software Engineering Confere...
arXiv 2020
-
[6]
Zhuangbin Chen, Jinyang Liu, Wenwei Gu, Yuxin Su, and Michael R. Lyu. 2021. Experience Report: Deep Learning- based System Log Analysis for Anomaly Detection. CoRR abs/2107.05908 (2021). arXiv:2107.05908 https://arxiv.org/ abs/2107.05908
arXiv 2021
-
[7]
Zhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang, Xiao Ling, and Michael R. Lyu. 2022. Adaptive Performance Anomaly Detection for Online Service Systems via Pattern Sketching. In 44th IEEE/ACM 44th International Conference on Software Engineering, ICSE 2022, Pittsburgh, PA, USA, May 25-27, 2022 . ACM, 61–72. doi:10.1145/3510003.3510085
arXiv 2022
-
[8]
Zhuangbin Chen, Jinyang Liu, Yuxin Su, Hongyu Zhang, Xuemin Wen, Xiao Ling, Yongqiang Yang, and Michael R. Lyu. 2021. Graph-based Incident Aggregation for Large-Scale Online Service Systems. In 36th IEEE/ACM International Conference on Automated Software Engineering, ASE 2021, Melbourne, Australia, November 15-19, 2021 . IEEE, 430–442. doi:10.1109/ASE5152...
arXiv 2021
Show all 57 references
-
[9]
Yingnong Dang, Qingwei Lin, and Peng Huang. 2019. AIOps: real-world challenges and research innovations. In Proceedings of the 41st International Conference on Software Engineering: Companion Proceedings, ICSE 2019, Montreal, QC, Canada, May 25-31, 2019, Joanne M. Atlee, Tevfi...
2019
-
[10]
Min Du, Feifei Li, Guineng Zheng, and Vivek Srikumar. 2017. DeepLog: Anomaly Detection and Diagnosis from System Logs through Deep Learning. In Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Proc. ACM Softw. Eng., Vol. 2, No. ISSTA, Article ISSTA0...
2017
-
[11]
Katz, Scott Shenker, and Ion Stoica
Rodrigo Fonseca, George Porter, Randy H. Katz, Scott Shenker, and Ion Stoica. 2007. X-Trace: A Pervasive Network Tracing Framework. In 4th Symposium on Networked Systems Design and Implementation (NSDI 2007), April 11-13, 2007, Cambridge, Massachusetts, USA, Proceedings , Hari...
2007
-
[12]
Yu Gan, Yanqi Zhang, Kelvin Hu, Dailun Cheng, Yuan He, Meghna Pancholi, and Christina Delimitrou. 2019. Seer: Leveraging Big Data to Navigate the Complexity of Performance Debugging in Cloud Microservices. In Proceedings of the Twenty-Fourth International Conference on Archite...
2019
-
[13]
Perusquía, Owen O’Brien, and Giuliano Casale
Alim Ul Gias, Yicheng Gao, Matthew Sheldon, José A. Perusquía, Owen O’Brien, and Giuliano Casale. 2023. SampleHST: Efficient On-the-Fly Selection of Distributed Traces. In NOMS 2023, IEEE/IFIP Network Operations and Management Symposium, Miami, FL, USA, May 8-12, 2023 . IEEE, ...
2023
-
[14]
Shilin He, Botao Feng, Liqun Li, Xu Zhang, Yu Kang, Qingwei Lin, Saravan Rajmohan, and Dongmei Zhang. 2023. STEAM: Observability-Preserving Trace Sampling. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Softwar...
2023
-
[15]
Shilin He, Pinjia He, Zhuangbin Chen, Tianyi Yang, Yuxin Su, and Michael R. Lyu. 2022. A Survey on Automated Log Analysis for Reliability Engineering. ACM Comput. Surv. 54, 6 (2022), 130:1–130:37. doi:10.1145/3460345
2022 doi
-
[16]
Lorch, Lidong Zhou, and Yingnong Dang
Peng Huang, Chuanxiong Guo, Jacob R. Lorch, Lidong Zhou, and Yingnong Dang. 2018. Capturing and Enhancing In Situ System Observability for Failure Detection. In 13th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2018, Carlsbad, CA, USA, October 8-10, 20...
2018
-
[17]
Zicheng Huang, Pengfei Chen, Guangba Yu, Hongyang Chen, and Zibin Zheng. 2021. Sieve: Attention-based Sampling of End-to-End Trace Data in Distributed Microservice Systems. In 2021 IEEE International Conference on Web Services, ICWS 2021, Chicago, IL, USA, September 5-10, 2021...
2021
-
[18]
Sambasivan
Darby Huye, Yuri Shkuro, and Raja R. Sambasivan. 2023. Lifting the veil on Meta’s microservice architecture: Analyses of topology and request workflows. In 2023 USENIX Annual Technical Conference, USENIX ATC 2023, Boston, MA, USA, July 10-12, 2023 , Julia Lawall and Dan Willia...
2023
-
[19]
Jaeger. 2024. An open source, distributed tracing platform . Retrieved March, 2024 from https://www.jaegertracing.io/
2024
-
[20]
Jonathan Kaldor, Jonathan Mace, Michal Bejda, Edison Gao, Wiktor Kuropatwa, Joe O’Neill, Kian Win Ong, Bill Schaller, Pingjia Shan, Brendan Viscomi, Vinod Venkataraman, Kaushik Veeraraghavan, and Yee Jiun Song. 2017. Canopy: An End-to-End Performance Tracing And Analysis Syste...
2017
-
[21]
Myunghwan Kim, Roshan Sumbaly, and Sam Shah. 2013. Root cause detection in a service-oriented architecture. In ACM SIGMETRICS / International Conference on Measurement and Modeling of Computer Systems, SIGMETRICS ’13, Pittsburgh, PA, USA, June 17-21, 2013 , Mor Harchol-Balter,...
2013
-
[22]
Las-Casas, Giorgi Papakerashvili, Vaastav Anand, and Jonathan Mace
Pedro Henrique B. Las-Casas, Giorgi Papakerashvili, Vaastav Anand, and Jonathan Mace. 2019. Sifter: Scalable Sampling for Distributed Traces, without Feature Engineering. In Proceedings of the ACM Symposium on Cloud Computing, SoCC 2019, Santa Cruz, CA, USA, November 20-23, 20...
2019
-
[23]
Cheryl Lee, Tianyi Yang, Zhuangbin Chen, Yuxin Su, and Michael R. Lyu. 2023. Eadro: An End-to-End Troubleshooting Framework for Microservices on Multi-source Data. In45th IEEE/ACM International Conference on Software Engineering, ICSE 2023, Melbourne, Australia, May 14-20, 202...
2023
-
[24]
Xiaoyun Li, Hongyu Zhang, Van-Hoang Le, and Pengfei Chen. 2024. LogShrink: Effective Log Compression by Leveraging Commonality and Variability of Log Data. In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering, ICSE 2024, Lisbon, Portugal, April ...
2024
-
[25]
Jinyang Liu, Zhihan Jiang, Jiazhen Gu, Junjie Huang, Zhuangbin Chen, Cong Feng, Zengyin Yang, Yongqiang Yang, and Michael R. Lyu. 2023. Prism: Revealing Hidden Functional Clusters from Massive Instances in Cloud Systems. In 38th IEEE/ACM International Conference on Automated S...
2023
-
[26]
Jinyang Liu, Jieming Zhu, Shilin He, Pinjia He, Zibin Zheng, and Michael R. Lyu. 2019. Logzip: Extracting Hidden Structures via Iterative Clustering for Log Compression. In 34th IEEE/ACM International Conference on Automated Software Engineering, ASE 2019, San Diego, CA, USA, ...
2019
-
[27]
Locust. 2024. An open source load testing tool . Retrieved August, 2024 from https://locust.io/
2024
-
[28]
Shutian Luo, Huanle Xu, Chengzhi Lu, Kejiang Ye, Guoyao Xu, Liping Zhang, Yu Ding, Jian He, and Chengzhong Xu. 2021. Characterizing Microservice Dependency and Performance: Alibaba Trace Analysis. In SoCC ’21: ACM Symposium on Cloud Computing, Seattle, W A, USA, November 1 - 4...
2021
-
[29]
Jonathan Mace, Ryan Roelke, and Rodrigo Fonseca. 2015. Pivot tracing: dynamic causal monitoring for distributed systems. In Proceedings of the 25th Symposium on Operating Systems Principles, SOSP 2015, Monterey, CA, USA, October 4-7, 2015, Ethan L. Miller and Steven Hand (Eds....
2015
-
[30]
OpenTelemetry. 2024. High-quality, ubiquitous, and portable telemetry to enable effective observability . Retrieved August, 2024 from https://opentelemetry.io/
2024
-
[31]
OpenTelemetry. 2024. OpenTelemetry Traces. Retrieved July, 2024 from https://opentelemetry.io/docs/concepts/signals/ traces/#spans
2024
-
[32]
OpenTelemetry. 2024. OpenTelemetry Zero-code Instrumentation. Retrieved August, 2024 from https://opentelemetry. io/docs/zero-code/
2024
-
[33]
OpenTelemetry. 2024. Trace Semantic Conventions. Retrieved October, 2024 from https://opentelemetry.io/docs/specs/ semconv/general/trace/
2024
-
[34]
OpenTelemetry. 2024. Vendor-agnostic way to receive, process and export telemetry data . Retrieved August, 2024 from https://opentelemetry.io/docs/collector/
2024
-
[35]
Xin Peng, Chenxi Zhang, Zhongyuan Zhao, Akasaka Isami, Xiaofeng Guo, and Yunna Cui. 2022. Trace analysis based microservice architecture measurement. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engi...
2022
-
[36]
Prometheus. 2024. Monitoring system & time series database . Retrieved August, 2024 from https://prometheus.io/
2024
-
[37]
Kirk Rodrigues, Yu Luo, and Ding Yuan. 2021. CLP: Efficient and Scalable Search on Compressed Text Logs. In 15th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2021, July 14-16, 2021 , Angela Demke Brown and Jay R. Lorch (Eds.). USENIX Association, 183–1...
2021
-
[38]
Junxian Shen, Han Zhang, Yang Xiang, Xingang Shi, Xinrui Li, Yunxi Shen, Zijian Zhang, Yongxiang Wu, Xia Yin, Jilong Wang, Mingwei Xu, Yahui Li, Jiping Yin, Jianchang Song, Zhuofeng Li, and Runjie Nie. 2023. Network-Centric Distributed Tracing with DeepFlow: Troubleshooting Yo...
2023
-
[39]
Benjamin H Sigelman, Luiz André Barroso, Mike Burrows, Pat Stephenson, Manoj Plakal, Donald Beaver, Saul Jaspan, and Chandan Shanbhag. 2010. Dapper, a large-scale distributed systems tracing infrastructure. (2010)
2010
-
[40]
Train Ticket. 2024. Train Ticket: A Benchmark Microservice System . Retrieved August, 2024 from https://github.com/ FudanSELab/train-ticket
2024
-
[41]
Kangjin Wang, Ying Li, Cheng Wang, Tong Jia, Kingsum Chow, Yang Wen, Yaoyong Dou, Guoyao Xu, Chuanjia Hou, Jie Yao, and Liping Zhang. 2022. Characterizing Job Microarchitectural Profiles at Scale: Dataset and Analysis. In Proceedings of the 51st International Conference on Par...
2022
-
[42]
Rui Wang, Devin Gibson, Kirk Rodrigues, Yu Luo, Yun Zhang, Kaibo Wang, Yupeng Fu, Ting Chen, and Ding Yuan
-
[43]
Junyu Wei, Guangyan Zhang, Yang Wang, Zhiwei Liu, Zhanyang Zhu, Junchao Chen, Tingtao Sun, and Qi Zhou. 2021. On the Feasibility of Parser-based Log Compression in Large-Scale Cloud Systems. In 19th USENIX Conference on File and Storage Technologies, FAST 2021, February 23-25,...
2021
-
[44]
Yang Wu, Ang Chen, and Linh Thi Xuan Phan. 2019. Zeno: Diagnosing Performance Problems with Temporal Provenance. In 16th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2019, Boston, MA, February 26-28, 2019 , Jay R. Lorch and Minlan Yu (Eds.). USENIX Ass...
2019
-
[45]
Guangba Yu, Pengfei Chen, Hongyang Chen, Zijie Guan, Zicheng Huang, Linxiao Jing, Tianjun Weng, Xinmeng Sun, and Xiaoyun Li. 2021. MicroRank: End-to-End Latency Issue Localization with Extended Spectrum Analysis in Microservice Environments. In WWW ’21: The Web Conference 2021...
2021
-
[46]
Guangba Yu, Pengfei Chen, Yufeng Li, Hongyang Chen, Xiaoyun Li, and Zibin Zheng. 2023. Nezha: Interpretable Fine-Grained Root Causes Analysis for Microservices on Multi-modal Observability Data. In Proceedings of the 31st ACM Joint European Software Engineering Conference and ...
2023
-
[47]
Guangba Yu, Zicheng Huang, and Pengfei Chen. 2023. TraceRank: Abnormal service localization with dis-aggregated end-to-end tracing data in cloud native systems. J. Softw. Evol. Process. 35, 10 (2023). doi:10.1002/SMR.2413
2023 doi
-
[48]
Lei Zhang, Zhiqiang Xie, Vaastav Anand, Ymir Vigfusson, and Jonathan Mace. 2023. The Benefit of Hindsight: Tracing Edge-Cases in Distributed Systems. In 20th USENIX Symposium on Networked Systems Design and Implementation, NSDI 2023, Boston, MA, April 17-19, 2023 , Mahesh Bala...
2023
-
[49]
Zhizhou Zhang, Murali Krishna Ramanathan, Prithvi Raj, Abhishek Parwal, Timothy Sherwood, and Milind Chabbi
- [50]
-
[51]
Tong Zhou, Chenxi Zhang, Xin Peng, Zhenghui Yan, Pairui Li, Jianming Liang, Haibing Zheng, Wujie Zheng, and Yuetang Deng. 2023. TraceStream: Anomalous Service Localization based on Trace Stream Clustering with Online Feedback. In 34th IEEE International Symposium on Software R...
2023
-
[52]
Xiang Zhou, Xin Peng, Tao Xie, Jun Sun, Chao Ji, Wenhai Li, and Dan Ding. 2021. Fault Analysis and Debugging of Microservice Systems: Industrial Survey, Benchmark System, and Empirical Study. IEEE Trans. Software Eng. 47, 2 (2021), 243–260. doi:10.1109/TSE.2018.2887384
2021
-
[53]
Xiang Zhou, Xin Peng, Tao Xie, Jun Sun, Chao Ji, Dewei Liu, Qilin Xiang, and Chuan He. 2019. Latent error prediction and fault localization for microservice applications by learning from system trace logs. In Proceedings of the ACM Joint Meeting on European Software Engineerin...
2019
-
[54]
Zipkin. 2024. A distributed tracing system . Retrieved March, 2024 from https://zipkin.io/ Proc. ACM Softw. Eng., Vol. 2, No. ISSTA, Article ISSTA019. Publication date: July 2025
2024
-
[2022]
In 2022 USENIX Annual Technical Conference, USENIX ATC 2022, Carlsbad, CA, USA, July 11-13, 2022 , Jiri Schindler and Noa Zilberman (Eds.)
CRISP: Critical Path Analysis of Large-Scale Microservice Architectures. In 2022 USENIX Annual Technical Conference, USENIX ATC 2022, Carlsbad, CA, USA, July 11-13, 2022 , Jiri Schindler and Noa Zilberman (Eds.). USENIX Association, 655–672. https://www.usenix.org/conference/a...
2022
-
[2023]
doi:10.1109/ASE56229.2023.00077
IEEE, 268–280. doi:10.1109/ASE56229.2023.00077
2023
-
[2024]
In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24)
uSlope: High Compression and Fast Search on Semi-Structured Logs. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, Santa Clara, CA, 529–544. https://www.usenix. org/conference/osdi24/presentation/wang-rui
Reviewed August 8, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.