{"work":{"id":"3392f2c8-a1cb-4d6c-8c82-2cdccffa33f9","openalex_id":null,"doi":null,"arxiv_id":"2504.17761","raw_key":null,"title":"Step1X-Edit: A Practical Framework for General Image Editing","authors":null,"authors_text":"Shiyu Liu, Yucheng Han, Peng Xing, Fukun Yin, Rui Wang, Wei Cheng","year":2025,"venue":"cs.CV","abstract":"In recent years, image editing models have witnessed remarkable and rapid development. The recent unveiling of cutting-edge multimodal models such as GPT-4o and Gemini2 Flash has introduced highly promising image editing capabilities. These models demonstrate an impressive aptitude for fulfilling a vast majority of user-driven editing requirements, marking a significant advancement in the field of image manipulation. However, there is still a large gap between the open-source algorithm with these closed-source models. Thus, in this paper, we aim to release a state-of-the-art image editing model, called Step1X-Edit, which can provide comparable performance against the closed-source models like GPT-4o and Gemini2 Flash. More specifically, we adopt the Multimodal LLM to process the reference image and the user's editing instruction. A latent embedding has been extracted and integrated with a diffusion image decoder to obtain the target image. To train the model, we build a data generation pipeline to produce a high-quality dataset. For evaluation, we develop the GEdit-Bench, a novel benchmark rooted in real-world user instructions. Experimental results on GEdit-Bench demonstrate that Step1X-Edit outperforms existing open-source baselines by a substantial margin and approaches the performance of leading proprietary models, thereby making significant contributions to the field of image editing.","external_url":"https://arxiv.org/abs/2504.17761","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-10T01:46:41.223286+00:00","pith_arxiv_id":"2504.17761","created_at":"2026-05-09T06:05:34.730852+00:00","updated_at":"2026-07-10T01:46:41.223286+00:00","title_quality_ok":true,"display_title":"Step1X-Edit: A Practical Framework for General Image Editing","render_title":"Step1X-Edit: A Practical Framework for General Image Editing"},"hub":{"state":{"work_id":"3392f2c8-a1cb-4d6c-8c82-2cdccffa33f9","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":104,"external_cited_by_count":null,"distinct_field_count":5,"first_pith_cited_at":"2025-04-29T12:14:47+00:00","last_pith_cited_at":"2026-07-09T17:59:06+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T17:59:25.399880+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":12},{"context_role":"dataset","n":12},{"context_role":"baseline","n":9},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":12},{"context_polarity":"use_dataset","n":12},{"context_polarity":"baseline","n":10}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Step1X-Edit: A Practical Framework for General Image Editing","claims":[{"claim_text":"In recent years, image editing models have witnessed remarkable and rapid development. The recent unveiling of cutting-edge multimodal models such as GPT-4o and Gemini2 Flash has introduced highly promising image editing capabilities. These models demonstrate an impressive aptitude for fulfilling a vast majority of user-driven editing requirements, marking a significant advancement in the field of image manipulation. However, there is still a large gap between the open-source algorithm with these closed-source models. Thus, in this paper, we aim to release a state-of-the-art image editing mode","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"\"medium view\" to \"close facial framing\"), our model follows the prompt accurately and produces videos with stable visual texture and consistent temporal evolution. These examples further demonstrate the effectiveness of the unified architecture for high-quality text-to-video generation. 5.2.3 Multimodal Editing Quantitative Results.We evaluate the image editing capability of our model on GEdit-Bench [77]. As shown in Table 7, our model achieves the best Avg/G_O score (7.30) among unified models,","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"405 0.279 0.548 Gemini-2.5-Flash-Image [92] 0.825 0.276 0.298 0.427 0.198 0.337 Emu3.5 (Ours) 0.8530.9410.3000.386 0.166 0.529 Table 6: Quantitative evaluation results on OneIG-ZH [12]. The overall score is the average of the five dimensions. 6.2 Any-to-Image Quantitative Analysis.We conducted quantitative evaluations on ImgEdit [117], GEdit-Bench [58], OmniCon- text [107], and ICE-Bench [68] using GPT-4o, GPT-4.1, and open-sourced models required for X2I tasks (across zero, one, or multiple inp","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"2B [16] as text encoder (5K steps), (2) full fine-tuning on 4M editing samples (50K steps), and (3) high-quality re- finement for human preference alignment (1K steps). Ad- ditional details are available in the main paper [5]. 3. Experiments We evaluateEditMGTon standard benchmarks: Emu Edit [14], MagicBrush [22], AnyBench [21], and GEdit- EN-full [10]. Details are available in the main paper [5]. Quantitative Results.As reported in Table 1, EditMGT achieves state-of-the-art CLIP im scores on bo","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"models, including Qwen3-VL [1] and InternVL3.5 [42], across standard multimodal understanding benchmarks. It also outperforms strong public image generators such as Z-Image [4], Flux [28], and Qwen-Image [43] on the GenEval2 benchmark, and demonstrates consistently stronger qualitative performance on classical image editing tasks compared to Step1X [29], Qwen-Image-Edit [43], and Z-Image-Edit [4]. Beyond static image manipulation, O mni extends unified multimodal modeling to video generation and","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"First, many spatially-conditioned world-model or video- generation pipelines require expert inputs such as full 6-DoF camera trajecto- ries [34,54], creating a steep usability barrier for typical image-editing scenarios. Benchmark Object Object Object Camera VLM Precise Translation Scaling Rotation Manipulation Metric Metric ImgEdit [65] ✗ ✗ ✗ ✗ ✔ ✗ GEdit [32] ✔ ✗ ✗ ✗ ✔ ✗ CEdit [46] ✔ ✗ ✗ ✔ ✔ ✗ SpatialEdit-Bench ✔ ✔ ✔ ✔ ✔ ✔ Table 1: Characteristics comparison with other editing benchmarks. Spati","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"which supports multiple image editing operations, produces a re-edited image that progressively corrects flaws and better aligns with the desired output. 4.5 Evaluation Agent The Evaluation Agent assesses the quality of re-edited images along three human-aligned dimen- sions: perceptual quality, instruction following, and visual consistency. It is built upon a VLM and 6 Table 1: Performance ofEditRefineron EditFHF-15K, GEdit-Bench [30], I2I-Bench [47], and KRIS-Bench [58].PQ: Perceptual Quality,","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Step1X-Edit: A Practical Framework for General Image Editing because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (12 contexts).","role_counts":[{"n":12,"context_role":"background"},{"n":12,"context_role":"dataset"},{"n":9,"context_role":"baseline"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-07-04T10:26:50.028894+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"dc5f261e-5ad8-416c-8c14-8ed305a8e25d","orcid":null,"display_name":"Shiyu Liu"},{"id":"d016a0dc-a943-40da-b3e1-0e03a927b1a4","orcid":null,"display_name":"Yucheng Han"},{"id":"3bf649eb-946c-4bc4-80e4-c930ccd35d5a","orcid":null,"display_name":"Peng Xing"},{"id":"8e310054-cfd5-4cbf-be3f-1bd5a007fea4","orcid":null,"display_name":"Fukun Yin"},{"id":"ff40bdbb-2ca4-4a95-8c8c-2cb7b2e54331","orcid":null,"display_name":"Rui Wang"},{"id":"6a1922a3-5f70-4700-935f-723b20580820","orcid":null,"display_name":"Wei Cheng"}]},"error":null,"updated_at":"2026-07-04T10:26:50.343833+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T15:21:52.765422+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":29},{"title":"Emerging Properties in Unified Multimodal Pretraining","work_id":"e0cfd82c-f5d4-44fd-b531-ec73ab0a805b","shared_citers":27},{"title":"OmniGen2: Towards Instruction-Aligned Multimodal Generation","work_id":"d3153e5f-b6e2-4ab3-9f41-e24e24d64496","shared_citers":26},{"title":"ImgEdit: A Unified Image Editing Dataset and Benchmark","work_id":"059b5c3a-404c-4d30-a631-68c1d88a08a7","shared_citers":19},{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","work_id":"5dfe19d5-3541-4803-8fe9-3c8b9e29b281","shared_citers":18},{"title":"UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation","work_id":"488a273e-95d8-46f1-87c7-2244068d00d0","shared_citers":18},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":17},{"title":"Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling","work_id":"67d9e391-26d1-459e-ab56-07e60511c886","shared_citers":14},{"title":"BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset","work_id":"86d896d2-592f-4d9b-938e-dfeb11f9388f","shared_citers":12},{"title":"ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment","work_id":"94248955-4bc5-4517-98a0-66224a36d865","shared_citers":11},{"title":"In-context edit: Enabling instructional image editing with in-context generation in large scale diffusion transformer","work_id":"8cafd680-1279-4695-8156-9079db325469","shared_citers":11},{"title":"Seedream 4.0: Toward Next-generation Multimodal Image Generation","work_id":"15c839a0-48a3-4218-82b6-cac5b7f66e13","shared_citers":11},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":11},{"title":"Emu3: Next-Token Prediction is All You Need","work_id":"720d288e-fac0-464c-9929-19efd9a52afc","shared_citers":9},{"title":"Prompt-to-Prompt Image Editing with Cross Attention Control","work_id":"196f7eef-d65a-47e4-b815-9a188f6aedcf","shared_citers":9},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":9},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":9},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":9},{"title":"arXiv preprint arXiv:2507.22058 (2025)","work_id":"3ee0ee57-31b9-49d4-98fc-c12c499a14b9","shared_citers":8},{"title":"Flow-GRPO: Training Flow Matching Models via Online RL","work_id":"bf1e8e81-ff31-401a-a5dc-d9c49df168ab","shared_citers":8},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":8},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":8},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":8},{"title":"Seedream 3.0 Technical Report","work_id":"013e56d0-7f47-4d0e-bbca-e9540fc0e0cc","shared_citers":8}],"time_series":[{"n":6,"year":2025},{"n":38,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T15:21:56.881765+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T15:21:59.165016+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Step1X-Edit: A Practical Framework for General Image Editing","claims":[{"claim_text":"In recent years, image editing models have witnessed remarkable and rapid development. The recent unveiling of cutting-edge multimodal models such as GPT-4o and Gemini2 Flash has introduced highly promising image editing capabilities. These models demonstrate an impressive aptitude for fulfilling a vast majority of user-driven editing requirements, marking a significant advancement in the field of image manipulation. However, there is still a large gap between the open-source algorithm with these closed-source models. Thus, in this paper, we aim to release a state-of-the-art image editing mode","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"\"medium view\" to \"close facial framing\"), our model follows the prompt accurately and produces videos with stable visual texture and consistent temporal evolution. These examples further demonstrate the effectiveness of the unified architecture for high-quality text-to-video generation. 5.2.3 Multimodal Editing Quantitative Results.We evaluate the image editing capability of our model on GEdit-Bench [77]. As shown in Table 7, our model achieves the best Avg/G_O score (7.30) among unified models,","claim_type":"dataset","confidence":0.95,"evidence_strength":"citation_context"},{"claim_text":"405 0.279 0.548 Gemini-2.5-Flash-Image [92] 0.825 0.276 0.298 0.427 0.198 0.337 Emu3.5 (Ours) 0.8530.9410.3000.386 0.166 0.529 Table 6: Quantitative evaluation results on OneIG-ZH [12]. The overall score is the average of the five dimensions. 6.2 Any-to-Image Quantitative Analysis.We conducted quantitative evaluations on ImgEdit [117], GEdit-Bench [58], OmniCon- text [107], and ICE-Bench [68] using GPT-4o, GPT-4.1, and open-sourced models required for X2I tasks (across zero, one, or multiple inp","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"2B [16] as text encoder (5K steps), (2) full fine-tuning on 4M editing samples (50K steps), and (3) high-quality re- finement for human preference alignment (1K steps). Ad- ditional details are available in the main paper [5]. 3. Experiments We evaluateEditMGTon standard benchmarks: Emu Edit [14], MagicBrush [22], AnyBench [21], and GEdit- EN-full [10]. Details are available in the main paper [5]. Quantitative Results.As reported in Table 1, EditMGT achieves state-of-the-art CLIP im scores on bo","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"models, including Qwen3-VL [1] and InternVL3.5 [42], across standard multimodal understanding benchmarks. It also outperforms strong public image generators such as Z-Image [4], Flux [28], and Qwen-Image [43] on the GenEval2 benchmark, and demonstrates consistently stronger qualitative performance on classical image editing tasks compared to Step1X [29], Qwen-Image-Edit [43], and Z-Image-Edit [4]. Beyond static image manipulation, O mni extends unified multimodal modeling to video generation and","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"First, many spatially-conditioned world-model or video- generation pipelines require expert inputs such as full 6-DoF camera trajecto- ries [34,54], creating a steep usability barrier for typical image-editing scenarios. Benchmark Object Object Object Camera VLM Precise Translation Scaling Rotation Manipulation Metric Metric ImgEdit [65] ✗ ✗ ✗ ✗ ✔ ✗ GEdit [32] ✔ ✗ ✗ ✗ ✔ ✗ CEdit [46] ✔ ✗ ✗ ✔ ✔ ✗ SpatialEdit-Bench ✔ ✔ ✔ ✔ ✔ ✔ Table 1: Characteristics comparison with other editing benchmarks. Spati","claim_type":"baseline","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"which supports multiple image editing operations, produces a re-edited image that progressively corrects flaws and better aligns with the desired output. 4.5 Evaluation Agent The Evaluation Agent assesses the quality of re-edited images along three human-aligned dimen- sions: perceptual quality, instruction following, and visual consistency. It is built upon a VLM and 6 Table 1: Performance ofEditRefineron EditFHF-15K, GEdit-Bench [30], I2I-Bench [47], and KRIS-Bench [58].PQ: Perceptual Quality,","claim_type":"dataset","confidence":0.9,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Step1X-Edit: A Practical Framework for General Image Editing because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (12 contexts).","role_counts":[{"n":12,"context_role":"background"},{"n":12,"context_role":"dataset"},{"n":9,"context_role":"baseline"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-07-04T10:26:50.031373+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Step1X-Edit: A Practical Framework for General Image Editing","claims":[{"claim_text":"In recent years, image editing models have witnessed remarkable and rapid development. The recent unveiling of cutting-edge multimodal models such as GPT-4o and Gemini2 Flash has introduced highly promising image editing capabilities. These models demonstrate an impressive aptitude for fulfilling a vast majority of user-driven editing requirements, marking a significant advancement in the field of image manipulation. However, there is still a large gap between the open-source algorithm with these closed-source models. Thus, in this paper, we aim to release a state-of-the-art image editing mode","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Step1X-Edit: A Practical Framework for General Image Editing because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T15:22:00.972605+00:00"}},"summary":{"title":"Step1X-Edit: A Practical Framework for General Image Editing","claims":[{"claim_text":"In recent years, image editing models have witnessed remarkable and rapid development. The recent unveiling of cutting-edge multimodal models such as GPT-4o and Gemini2 Flash has introduced highly promising image editing capabilities. These models demonstrate an impressive aptitude for fulfilling a vast majority of user-driven editing requirements, marking a significant advancement in the field of image manipulation. However, there is still a large gap between the open-source algorithm with these closed-source models. Thus, in this paper, we aim to release a state-of-the-art image editing mode","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Step1X-Edit: A Practical Framework for General Image Editing because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"Qwen-Image Technical Report","work_id":"d06d7ecc-7579-4f89-a60b-4278a0f3c562","shared_citers":29},{"title":"Emerging Properties in Unified Multimodal Pretraining","work_id":"e0cfd82c-f5d4-44fd-b531-ec73ab0a805b","shared_citers":27},{"title":"OmniGen2: Towards Instruction-Aligned Multimodal Generation","work_id":"d3153e5f-b6e2-4ab3-9f41-e24e24d64496","shared_citers":26},{"title":"ImgEdit: A Unified Image Editing Dataset and Benchmark","work_id":"059b5c3a-404c-4d30-a631-68c1d88a08a7","shared_citers":19},{"title":"FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space","work_id":"5dfe19d5-3541-4803-8fe9-3c8b9e29b281","shared_citers":18},{"title":"UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation","work_id":"488a273e-95d8-46f1-87c7-2244068d00d0","shared_citers":18},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":17},{"title":"Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling","work_id":"67d9e391-26d1-459e-ab56-07e60511c886","shared_citers":14},{"title":"BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset","work_id":"86d896d2-592f-4d9b-938e-dfeb11f9388f","shared_citers":12},{"title":"ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment","work_id":"94248955-4bc5-4517-98a0-66224a36d865","shared_citers":11},{"title":"In-context edit: Enabling instructional image editing with in-context generation in large scale diffusion transformer","work_id":"8cafd680-1279-4695-8156-9079db325469","shared_citers":11},{"title":"Seedream 4.0: Toward Next-generation Multimodal Image Generation","work_id":"15c839a0-48a3-4218-82b6-cac5b7f66e13","shared_citers":11},{"title":"Wan: Open and Advanced Large-Scale Video Generative Models","work_id":"ad3ebc3b-4224-46c9-b61d-bcf135da0a7c","shared_citers":11},{"title":"Emu3: Next-Token Prediction is All You Need","work_id":"720d288e-fac0-464c-9929-19efd9a52afc","shared_citers":9},{"title":"Prompt-to-Prompt Image Editing with Cross Attention Control","work_id":"196f7eef-d65a-47e4-b815-9a188f6aedcf","shared_citers":9},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":9},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":9},{"title":"SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis","work_id":"8034c587-fba6-4941-87ba-c98f2ac962cb","shared_citers":9},{"title":"arXiv preprint arXiv:2507.22058 (2025)","work_id":"3ee0ee57-31b9-49d4-98fc-c12c499a14b9","shared_citers":8},{"title":"Flow-GRPO: Training Flow Matching Models via Online RL","work_id":"bf1e8e81-ff31-401a-a5dc-d9c49df168ab","shared_citers":8},{"title":"Flow Matching for Generative Modeling","work_id":"6edb71c4-5d64-40af-a394-9757ea051a36","shared_citers":8},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":8},{"title":"HunyuanVideo: A Systematic Framework For Large Video Generative Models","work_id":"881efa7e-7e73-4c66-9cc3-2803e551061c","shared_citers":8},{"title":"Seedream 3.0 Technical Report","work_id":"013e56d0-7f47-4d0e-bbca-e9540fc0e0cc","shared_citers":8}],"time_series":[{"n":6,"year":2025},{"n":38,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"8e310054-cfd5-4cbf-be3f-1bd5a007fea4","orcid":null,"display_name":"Fukun Yin","source":"manual","import_confidence":0.72},{"id":"3bf649eb-946c-4bc4-80e4-c930ccd35d5a","orcid":null,"display_name":"Peng Xing","source":"manual","import_confidence":0.72},{"id":"ff40bdbb-2ca4-4a95-8c8c-2cb7b2e54331","orcid":null,"display_name":"Rui Wang","source":"manual","import_confidence":0.72},{"id":"dc5f261e-5ad8-416c-8c14-8ed305a8e25d","orcid":null,"display_name":"Shiyu Liu","source":"manual","import_confidence":0.72},{"id":"6a1922a3-5f70-4700-935f-723b20580820","orcid":null,"display_name":"Wei Cheng","source":"manual","import_confidence":0.72},{"id":"d016a0dc-a943-40da-b3e1-0e03a927b1a4","orcid":null,"display_name":"Yucheng Han","source":"manual","import_confidence":0.72}]}}