Publications

You can also find my articles on my Google Scholar profile.

Conference Papers


TArCS: Trusted and Attack-Resilient Clock Source with TEE and RDMA

Yunpeng Xu, Yuchen Fan, Yu Jin, Shuangjie Yao, Jia Zhang, Teng Ma, Shuwen Deng

17th International Symposium on Advanced Parallel Processing Technologies (APPT'26), 2026

As confidential computing technologies advance, the secure and trustworthy execution of applications has garnered increasing attention. However, recent time-based attacks have severely disrupted trust and pose significant threats to the normal functioning of applications across various fields, especially the safety of real-time systems. This highlights the critical importance of a robust trusted timekeeping infrastructure. We conduct a detailed analysis of the unique challenges in trusted timekeeping systems and propose the CIAT model, which integrates the classic Confidentiality, Integrity, and Availability framework by incorporating Timeliness. Within the CIAT model, we systematically analyze various threats, including Trust Attacks, Precision Attacks, and Availability Attacks.

SSBench: Automated Characterization of Memory Dependence Predictors on Modern CPUs

Chang Liu, Yu Jin, Yuchen Fan, Tianrui Xiao, Lingfeng Yin, Trevor E. Carlson, Shuwen Deng, Dongsheng Wang

53rd Annual International Symposium on Computer Architecture (ISCA'26), 2026

Memory Dependence Predictors (MDPs) improve the performance of modern CPUs by exposing additional parallelism through predicting data dependence between store and load instructions. Since the 1990s, various MDP designs have been proposed across architectures. Recent studies reveal that MDPs are widely deployed on modern CPUs and can be exploited as side channels to leak data. However, because MDP designs are undocumented, characterizing an MDP design still requires complicated manual analysis.

DROW: Training-Free Load Speculative Execution Attacks on Apple Silicon

Yuchen Fan, Yu Jin, Chang Liu, Minghong Sun, Xuanzeng Song, Tingting Yin, Shuwen Deng

2nd Microarchitecture Security Conference (uASC'26), 2026

Traditional speculative attacks rely on training hardware predictors, limiting practicality. We present DROW, exploiting blind bypassing on Apple silicon, where loads speculatively bypass stores without training. This primitive circumvents mistraining defenses to hijack data and control flow. DROW achieves $16.3\times$ higher bandwidth than prior art, enabling practical exploits including cross-page browser leakage, ASLR/KASLR bypassing, and PAC circumvention.

Ragnar: Exploring Volatile-Channel Vulnerabilities on RDMA NIC

Yunpeng Xu, Yuchen Fan, Teng Ma, Shuwen Deng

62nd ACM/IEEE Design Automation Conference (DAC'25), 2025

With the surge in data computation, Remote Direct Memory Access (RDMA) becomes crucial to offering low-latency and high-throughput communication for data centers, but it faces new security threats. This paper presents Ragnar, a comprehensive suite of hardware-contention-based volatile-channel attacks leveraging the under-explored security vulnerabilities in RDMA hardware. Through comprehensive microbenchmark reverse engineering, we analyze RDMA NICs at multiple granularity levels and then construct covert-channel attacks, achieving 3.2x the bandwidth of state-of-the-art RDMA-targeted attacks on CX-5. We apply side-channel attacks on real-world distributed databases and disaggregated memory, where we successfully fingerprint operations and recover sensitive address data with 95.6% accuracy.

Preprints


Application-Driven Architecture Exploration for Cross-Layer Heterogeneous Systems

Yuchen Fan, Minghong Sun, Jikui Ma, Yunpeng Xu, Shunyu Mao, Liu He, Shunan Dong, Jiahao Yang, Yu Zhu, Xinhao Yang, Tianyan Zhong, Haoran Sun, Daoqi Liu, Zongle Huang, Xinyuan Lin, Huazhong Yang, Maokun Li, Yongpan Liu, Yu Wang, Zhenhua Zhu, Hongyang Jia, Shuwen Deng

arXiv:2607.23042, 2026

AI and HPC infrastructure increasingly serves workload portfolios that combine dense tensor computation, sparse kernels, large memory footprints, and communication-intensive collectives. Supporting such portfolios requires coordinated choices across accelerators, memory tiers, scale-up fabrics, and cluster networks. The resulting Cross-layer Heterogeneous System (XHS) design space is difficult to explore: hardware choices change the legal task mappings, while rack power, switch radix, cabling, and cost constraints invalidate many candidates. Existing node-scale design-space exploration tools and fixed-platform distributed simulators address only parts of this problem.