About
I am currently a Postdoctoral Research Fellow at the National University of Singapore (NUS), working with Prof. Jiaheng Zhang. I received my Ph.D. from Zhejiang University under the supervision of Prof. Wenzhi Chen. My research primarily focuses on leveraging GPU to optimize Zero-Knowledge Proofs and AI workloads.
Employment
-
National University of Singapore, SingaporePostdoctoral Research FellowMainly work on GPU acceleration for ZKPs and AI workloads, hosted by Prof. Jiaheng Zhang.
-
Polyhedra NetworkSoftware Engineering InternWork on accelerating ZKPs for neural networks.
-
PrimeblockAlgorithm Engineering InternWork on accelerating multi-scalar multiplication on GPUs.
Project
-
GPU Inference Acceleration for Large Language ModelsDesigned a high-throughput GPU inference approach for LLMs based on moderately unstructured sparse weight matrices. Built an optimized sparse matrix-matrix multiplication (SpMM) kernel and corresponding scheduling strategy to improve memory efficiency and end-to-end inference throughput. [Code]
-
ZPrize (3rd Place)Implemented a high-performance GPU-based multi-scalar multiplication algorithm for ZKPs. Achieved a 1.67x speedup over the reference implementation. [Code]
Paper
- Tao Lu, Haoyu Wang, Zonghui Wang, Keshen Xiang, Jiaheng Zhang, Wenzhi Chen. "Accelerating GPU Inference of Large Language Models with Moderately Unstructured Sparse Weight Matrices." (DAC 2026). [Code]
- Xin Li, Shujun Tian, Tao Lu, Han Bao, Zonghui Wang, Wenzhi Chen. "Otil: Accelerating Diffusion Model Inference via Communication-Efficient Multi-GPU Parallelism." (CVPR 2026).
- Wenjie Qu, Yijun Sun, Xuanming Liu, Tao Lu, Yanpei Guo, Kai Chen, Jiaheng Zhang. "zkGPT: An Efficient Non-interactive Zero-knowledge Proof Framework for LLM Inference." (Usenix Security 2025).
- Tao Lu, Yuxun Chen, Zonghui Wang, Xiaohang Wang, Wenzhi Chen, Jiaheng Zhang. "BatchZK: A Fully Pipelined GPU-Accelerated System for Batch Generation of Zero-Knowledge Proofs." (ASPLOS 2025). [Paper]
- Tao Lu, Haoyu Wang, Wenjie Qu, Zonghui Wang, Jinye He, Tianyang Tao, Wenzhi Chen, Jiaheng Zhang. "zkNN: An Efficient and Extensible Zero-knowledge Proof Framework for Neural Networks." [Paper]
- Tao Lu, Chengkun Wei, Ruijing Yu, Chaochao Chen, Wenjing Fang, Lei Wang, Zeke Wang, and Wenzhi Chen. "cuZK: Accelerating Zero-Knowledge Proof with a Faster Parallel Multi-scalar Multiplication Algorithm on GPUs." (CHES 2023). [Paper] [Code]
- Peiyu Liu, Shouling Ji, Lirong Fu, Kangjie Lu, Xuhong Zhang, Wei-Han Lee, Tao Lu, Wenzhi Chen, and Raheem Beyah. "Understanding the security risks of Docker Hub." (ESORICS 2020).
Skill
Programming Languages: C/C++, CUDA, Python, Verilog, Rust.