Hongyuan Liu (hliu96)

Hongyuan Liu

Assistant Professor

School of Computing

Education

  • Ph.D. (2022) College of William & Mary (Computer Science)
  • M.S. (2016) The University of Hong Kong (Computer Science)
  • B.E. (2013) Shandong University (Computer Science and Technology)

Research

High-Performance Computing, GPU Computing, Computer Architecture

General Information

​​Hongyuan Liu is an Assistant Professor in the School of Computing at Stevens Institute of Technology. His research lies at the intersection of computer architecture, high-performance computing, and parallel systems, with a particular focus on improving the efficiency of GPUs for irregular and data-intensive workloads. His research has appeared in leading computer architecture and systems venues, including ASPLOS, MICRO, SIGMETRICS, PPoPP, and IPDPS, and has received an ASPLOS Best Paper Award. He earned his Ph.D. in Computer Science from William & Mary.

Institutional Service

  • Research Computing Committee Member

Professional Service

  • 2027 IEEE International Symposium on High-Performance Computer Architecture (HPCA) Program Committee Member
  • 2026 IEEE International Symposium on Workload Characterization (IISWC) Artifact Evaluation Chair
  • 2026 ACM/IEEE International Symposium on Computer Architecture (ISCA) Session Chair
  • 2026 ACM/IEEE International Symposium on Computer Architecture (ISCA) Program Committee Member
  • ACM Transactions on Architecture and Code Optimization (TACO) Journal Reviewer
  • ACM Transactions on Modeling and Performance Evaluation of Computing Systems (ToMPECS) Journal Reviewer
  • 2026 IEEE International Conference on Distributed Computing Systems (ICDCS) Program Committee Member
  • 2026 ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems (SIGMETRICS) Program Committee Member, Student Research Competition Track
  • 2026 IEEE International Symposium on High-Performance Computer Architecture (HPCA) Program Committee Member
  • 3rd Workshop on Computer Architecture Modeling and Simulation (CAMS) at MICRO 2025 Program Committee Member
  • 2025 IEEE International Symposium on High-Performance Computer Architecture (HPCA) Session Chair

Selected Publications

Conference Proceeding

  1. Chen, Z.; Ge, T.; Eeckhout, L.; Liu, H.; Huang, J. (2026). cuPTW: Leveraging Idle Compute Units for Massively Parallel GPU Page Table Walks. Proceedings of the ACM on Measurement and Analysis of Computing Systems (SIGMETRICS) (vol. 10, pp. 1-26). ACM.
    https://doi.org/10.1145/3805633.
  2. Zeng, B.; Wu, Z.; Liu, H.; Chen, X. (2026). DSTree: Data-Driven Synchronous Traversals for Decision Forests on GPUs. Proceedings of the IEEE International Parallel and Distributed Processing Symposium (IPDPS) (pp. 401-415). IEEE.
    https://doi.org/10.1109/IPDPS65963.2026.00043.
  3. Luo, W.; Chen, Y.; Yu, X.; Wang, Q.; Fan, R.; Liu, H.; Chu, X. (2026). ROME: Maximizing GPU Efficiency for All-Pairs Shortest Path via Taming Fine-Grained Irregularities. Proceedings of the 31st ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP) (pp. 204-217). ACM.
    https://doi.org/10.1145/3774934.3786461.
  4. Ge, T.; Chu, X.; Liu, H. (2025). Interleaved Bitstream Execution for Multi-Pattern Regex Matching on GPUs. Proceedings of the 58th IEEE/ACM International Symposium on Microarchitecture (MICRO) (pp. 385-400). IEEE/ACM.
    https://doi.org/10.1145/3725843.3756052.
  5. Zhao, H.; Huang, J.; Chen, Z.; Zhu, K.; Chen, D.; Ji, Z.; Liu, H. (2025). VESTA: A Secure and Efficient FHE-based Three-Party Vectorized Evaluation System for Tree Aggregation Models. Proceedings of the ACM on Measurement and Analysis of Computing Systems (SIGMETRICS) (1 ed., vol. 9, pp. 1-26). ACM.
    https://dl.acm.org/doi/10.1145/3711707.
  6. Zhu, K.; Wu, Z.; Liu, H. (2024). Efficient Point Cloud Analytics on Edge Devices. Proceedings of the 30th International Conference on Parallel and Distributed Systems (ICPADS) (pp. 382-389). IEEE.
    https://doi.org/10.1109/ICPADS63350.2024.00057.
  7. Ge, T.; Zhang, T.; Liu, H. (2024). ngAP: Non-blocking Large-scale Automata Processing on GPUs. Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), Volume 1 (pp. 268-285). ACM.
    https://dl.acm.org/doi/10.1145/3617232.3624848.
  8. Liu, H.; Pai, S.; Jog, A. (2023). Asynchronous Automata Processing on GPUs. Proceedings of the ACM on Measurement and Analysis of Computing Systems (SIGMETRICS) (1 ed., vol. 7, pp. 1-27). ACM.
    https://dl.acm.org/doi/10.1145/3579453.
  9. Liu, H.; Nicolae, B.; Di, S.; Cappello, F.; Jog, A. (2021). Accelerating DNN Architecture Search at Scale Using Selective Weight Transfer. Proceedings of the IEEE International Conference on Cluster Computing (CLUSTER) (pp. 82-93). IEEE.
    https://doi.org/10.1109/Cluster48925.2021.00051.
  10. Liu, H.; Pai, S.; Jog, A. (2020). Why GPUs are Slow at Executing NFAs and How to Make them Faster. Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS) (pp. 251-265). ACM.
    https://dl.acm.org/doi/10.1145/3373376.3378471.
  11. Ibrahim, M. A.; Liu, H.; Kayiran, O.; Jog, A. (2019). Analyzing and Leveraging Remote-core Bandwidth for Enhanced Performance in GPUs. Proceedings of the 28th International Conference on Parallel Architectures and Compilation Techniques (PACT) (pp. 258-271). IEEE.
    https://doi.org/10.1109/PACT.2019.00028.
  12. Liu, H.; Ibrahim, M. A.; Kayiran, O.; Pai, S.; Jog, A. (2018). Architectural Support for Efficient Large-Scale Automata Processing. Proceedings of the 51st Annual IEEE/ACM International Symposium on Microarchitecture (MICRO) (pp. 908-920). IEEE.
    https://dl.acm.org/doi/10.1109/MICRO.2018.00078.
  13. Liu, H.; Lam, K.; Lin, H.; Wang, C.; Ma, J. (2016). Lightweight Dependency Checking for Parallelizing Loops with Non-Deterministic Dependency on GPU. Proceedings of the IEEE 22nd International Conference on Parallel and Distributed Systems (ICPADS) (pp. 884-893). IEEE.
    https://doi.org/10.1109/icpads.2016.0119.

Journal Article

  1. Ge, T.; Zhang, T.; Liu, H. (2026). Towards Scalable and Non-blocking Automata Processing on GPUs with ngAP. ACM Transactions on Computer Systems (TOCS) (3 ed., vol. 44, pp. 1-33). ACM.
    https://doi.org/10.1145/3748646.
  2. Wu, Z.; Ge, T.; Li, J.; Chen, X.; Liu, H. (2025). Advancing Matrix Operations for High-Performance and Memory-Efficient Automata Processing on GPUs. ACM Transactions on Architecture and Code Optimization (TACO) (4 ed., vol. 22, pp. 1-26). ACM.
    https://doi.org/10.1145/3774656.
  3. Wu, Z.; Zhao, H.; Liu, H.; Wen, W.; Li, J. (2025). gHyPart: GPU-friendly End-to-End Hypergraph Partitioner. ACM Transactions on Architecture and Code Optimization (TACO) (1 ed., vol. 22, pp. 1-25). ACM.
    https://dl.acm.org/doi/10.1145/3711925.
  4. Lin, H.; Wang, C.; Liu, H. (2018). On-GPU Thread-Data Remapping for Branch Divergence Reduction. ACM Transactions on Architecture and Code Optimization (TACO) (3 ed., vol. 15, pp. 1-24). ACM.
    https://dl.acm.org/doi/10.1145/3242089.

Courses

[2026F CS 550-A] Computer Organization and Programming
[2026S CS 521-A] TCP/IP Networking
[2025F CPE 550-A] Computer Organization and Programming
[2025F CS 521-CA] TCP/IP Networking