I am a Ph.D. candidate in Computer Engineering at the University of California, Santa Barbara, where I also received my M.S. in Computer Engineering. Before that, I received my B.Eng. in Electrical Engineering from Nanjing University.
My research focuses on efficient large language model (LLM) training and inference and efficient chain-of-thought (CoT) reasoning via algorithm-system co-design. I work across the stack — from compression algorithms to the systems that run them — on topics including model compression, KV cache compression, pruning, low-rank decomposition, early exit, knowledge distillation, quantization, speculative decoding/reasoning, and efficient agentic systems. My work has appeared at COLM, MLSys, EACL, EMNLP, ICASSP, and IEEE TCAD, and was developed in part during research internships at Intel and AMD-Xilinx.
You can find my publications on Google Scholar, with total citations 66 .
🔥 News
- 2026.08: 🎉 LoRi: Low-Rank Distillation for Implicit Reasoning accepted to Findings of the Association for Computational Linguistics: EMNLP 2026.
- 2026.07: 🎓 Advanced to Ph.D. candidacy in Computer Engineering at UC Santa Barbara.
- 2026.07: 🎉 RankGuide: Tensor-Rank-Guided Routing and Steering for Efficient Reasoning accepted to the Third Conference on Language Modeling (COLM 2026).
- 2026.02: 🎉 SkipKV: Selective Skipping of KV Generation and Storage for Efficient Inference with Large Reasoning Models accepted to the Ninth Annual Conference on Machine Learning and Systems (MLSys 2026).
- 2026.01: 🎉 FLAT-LLM: Fine-grained Low-rank Activation Space Transformation for Large Language Model Compression accepted to Findings of the Association for Computational Linguistics: EACL 2026.
- 2026.01: 🎉 FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training published in IEEE TCAD.
- 2025.09: 🎓 Received my M.S. in Computer Engineering from UC Santa Barbara.
- 2025.08: 🎉 Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization accepted to IEEE TCAD.
- 2025.06: 🛠️ Started my summer research internship at Intel, Hillsboro, OR.
- 2024.06: 🛠️ Started my summer research internship at Intel, Hillsboro, OR.
📝 Publications

RankGuide: Tensor-Rank-Guided Routing and Steering for Efficient Reasoning
Jiayi Tian, Yupeng Su, Ryan Solgi, Souvik Kundu, Zheng Zhang
- Leverages tensor-rank signals from hidden states to accelerate large reasoning model inference, yielding up to 1.75× and 1.36× latency benefit over LRM and SoTA collaborative inference while maintaining or improving accuracy.
- A tensor-rank scoring metric on step-level hidden states detects low-quality reasoning steps and selectively routes them to larger models.

Jiayi Tian, Seyedarmin Azizi, Yequan Zhao, Erfan Baghaei Potraghloo, Sean McPherson, Sharath Nittur Sridhar, Zhengyang Wang, Zheng Zhang, Massoud Pedram, Souvik Kundu
- A training-free KV-cache compression framework with sentence-level selective eviction and dynamic thoughts-type control for efficient CoT reasoning in multi-batch serving.
- Up to 26.7% higher accuracy, 1.6× shorter generation, and 1.7× higher throughput vs. SoTA under equal compression.

FLAT-LLM: Fine-grained Low-rank Activation Space Transformation for Large Language Model Compression
Jiayi Tian, Ryan Solgi, Jinming Lu, Yifan Yang, Hai Li, Zheng Zhang
- A training-free, fine-grained compression method that exploits the low-rank structure of the activation space to transform and compress model weights.
- A training-free rank selection algorithm based on greedy rank redistribution; strong results on LLaMA-2/3 and Mistral with calibration overhead in minutes.
- LoRi: Low-Rank Distillation for Implicit Reasoning, Ryan Solgi, Jiayi Tian, Zheng Zhang, Findings of EMNLP 2026
- IAPO: Input Attribution-Aware Policy Optimization for Tool Use in Small Multimodal Agents, Yifan Yang, Zhen Zhang, Jiayi Tian, Liyan Tan, Zheng Zhang, arXiv 2026
- FETTA: Flexible and Efficient Hardware Accelerator for Tensorized Neural Network Training, Jinming Lu, Jiayi Tian, Hai Li, Ian Young, Zheng Zhang, IEEE TCAD 2026
- Ultra Memory-Efficient On-FPGA Training of Transformers via Tensor-Compressed Optimization, Jiayi Tian, Jinming Lu, Hai Li, Xiangwei Wang, Cong (Callie) Hao, Ian Young, Zheng Zhang, IEEE TCAD 2025
- Tensor-Compressed and Fully-Quantized Training of Neural PDE Solvers, Jinming Lu, Jiayi Tian, Yequan Zhao, Hai Li, Zheng Zhang, arXiv 2025
- Activation-Informed Pareto-Guided Low-Rank Compression for Efficient LLM/VLM, Ryan Solgi, Parsa Madinei, Jiayi Tian, Rupak Swaminathan, Jing Liu, Nathan Susanj, Zheng Zhang, arXiv 2025
- BEBERT: Efficient and Robust Binary Ensemble BERT, Jiayi Tian, Chao Fang, Haonan Wang, Zhongfeng Wang, ICASSP 2023 | Code
Mentorship
- Comprehensive Design Space Exploration for Tensorized Neural Network Hardware Accelerators, Jinsong Zhang, Minghe Li, Jiayi Tian, Jinming Lu, Zheng Zhang, arXiv 2025
📖 Educations
- 2023.09 - now, Ph.D. in Computer Engineering, University of California, Santa Barbara, CA, USA. (GPA 3.93/4.0)
- 2023.09 - 2025.09, M.S. in Computer Engineering, University of California, Santa Barbara, CA, USA. (GPA 3.93/4.0)
- 2019.09 - 2023.06, B.Eng. in Electrical Engineering, Nanjing University, China. (GPA 4.50/5.0)
💻 Internships
- 2025.06 - 2025.09, Research Intern, Intel Corporation, Hillsboro, OR, USA. Mentor: Souvik Kundu. Proposed SkipKV, a training-free KV-cache compression framework with a dynamic latent thoughts-type steering mechanism for concise and stable reasoning (MLSys 2026).
- 2024.06 - 2024.09, Research Intern, Intel Corporation, Hillsboro, OR, USA. Mentor: Hai Li. Built a tensor-compressed Transformer training accelerator on FPGA with a bidirectional tensor contraction scheme, reaching up to 51× memory efficiency and 4× energy efficiency vs. an NVIDIA RTX 3090 (IEEE TCAD 2025).
- 2023.06 - 2023.09, Co-Op/Intern, AMD-Xilinx Technology, Beijing, China. Developed a C++/HLS Transformer training framework with custom tensorized linear layers and nonlinear operators, achieving 30×~52× model size savings for end-to-end Transformer training.