Youwei Xiao
Youwei Xiao 肖有为
School of Integrated Circuits
Peking University
Beijing, China
I am a Ph.D. candidate at the School of Integrated Circuits, Peking University, advised by Prof. Yun Liang, and a Large Language Model Algorithm Intern at ByteDance Seed. My research pursues compiler-driven hardware-software co-design for agile chips. My current work creates an extensible benchmark of kernel programming and optimization tasks for dataflow-architecture chips, develops a layered knowledge and skill system across multiple compilation paths, and builds full-spectrum data collection and training workflows for continual model evolution. The ultimate goal is a methodology and workflow that enable models and agents to rapidly acquire expert-level kernel optimization capabilities for new chip architectures.
In architecture description and exploration, I combine compiler analysis, formal methods, and architecture DSLs to represent and search design spaces while connecting customized hardware capabilities to programmable software. ISAMORE (ASPLOS 2026 Best Paper) discovers reusable custom instructions from equivalent program fragments, while Cayman identifies accelerator opportunities in complete applications while co-optimizing control flow and data access. APS/Aquas describes customized architectures through DSLs and exposes them through complete hardware/software compiler stacks.
In chip DSL and compilation, I create DSLs and intermediate representations at multiple abstraction levels, then develop compilation and synthesis passes that optimize timing, microarchitecture, implementation selection, mapping, and scheduling. Hector provides a multi-level MLIR foundation for hardware synthesis. Cement couples the cycle-deterministic CmtHDL DSL with the CmtC compiler for timing analysis and control synthesis. Clay and SkyEgg further automate microarchitecture-aware implementation selection and scheduling.
In compiler optimization and LLM systems, EggMind synthesizes equality-saturation strategies with LLM guidance, and IntelliC studies inspectable compiler representations for human-model collaboration. Spine organizes verification-bounded co-synthesis across design intent, architecture, compiler, hardware, and execution evidence. PTO Runtime supports dynamic kernel fusion and distributed execution of compiled task graphs on Ascend chips and LingQu SuperPods, while Hive provides infrastructure for multi-agent inference workloads.
Academic service. I serve on the Student Technical Program Committee of MICRO 2026 and the Technical Program Committee of MLCAD 2026. I also reviewed journal submissions for IEEE TCAD and ACM TECS in 2024.
news
| Aug 12, 2026 | Our paper Aquas has been accepted to ICCAD 2026. It presents holistic hardware-software co-optimization based on MLIR. |
|---|---|
| May 28, 2026 | Awarded 博士研究生校长奖学金 for the 2026-2027 academic year, with 10 recipients selected from the School of Integrated Circuits. |
| May 09, 2026 | Awarded 学术之芯, a 2026 academic honor from Peking University’s School of Integrated Circuits granted to 8 recipients. |
selected publications
- ICCADAquas: Enhancing Domain Specialization through Holistic Hardware-Software Co-Optimization based on MLIRIn Proceedings of the 45th IEEE/ACM International Conference on Computer-Aided Design (ICCAD ’26), 2026