Posts by Collection
portfolio
preprints
Hybrid SLM and LLM for Edge-Cloud Collaborative Inference
Published in EdgeFM'24 Workshop (Colocated with MobiCom'24), 2024 | Download Paper
Bulk Bitwise Accumulation in Commercial DRAM
Published in NeurIPS 2024 Workshop Machine Learning with new Compute Paradigms, 2024 | Download Paper
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM
Published in IEEE Computer Architecture Letters, 2025 | Download Paper
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
Published in arXiv, 2025 | Download Paper
Scaling Up On-Device LLMs via Active-Weight Swapping Between DRAM and Flash
Published in arXiv, 2025 | Download Paper
SwarmThinkers: Learning Physically Consistent Atomic KMC Transitions at Scale
Published in arXiv, 2025 | Download Paper
MemCompiler: Compile, Don’t Inject – State-Conditioned Memory for Embodied Agents
Published in arXiv, 2026 | Download Paper
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
Published in arXiv, 2026 | Download Paper
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models
Published in arXiv, 2026 | Download Paper
ActProbe: Action-Space Probe for Early Failure Detection of Generative Robot Policies
Published in arXiv, 2026 | Download Paper
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots
Published in arXiv, 2026 | Download Paper
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
Published in arXiv, 2026 | Download Paper
Zeva: In-Context Causal Learning for Generalizable Embodied Manipulation
Published in arXiv, 2026 | Download Paper
Zeva-Ego: Egocentric Mid-Training with In-Context Causal Learning for Robot Manipulation
Published in arXiv, 2026 | Download Paper
EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management
Published in arXiv, 2026 | Download Paper
publications
Panthera: Holistic Memory Management for Big Data Processing over Hybrid Memories
Published in ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI), 2019 | Download Paper
Profiling and optimizing deep learning inference on mobile GPUs
Published in Proceedings of the 11th ACM SIGOPS Asia-Pacific Workshop on Systems (APSys), 2020 | Download Paper
To Bridge Neural Network Design and Real-World Performance: A Behaviour Study for Neural Networks
Published in Proceedings of Machine Learning and Systems (MLSys), 2021 | Download Paper
nn-Meter: towards accurate latency prediction of deep-learning model inference on diverse edge devices
Published in 19th International Conference on Mobile Systems, Applications, and Services (MobiSys), 2021 | Download Paper
MobiSys 2021 Best Paper Award
AsyMo: Scalable and Efficient Deep-Learning Inference on Asymmetric Mobile CPUs
Published in Proceedings of the 27th Annual International Conference on Mobile Computing and Networking (MobiCom), 2021 | Download Paper
Unified Holistic Memory Management Supporting Multiple Big Data Processing Frameworks over Hybrid Memories
Published in ACM Transactions on Computer Systems (TOCS), 2021 | Download Paper
nn-Meter: towards accurate latency prediction of DNN inference on diverse edge devices
Published in GetMobile: Mobile Computing and Communications, Research Highlights, 2021 | Download Paper
ACM SigMobile Research Highlight
CoDL: Efficient CPU-GPU Co-execution for Deep Learning Inference on Mobile Devices
Published in 20th International Conference on Mobile Systems, Applications, and Services (MobiSys), 2022 | Download Paper
SwiftPruner: Reinforced Evolutionary Pruning for Efficient Ad Relevance
Published in ACM International Conference on Information and Knowledge Management (CIKM), 2022 | Download Paper
MobiDepth: Real-Time Depth Estimation Using On-Device Dual Cameras
Published in Proceedings of the 28th Annual International Conference on Mobile Computing and Networking (MobiCom), 2022 | Download Paper
Romou: Rapidly Generate High-Performance Tensor Kernels for Mobile GPUs
Published in Proceedings of the 28th Annual International Conference on Mobile Computing and Networking (MobiCom), 2022 | Download Paper
Hyperion: A Generic and Distributed Mobile Offloading Framework on OpenCL
Published in The 20th ACM Conference on Embedded Networked Sensor Systems (SenSys), 2022 | Download Paper
Turbo: Opportunistic Enhancement for Edge Video Analytics
Published in The 20th ACM Conference on Embedded Networked Sensor Systems (SenSys), 2022 | Download Paper
Efficient GPU Kernels for N:M-SPARSE Weights in Deep Learning
Published in Sixth Conference on Machine Learning and Systems (MLSys), 2023 | Download Paper
Boosting DNN Cold Inference on Devices
Published in The 21st Annual International Conference on Mobile Systems, Applications and Services (MobiSys), 2023 | Download Paper
NN-Stretch: Automatic Neural Network Branching for Parallel Inference on Heterogeneous Multi-Processors
Published in The 21st International Conference on Mobile Systems, Applications, and Services (MobiSys), 2023 | Download Paper
VSPIM: SRAM Processing-in-Memory DNN Acceleration via Vector-Scalar Operations
Published in IEEE Transactions on Computers (TC), 2023 | Download Paper
HiMoDepth: Efficient Training-Free High-Resolution On-Device Depth Perception
Published in IEEE Transactions on Mobile Computing (TMC), 2023 | Download Paper
Accurate and Structured Pruning for Efficient Automatic Speech Recognition
Published in Conference of the International Speech Communication Association (INTERSPEECH), 2023 | Download Paper
Adam Accumulation to Reduce Memory Footprints of both Activations and Gradients for Large-scale DNN Training
Published in 26th European Conference on Artificial Intelligence (ECAI), 2023 | Download Paper
ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile Devices
Published in International Conference on Computer Vision (ICCV), 2023 | Download Paper
SpaceEvo: Hardware-Friendly Search Space Design for Efficient INT8 Inference
Published in International Conference on Computer Vision (ICCV), 2023 | Download Paper
Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer Inference
Published in ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2023 | Download Paper
LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table Lookup
Published in The 29th Annual International Conference On Mobile Computing And Networking (MobiCom), 2023 | Download Paper
ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor Cores
Published in ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), 2024 | Download Paper
PPoPP 2024 Best Paper Award
LitePred: Transferable and Scalable Latency Prediction for Hardware-Aware Neural Architecture Search
Published in USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2024 | Download Paper
PIM-DL: Expanding the Applicability of Commodity DRAM-PIMs for Deep Learning via Algorithm-System Co-Optimization
Published in ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2024 | Download Paper
FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge Devices
Published in The 30th Annual International Conference On Mobile Computing And Networking (MobiCom), 2024 | Download Paper
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
Published in The 51st Annual International Symposium on Computer Architecture 2024 (ISCA’24), 2024 | Download Paper
Empowering In-Browser Deep Learning Inference on Edge Through Just-In-Time Kernel Optimization
Published in The 22nd Annual International Conference on Mobile Systems, Applications and Services (MobiSys), 2024 | Download Paper
Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models
Published in IEEE International Conference on Multimedia and Expo (ICME’24), 2024 | Download Paper
Ladder: Enabling Efficient Low-Precision Deep Learning Computing through Hardware-aware Tensor Transformation
Published in The 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 2024 | Download Paper
AFPQ: Asymmetric Floating Point Quantization for LLMs
Published in 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024 Finding short paper), 2024 | Download Paper
BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
Published in 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024 Main Conference, Long paper), 2024 | Download Paper
PruneAug: Bridging DNN Pruning and Inference Latency on Diverse Sparse Platforms Using Automatic Layerwise Block Pruning
Published in IEEE Transactions on Computers (TC), 2024 | Download Paper
VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models
Published in The 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024 | Download Paper
Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC’24), 2024 | Download Paper
LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor Cores
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC’24), 2024 | Download Paper
SC 2025 Reproducibility Challenge Finalist
Anatomizing Deep Learning Inference in Web Browsers
Published in ACM Transactions on Software Engineering and Methodology (TOSEM), 2025 | Download Paper
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
Published in 31st IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2025 | Download Paper
FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Tensor Core Units
Published in 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), 2025 | Download Paper
Jigsaw: Toward Conflict-free Vectorized Stencil Computation by Tessellating Swizzled Registers
Published in 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), 2025 | Download Paper
Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile Devices
Published in IEEE Transactions on Mobile Computing (TMC), 2025 | Download Paper
T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge
Published in The 2025 ACM European Conference on Computer Systems (EuroSys), 2025 | Download Paper
Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment
Published in The 23rd ACM Conference on Embedded Networked Sensor Systems (SenSys), 2025 | Download Paper
LUTensor: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
Published in The 52nd Annual International Symposium on Computer Architecture (ISCA), 2025 | Download Paper
Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
Published in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), 2025 | Download Paper
Jenga: Enhancing Long-Context Fine-tuning of LLMs with Contextual Token Sparsity
Published in USENIX Annual Technical Conference (ATC'25), 2025 | Download Paper
StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
Published in International Conference on Computer Vision (ICCV'25), 2025 | Download Paper
SeerAttention: Self-distilled Attention Gating for Efficient Long-context Prefilling
Published in The Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS), 2025 | Download Paper
Matrix Is All You Need: Rearchitecting Quantum Chemistry to Scale on AI Accelerators
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2025 | Download Paper
SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity Transformation
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2025 | Download Paper
SC 2025 Best Student Paper Award Finalist
DWI: Efficient On-Device LLM Decoding Through Dynamic Width Selection on Nested Models
Published in IEEE Transactions on Mobile Computing (TMC), 2026 | Download Paper
MatXtract: Sparsity-Aware Matrix Transformation via Cascaded Compute Density EXtraction for SpMV
Published in ACM Transactions on Architecture and Code Optimization (TACO), 2026 | Download Paper
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs Decoding with Low-Bit KV Cache
Published in The 32nd International Symposium on High-Performance Computer Architecture (HPCA), 2026 | Download Paper
Neuralink: Fast on-Device LLM Inference with Neuron Co-Activation Linking
Published in ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2026 | Download Paper
Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management
Published in International Conference on Learning Representations (ICLR), 2026 | Download Paper
ProRe: A Proactive Reward System for GUI Agents via Reasoner–Actor Collaboration
Published in International Conference on Learning Representations (ICLR), 2026 | Download Paper
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Published in International Conference on Learning Representations (ICLR), 2026 | Download Paper
AVA: Towards Agentic Video Analytics Systems with Video Language Models
Published in USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2026 | Download Paper
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
Published in The 2026 ACM European Conference on Computer Systems (EuroSys), 2026 | Download Paper
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
Published in The 24th International Conference on Mobile Systems, Applications, and Services (MobiSys), 2026 | Download Paper
This paper wins MobiSys 2026 Best Paper Runner-Up Award, and also selected as the featured paper for the On-Device AI session at MobiSys 2026.
V-Droid: Advancing Mobile GUI Agent Through Generative Verifiers
Published in The 32nd Annual International Conference On Mobile Computing And Networking (MobiCom), 2026 | Download Paper
Pushing a Single GPU to Its Limits and Scaling to Tens of Thousands: RL-Guided, Physically Consistent KMC for Nuclear Materials Simulation
Published in ISC High Performance (ISC), 2026 | Download Paper
AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation
Published in International Conference on Machine Learning (ICML), 2026 | Download Paper
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
Published in International Conference on Machine Learning (ICML), 2026 | Download Paper
PyramidFFT: Rearchitecting FFT with Matrix-Aligned Nested-Radix for Hierarchical Scratchpad Memory on AI Accelerators
Published in ACM Transactions on Architecture and Code Optimization (TACO), 2026 | Download Paper
Efficient Remote Prefix Fetching with GPU-native Media ASICs
Published in ACM SIGCOMM Conference, 2026 | Download Paper
Efficient VLM Inference System With LoRA Adapters At the Edge
Published in IEEE Transactions on Mobile Computing (TMC), 2026 | Download Paper
Em-garde: A propose-match framework for proactive streaming video understanding
Published in The European Conference on Computer Vision (ECCV), 2026 | Download Paper
Highlighted by Tsinghua University official accounts
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
Published in IEEE/ACM International Symposium on Microarchitecture (MICRO), 2026 | Download Paper
FluidGPU: Fine-Grained Kernel Disaggregation for Large Model Inference on Heterogeneous GPUs
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2026 | Download Paper
MakoXC: Rearchitecting DFT Exchange-Correlation with Matrix-Aligned and Knowledge-Organized Sparsity
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2026 | Download Paper
SPIKe: Generating Practical SpMV Kernels for Sparse Iterative Methods on GPUs
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2026 | Download Paper
AtomWorld-Mem: Memory-Restored World States for Long-Horizon Atomistic Evolution
Published in Conference on Neural Information Processing Systems (NeurIPS), 2026 | Download Paper
Learning to Commit: Generating Organic Pull Requests via Online Repository Memory
Published in Conference on Neural Information Processing Systems (NeurIPS), 2026 | Download Paper
OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism
Published in Conference on Neural Information Processing Systems (NeurIPS), 2026 | Download Paper
talks
Talk 1 on Relevant Topic in Your Field
Published:
teaching
Teaching experience 1
Undergraduate course, University 1, Department, 2014
Teaching experience 2
Workshop, University 1, Department, 2015
