Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Posts
Future Blog Post
Published:
Blog Post number 4
Published:
Blog Post number 3
Published:
Blog Post number 2
Published:
Blog Post number 1
Published:
portfolio
preprints
Hybrid SLM and LLM for Edge-Cloud Collaborative Inference
Published in EdgeFM'24 Workshop (Colocated with MobiCom'24), 2024 | Download Paper
Bulk Bitwise Accumulation in Commercial DRAM
Published in NeurIPS 2024 Workshop Machine Learning with new Compute Paradigms, 2024 | Download Paper
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM
Published in IEEE Computer Architecture Letters, 2025 | Download Paper
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
Published in arXiv, 2025 | Download Paper
Scaling Up On-Device LLMs via Active-Weight Swapping Between DRAM and Flash
Published in arXiv, 2025 | Download Paper
SwarmThinkers: Learning Physically Consistent Atomic KMC Transitions at Scale
Published in arXiv, 2025 | Download Paper
MemCompiler: Compile, Don’t Inject – State-Conditioned Memory for Embodied Agents
Published in arXiv, 2026 | Download Paper
EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents
Published in arXiv, 2026 | Download Paper
GRIP-VLM: Group-Relative Importance Pruning for Efficient Vision-Language Models
Published in arXiv, 2026 | Download Paper
ActProbe: Action-Space Probe for Early Failure Detection of Generative Robot Policies
Published in arXiv, 2026 | Download Paper
Embodied.cpp: A Portable Inference Runtime of Embodied AI Models on Heterogeneous Robots
Published in arXiv, 2026 | Download Paper
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence
Published in arXiv, 2026 | Download Paper
Zeva: In-Context Causal Learning for Generalizable Embodied Manipulation
Published in arXiv, 2026 | Download Paper
Zeva-Ego: Egocentric Mid-Training with In-Context Causal Learning for Robot Manipulation
Published in arXiv, 2026 | Download Paper
EBRL: Asynchronous Embodied RL by Multi-Grained Resource Management
Published in arXiv, 2026 | Download Paper
publications
Panthera: Holistic Memory Management for Big Data Processing over Hybrid Memories
Published in ACM SIGPLAN Conference on Programming Language Design and Implementation (PLDI), 2019 | Download Paper
Profiling and optimizing deep learning inference on mobile GPUs
Published in Proceedings of the 11th ACM SIGOPS Asia-Pacific Workshop on Systems (APSys), 2020 | Download Paper
To Bridge Neural Network Design and Real-World Performance: A Behaviour Study for Neural Networks
Published in Proceedings of Machine Learning and Systems (MLSys), 2021 | Download Paper
nn-Meter: towards accurate latency prediction of deep-learning model inference on diverse edge devices
Published in 19th International Conference on Mobile Systems, Applications, and Services (MobiSys), 2021 | Download Paper
MobiSys 2021 Best Paper Award
AsyMo: Scalable and Efficient Deep-Learning Inference on Asymmetric Mobile CPUs
Published in Proceedings of the 27th Annual International Conference on Mobile Computing and Networking (MobiCom), 2021 | Download Paper
Unified Holistic Memory Management Supporting Multiple Big Data Processing Frameworks over Hybrid Memories
Published in ACM Transactions on Computer Systems (TOCS), 2021 | Download Paper
nn-Meter: towards accurate latency prediction of DNN inference on diverse edge devices
Published in GetMobile: Mobile Computing and Communications, Research Highlights, 2021 | Download Paper
ACM SigMobile Research Highlight
CoDL: Efficient CPU-GPU Co-execution for Deep Learning Inference on Mobile Devices
Published in 20th International Conference on Mobile Systems, Applications, and Services (MobiSys), 2022 | Download Paper
SwiftPruner: Reinforced Evolutionary Pruning for Efficient Ad Relevance
Published in ACM International Conference on Information and Knowledge Management (CIKM), 2022 | Download Paper
MobiDepth: Real-Time Depth Estimation Using On-Device Dual Cameras
Published in Proceedings of the 28th Annual International Conference on Mobile Computing and Networking (MobiCom), 2022 | Download Paper
Romou: Rapidly Generate High-Performance Tensor Kernels for Mobile GPUs
Published in Proceedings of the 28th Annual International Conference on Mobile Computing and Networking (MobiCom), 2022 | Download Paper
Hyperion: A Generic and Distributed Mobile Offloading Framework on OpenCL
Published in The 20th ACM Conference on Embedded Networked Sensor Systems (SenSys), 2022 | Download Paper
Turbo: Opportunistic Enhancement for Edge Video Analytics
Published in The 20th ACM Conference on Embedded Networked Sensor Systems (SenSys), 2022 | Download Paper
Efficient GPU Kernels for N:M-SPARSE Weights in Deep Learning
Published in Sixth Conference on Machine Learning and Systems (MLSys), 2023 | Download Paper
Boosting DNN Cold Inference on Devices
Published in The 21st Annual International Conference on Mobile Systems, Applications and Services (MobiSys), 2023 | Download Paper
NN-Stretch: Automatic Neural Network Branching for Parallel Inference on Heterogeneous Multi-Processors
Published in The 21st International Conference on Mobile Systems, Applications, and Services (MobiSys), 2023 | Download Paper
VSPIM: SRAM Processing-in-Memory DNN Acceleration via Vector-Scalar Operations
Published in IEEE Transactions on Computers (TC), 2023 | Download Paper
HiMoDepth: Efficient Training-Free High-Resolution On-Device Depth Perception
Published in IEEE Transactions on Mobile Computing (TMC), 2023 | Download Paper
Accurate and Structured Pruning for Efficient Automatic Speech Recognition
Published in Conference of the International Speech Communication Association (INTERSPEECH), 2023 | Download Paper
Adam Accumulation to Reduce Memory Footprints of both Activations and Gradients for Large-scale DNN Training
Published in 26th European Conference on Artificial Intelligence (ECAI), 2023 | Download Paper
ElasticViT: Conflict-aware Supernet Training for Deploying Fast Vision Transformer on Diverse Mobile Devices
Published in International Conference on Computer Vision (ICCV), 2023 | Download Paper
SpaceEvo: Hardware-Friendly Search Space Design for Efficient INT8 Inference
Published in International Conference on Computer Vision (ICCV), 2023 | Download Paper
Constraint-aware and Ranking-distilled Token Pruning for Efficient Transformer Inference
Published in ACM SIGKDD Conference on Knowledge Discovery and Data Mining (KDD), 2023 | Download Paper
LUT-NN: Empower Efficient Neural Network Inference with Centroid Learning and Table Lookup
Published in The 29th Annual International Conference On Mobile Computing And Networking (MobiCom), 2023 | Download Paper
ConvStencil: Transform Stencil Computation to Matrix Multiplication on Tensor Cores
Published in ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), 2024 | Download Paper
PPoPP 2024 Best Paper Award
LitePred: Transferable and Scalable Latency Prediction for Hardware-Aware Neural Architecture Search
Published in USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2024 | Download Paper
PIM-DL: Expanding the Applicability of Commodity DRAM-PIMs for Deep Learning via Algorithm-System Co-Optimization
Published in ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2024 | Download Paper
FlexNN: Efficient and Adaptive DNN Inference on Memory-Constrained Edge Devices
Published in The 30th Annual International Conference On Mobile Computing And Networking (MobiCom), 2024 | Download Paper
Pre-gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert Inference
Published in The 51st Annual International Symposium on Computer Architecture 2024 (ISCA’24), 2024 | Download Paper
Empowering In-Browser Deep Learning Inference on Edge Through Just-In-Time Kernel Optimization
Published in The 22nd Annual International Conference on Mobile Systems, Applications and Services (MobiSys), 2024 | Download Paper
Integer or Floating Point? New Outlooks for Low-Bit Quantization on Large Language Models
Published in IEEE International Conference on Multimedia and Expo (ICME’24), 2024 | Download Paper
Ladder: Enabling Efficient Low-Precision Deep Learning Computing through Hardware-aware Tensor Transformation
Published in The 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI), 2024 | Download Paper
AFPQ: Asymmetric Floating Point Quantization for LLMs
Published in 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024 Finding short paper), 2024 | Download Paper
BitDistiller: Unleashing the Potential of Sub-4-Bit LLMs via Self-Distillation
Published in 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024 Main Conference, Long paper), 2024 | Download Paper
PruneAug: Bridging DNN Pruning and Inference Latency on Diverse Sparse Platforms Using Automatic Layerwise Block Pruning
Published in IEEE Transactions on Computers (TC), 2024 | Download Paper
VPTQ: Extreme Low-bit Vector Post-Training Quantization for Large Language Models
Published in The 2024 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024 | Download Paper
Long Exposure: Accelerating Parameter-Efficient Fine-Tuning for LLMs under Shadowy Sparsity
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC’24), 2024 | Download Paper
LoRAStencil: Low-Rank Adaptation of Stencil Computation on Tensor Cores
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC’24), 2024 | Download Paper
SC 2025 Reproducibility Challenge Finalist
Anatomizing Deep Learning Inference in Web Browsers
Published in ACM Transactions on Software Engineering and Methodology (TOSEM), 2025 | Download Paper
LUT-DLA: Lookup Table as Efficient Extreme Low-Bit Deep Learning Accelerator
Published in 31st IEEE International Symposium on High-Performance Computer Architecture (HPCA), 2025 | Download Paper
FlashFFTStencil: Bridging Fast Fourier Transforms to Memory-Efficient Stencil Computations on Tensor Core Units
Published in 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), 2025 | Download Paper
Jigsaw: Toward Conflict-free Vectorized Stencil Computation by Tessellating Swizzled Registers
Published in 30th ACM SIGPLAN Annual Symposium on Principles and Practice of Parallel Programming (PPoPP), 2025 | Download Paper
Efficient and Adaptive Diffusion Model Inference Through Lookup Table on Mobile Devices
Published in IEEE Transactions on Mobile Computing (TMC), 2025 | Download Paper
T-MAC: CPU Renaissance via Table Lookup for Low-Bit LLM Deployment on Edge
Published in The 2025 ACM European Conference on Computer Systems (EuroSys), 2025 | Download Paper
Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment
Published in The 23rd ACM Conference on Embedded Networked Sensor Systems (SenSys), 2025 | Download Paper
LUTensor: A Software-Hardware Co-Design for LUT-Based Low-Bit LLM Inference
Published in The 52nd Annual International Symposium on Computer Architecture (ISCA), 2025 | Download Paper
Bitnet.cpp: Efficient Edge Inference for Ternary LLMs
Published in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), 2025 | Download Paper
Jenga: Enhancing Long-Context Fine-tuning of LLMs with Contextual Token Sparsity
Published in USENIX Annual Technical Conference (ATC'25), 2025 | Download Paper
StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
Published in International Conference on Computer Vision (ICCV'25), 2025 | Download Paper
SeerAttention: Self-distilled Attention Gating for Efficient Long-context Prefilling
Published in The Thirty-Ninth Annual Conference on Neural Information Processing Systems (NeurIPS), 2025 | Download Paper
Matrix Is All You Need: Rearchitecting Quantum Chemistry to Scale on AI Accelerators
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2025 | Download Paper
SparStencil: Retargeting Sparse Tensor Cores to Scientific Stencil Computations via Structured Sparsity Transformation
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2025 | Download Paper
SC 2025 Best Student Paper Award Finalist
DWI: Efficient On-Device LLM Decoding Through Dynamic Width Selection on Nested Models
Published in IEEE Transactions on Mobile Computing (TMC), 2026 | Download Paper
MatXtract: Sparsity-Aware Matrix Transformation via Cascaded Compute Density EXtraction for SpMV
Published in ACM Transactions on Architecture and Code Optimization (TACO), 2026 | Download Paper
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs Decoding with Low-Bit KV Cache
Published in The 32nd International Symposium on High-Performance Computer Architecture (HPCA), 2026 | Download Paper
Neuralink: Fast on-Device LLM Inference with Neuron Co-Activation Linking
Published in ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2026 | Download Paper
Sculptor: Empowering LLMs with Cognitive Agency via Active Context Management
Published in International Conference on Learning Representations (ICLR), 2026 | Download Paper
ProRe: A Proactive Reward System for GUI Agents via Reasoner–Actor Collaboration
Published in International Conference on Learning Representations (ICLR), 2026 | Download Paper
SeerAttention-R: Sparse Attention Adaptation for Long Reasoning
Published in International Conference on Learning Representations (ICLR), 2026 | Download Paper
AVA: Towards Agentic Video Analytics Systems with Video Language Models
Published in USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2026 | Download Paper
Scaling LLM Test-Time Compute with Mobile NPU on Smartphones
Published in The 2026 ACM European Conference on Computer Systems (EuroSys), 2026 | Download Paper
Vec-LUT: Vector Table Lookup for Parallel Ultra-Low-Bit LLM Inference on Edge Devices
Published in The 24th International Conference on Mobile Systems, Applications, and Services (MobiSys), 2026 | Download Paper
This paper wins MobiSys 2026 Best Paper Runner-Up Award, and also selected as the featured paper for the On-Device AI session at MobiSys 2026.
V-Droid: Advancing Mobile GUI Agent Through Generative Verifiers
Published in The 32nd Annual International Conference On Mobile Computing And Networking (MobiCom), 2026 | Download Paper
Pushing a Single GPU to Its Limits and Scaling to Tens of Thousands: RL-Guided, Physically Consistent KMC for Nuclear Materials Simulation
Published in ISC High Performance (ISC), 2026 | Download Paper
AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation
Published in International Conference on Machine Learning (ICML), 2026 | Download Paper
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
Published in International Conference on Machine Learning (ICML), 2026 | Download Paper
PyramidFFT: Rearchitecting FFT with Matrix-Aligned Nested-Radix for Hierarchical Scratchpad Memory on AI Accelerators
Published in ACM Transactions on Architecture and Code Optimization (TACO), 2026 | Download Paper
Efficient Remote Prefix Fetching with GPU-native Media ASICs
Published in ACM SIGCOMM Conference, 2026 | Download Paper
Efficient VLM Inference System With LoRA Adapters At the Edge
Published in IEEE Transactions on Mobile Computing (TMC), 2026 | Download Paper
Em-garde: A propose-match framework for proactive streaming video understanding
Published in The European Conference on Computer Vision (ECCV), 2026 | Download Paper
Highlighted by Tsinghua University official accounts
TENET: An Efficient Sparsity-Aware LUT-Centric Architecture for Ternary LLM Inference On Edge
Published in IEEE/ACM International Symposium on Microarchitecture (MICRO), 2026 | Download Paper
FluidGPU: Fine-Grained Kernel Disaggregation for Large Model Inference on Heterogeneous GPUs
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2026 | Download Paper
MakoXC: Rearchitecting DFT Exchange-Correlation with Matrix-Aligned and Knowledge-Organized Sparsity
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2026 | Download Paper
SPIKe: Generating Practical SpMV Kernels for Sparse Iterative Methods on GPUs
Published in International Conference for High Performance Computing, Networking, Storage, and Analysis (SC), 2026 | Download Paper
AtomWorld-Mem: Memory-Restored World States for Long-Horizon Atomistic Evolution
Published in Conference on Neural Information Processing Systems (NeurIPS), 2026 | Download Paper
Learning to Commit: Generating Organic Pull Requests via Online Repository Memory
Published in Conference on Neural Information Processing Systems (NeurIPS), 2026 | Download Paper
OxyGen: Unified KV Cache Management for VLA Inference under Multi-Task Parallelism
Published in Conference on Neural Information Processing Systems (NeurIPS), 2026 | Download Paper
talks
Talk 1 on Relevant Topic in Your Field
Published:
teaching
Teaching experience 1
Undergraduate course, University 1, Department, 2014
Teaching experience 2
Workshop, University 1, Department, 2015
