Research focus on Operating Systems, Cloud Computing, Network Systems, AI Infrastructure, and System Software at Korea University.
AI has been transitioning to essential everyday infrastructure. System software determines whether large models are practical at scale—it controls how efficiently hardware resources are used, how predictable runtimes are, and how easily workloads can scale across multiple machines. For example, in recent years, datacenters have evolved to support diverse, heterogeneous workloads and GPU devices. One prominent example is distributed deep learning training, which enables the development of large-scale AI models such as GPT, DALL-E, and LLaMA. These models can involve more than 530 billion parameters and rely on hundreds of GPU nodes. At the same time, large language model (LLM) inference and serving have emerged as critical workloads that demand efficient KV-Cache management, request scheduling, and low-latency execution. As a result, optimizing infrastructure efficiency and utilization—both in training and inference—has become increasingly important. However, recent reports from major cloud providers such as Microsoft and Alibaba show that average GPU utilization rates are only 20–50%. Such low utilization indicates a significant waste of datacenter resources, underscoring the need for more effective systems strategies to improve efficiency. Our research aims to make AI systems—spanning training, inference, and serving—faster, more affordable, energy-efficient, and reliable as models and datasets continue to grow.
System software for networking matters because modern computing is inherently connected. Every cloud service, data pipeline, and AI workload depends on timely and reliable communication. We enhance performance, isolation, and predictability through diverse approaches: 1) network virtualization that allows multiple tenants to share the same physical network without interference; 2) programmable, softwarized networks to enable the network to be managed as reconfigurable resources; and 3) kernel-level improvements to the networking stack. On the mobile side, we study flexible and high-performance networking systems by extending GPU- and software-based network stacks, particularly in emerging 6G architectures such as AI-RAN. In summary, our goal is to design structurally isolated and well-architected networking systems that ensure resource efficiency, predictable low latency, and a high quality of experience across diverse platforms—from edge devices and datacenter infrastructure to mobile systems like 6G.
We also apply well-designed system software to a variety of application domains. One promising area is digital healthcare. In this domain, personal and sensitive health data are collected by edge and wearable devices, and must be stored and processed securely and privately. Since user-side platforms often have limited computational resources, designing compact and efficient systems is essential. For example, providing personalized healthcare services directly on edge or wearable devices reduces round-trip latency and bandwidth usage while maintaining privacy and regulatory compliance. This capability is crucial for timely and reliable clinical tasks. Such application requirements highlight the importance of strong system software foundations. For instance, blockchain-based data management enhances resilience and privacy in data handling, while geo-distributed cloud management enables timely service delivery even on resource-constrained devices. Such system-level optimizations yield measurable improvements in reliability and cost efficiency in production environments. Beyond digital healthcare, we are developing diverse applications that leverage operating systems, virtualization, and system software to build practical technologies to benefit people and society.