Author
Nan Lin
Recent research
- AI & ComputingOpen access
CELLServe: An SLO-Aware and Cost Efficient LLMs Serving System for Serverless Computing Environments
Large Language Models (LLMs) have enabled diverse AI applications. However, LLMs inference impose unprecedented computational and memory overhead, creating an inherent trade-off between latency Service Level Objectives (SLOs) and resource constraints. Serverless computing, with o...