Institution

Kraftanlagen (Germany)

DEcompany

Recent research

  • AI & ComputingOpen access

    Multi-Bin Batching for Increasing LLM Inference Throughput

    As large language models (LLMs) grow in popularity for their diverse capabilities, improving the efficiency of their inference systems has become increasingly critical. Batching LLM requests is a critical step in scheduling the inference jobs on servers (e.g. GPUs), enabling the...

    ACM Transactions on Modeling and Performance Evaluation of Computing Systems2026-08-220 citationsDOI