arXiv 2024 / LLM Agents

P/D-Serve: Serving Disaggregated Large Language Model at Scale

A systems paper on disaggregated large language model serving, separating serving stages to support large-scale inference.

Research area

Skill retrieval, semantic identifiers, NL2SQL benchmarks, and LLM serving systems are treated as capability surfaces that should be inspected before they are trusted.

Publication record

Authors
Y Jin, T Wang, H Lin, M Song, P Li, Y Ma, Y Shan, Z Yuan, C Li, Y Sun, et al.
Venue
arXiv preprint arXiv:2408.08147
Year
2024
Area
LLM Agents