LLM Agents

I study how LLM-based systems expose, retrieve, evaluate, and govern capabilities when they become reusable infrastructure.

Evaluation before dependency

How can agent capabilities be selected, inspected, and evaluated before they become production dependencies?

Scope

Skill retrieval, semantic identifiers, NL2SQL benchmarks, and LLM serving systems are treated as capability surfaces that should be inspected before they are trusted.

Skill retrieval

Evaluating ambiguity when multiple skills expose similar capabilities and must be selected reliably.

Semantic-ID diagnostics

Inspecting semantic identifier mappings before downstream recommendation training.

NL2SQL and serving

Benchmarking business intelligence services and scaling LLM inference systems.

Papers in this area

LLM Agents

SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval

J Ding arXiv 2026

SIDInspector: A Mapping-First Diagnostic Resource for Semantic-ID Tokenizers

J Ding, H Chang, H Qin, T Liu arXiv 2026

BIS: NL2SQL Service Evaluation Benchmark for Business Intelligence Scenarios

B Caglayan, M Wang, JD Kelleher, S Fei, G Tong, J Ding, P Zhang LNCS 2025

P/D-Serve: Serving Disaggregated Large Language Model at Scale

Y Jin, T Wang, H Lin, M Song, P Li, Y Ma, Y Shan, Z Yuan, C Li, Y Sun, et al. arXiv 2024