LLM Agents

I study how LLM-based systems search, expose, retrieve, evaluate, and govern capabilities when they become reusable infrastructure.

Evaluation before dependency

How can agent capabilities and search policies be selected, inspected, and evaluated before they become production dependencies?

Scope

AI search, skill retrieval, NL2SQL benchmarks, LLM serving systems, and capability governance are treated as system surfaces that should be inspected before they are trusted.

AI search

Keeping search policies aligned with changing inventory and evidence conditions.

Skill retrieval

Evaluating ambiguity when multiple skills expose similar capabilities and must be selected reliably.

Capability governance

Studying how reusable agent capabilities are represented, selected, exposed, and evaluated.

NL2SQL and serving

Benchmarking business intelligence services and scaling LLM inference systems.

Papers in this area

LLM Agents

Inventory-Grounded Policy-Level Optimization for Training-Free AI Search

W Zhou, T Wu, J Ding, Z Fan, Y Cao EMNLP 2026

SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval

J Ding arXiv 2026

BIS: NL2SQL Service Evaluation Benchmark for Business Intelligence Scenarios

B Caglayan, M Wang, JD Kelleher, S Fei, G Tong, J Ding, P Zhang ICSOC 2025

P/D-Serve: Serving Disaggregated Large Language Model at Scale

Y Jin, T Wang, H Lin, M Song, P Li, Y Ma, Y Shan, Z Yuan, C Li, Y Sun, et al. arXiv 2024