LLM Agents
LLM Agents
I study how LLM-based systems expose, retrieve, evaluate, and govern capabilities when they become reusable infrastructure.
Evaluation before dependency
How can agent capabilities be selected, inspected, and evaluated before they become production dependencies?
Scope
Skill retrieval, semantic identifiers, NL2SQL benchmarks, and LLM serving systems are treated as capability surfaces that should be inspected before they are trusted.
Skill retrieval
Evaluating ambiguity when multiple skills expose similar capabilities and must be selected reliably.
Semantic-ID diagnostics
Inspecting semantic identifier mappings before downstream recommendation training.
NL2SQL and serving
Benchmarking business intelligence services and scaling LLM inference systems.
Papers in this area
LLM Agents