arXiv 2026 / LLM Agents

SkillResolve-Bench: Measuring and Resolving Same-Capability Ambiguity in Agent Skill Retrieval

A benchmark-focused study of same-capability ambiguity in agent skill retrieval, where similar-looking skills must be distinguished by execution constraints and task fit.

Research area

AI search, skill retrieval, NL2SQL benchmarks, LLM serving systems, and capability governance are treated as system surfaces that should be inspected before they are trusted.

Publication record

Authors
J Ding
Venue
arXiv preprint arXiv:2606.10388
Year
2026
Area
LLM Agents