What static scoring cannot tell
This score reads only the text of each description and schema: it says what is written, not what a model will do. Short tools whose names already say what they do, such as browser_close, score low here even when models use them well, and in our measurements current models chose the right tool in one step even from one-line descriptions, as long as the names and schemas differed. What the text cannot show, such as what happens over several steps and whether an added sentence helps or misleads, is what an evaluation measures, on the Pro and Team plans.