Read your scores
How a tool's score out of 100 is made, what the checks across the server look for, and every issue code with its severity.
On this page
Each tool gets a score out of 100, the sum of six items. The server's average is the mean of its tools' scores, with one decimal place. Each result also lists issues: what to fix, each with an issue code and a severity.
The six items
| Item | Points |
|---|---|
| Purpose | 20 |
| When to use | 20 |
| Arguments | 25 |
| Return value | 15 |
| Constraints and side effects | 10 |
| Examples | 10 |
- Purpose: the description exists, says what the tool does (starting with a verb helps), is long enough to explain it, and does more than repeat the tool's name.
- When to use: the description says when to use the tool, when not to, and what to use instead.
- Arguments: every argument has a description; string arguments have a constraint or an
example (
enum,format,pattern,examples,default,minLengthormaxLength); the schema lists therequiredarguments and setsadditionalPropertiestofalse. - Return value: the tool has an
outputSchema, or its description says what it returns. - Constraints and side effects: the tool has
annotations(readOnlyHint,destructiveHint,idempotentHint), and its description mentions limits, side effects, permissions, costs or rate limits. - Examples: the description, or an argument, gives an example.
The rules look for English words in the descriptions, so descriptions in other languages score lower than they deserve.
Score bands
A weather sign goes with each score, so that you can scan a long list:
70 and above: clear40–69: cloudybelow 40: rain
The number is what counts; the sign only groups it.
Across the server
Some problems only show when tools are seen together:
- Too many tools: above 25 tools, models start to tell them apart less well, and above 40 they choose less accurately.
- Tools that are easy to mix up: pairs whose descriptions share most of their words, or whose names are alike and descriptions partly alike. A result shows the 12 most similar pairs and counts all of them; each tool's card names the tools it is easily confused with.
- Identical descriptions: tools that share one description cannot be told apart at all.
Severity
- Critical: a model has nothing to choose the tool by. Fix these first.
- Major: a model is likely to choose the tool at the wrong time, or call it wrongly.
- Minor: details that would help a model, such as the return value or constraints.
The issues alone do not fail the CLI; set --fail-under to stop a CI job on a low average (see
Score with the CLI).
Issue codes
The codes are the same on the web, in the CLI and in its --json report. The sentences below are
examples: a result fills in the real numbers and names.
Issues of each tool
| Code | Severity | What a result says (example) |
|---|---|---|
no_description | Critical | No description. |
restates_name | Major | The description only restates the name (4 words). |
too_short | Major | The description is too short (5 words). |
no_usage_context | Major | Says nothing about when to use it, or when not to. |
destructive_unmarked | Major | It may be destructive, but destructiveHint is not set. |
generic_name | Major | The name is made of generic words only. |
bad_name_chars | Major | The name contains characters other than letters, digits, _, - and . |
param_no_description | MajorMinor | Arguments without a description: path, recursive. |
too_long | Minor | The description is long (430 words). It helps selection but costs steps and tokens. |
loose_string_param | Minor | String arguments with no constraint or example: query. |
no_required | Minor | No `required` list, so every argument is optional. |
params_not_in_description | Minor | The description mentions none of its 3 or more arguments. |
vague_boolean | Minor | The meaning of the boolean argument recursive is unclear. |
no_return_info | Minor | Says nothing about what it returns. |
no_constraints | Minor | No constraints, side effects, auth or rate limits described, and no annotations. |
long_name | Minor | The name is long (48 characters). |
Issues across the server
| Code | Severity | What a result says (example) |
|---|---|---|
identical_description | Critical | These tools share the same description: search_pages, search_docs. |
too_many_tools | Major | 52 tools. Above 40, models choose less accurately. |
confusable_pair | Major | get_block and retrieve_block are easy to mix up. |
confusable_pair_total | Major | 57 confusable pairs in all; the top 12 are shown. |
many_tools | Minor | 31 tools. Above 25, models start to tell tools apart less well. |
Versions of the scoring rules
Every result shows the version of the scoring rules it was scored with. Any change that can move a score, an item, an issue code or a severity comes with a new version, so compare scores only within one version. What static scoring cannot tell at all is in Getting started.