All docs

Read your scores

How a tool's score out of 100 is made, what the checks across the server look for, and every issue code with its severity.

On this page

Each tool gets a score out of 100, the sum of six items. The server's average is the mean of its tools' scores, with one decimal place. Each result also lists issues: what to fix, each with an issue code and a severity.

The six items

ItemPoints
Purpose20
When to use20
Arguments25
Return value15
Constraints and side effects10
Examples10
  • Purpose: the description exists, says what the tool does (starting with a verb helps), is long enough to explain it, and does more than repeat the tool's name.
  • When to use: the description says when to use the tool, when not to, and what to use instead.
  • Arguments: every argument has a description; string arguments have a constraint or an example (enum, format, pattern, examples, default, minLength or maxLength); the schema lists the required arguments and sets additionalProperties to false.
  • Return value: the tool has an outputSchema, or its description says what it returns.
  • Constraints and side effects: the tool has annotations (readOnlyHint, destructiveHint, idempotentHint), and its description mentions limits, side effects, permissions, costs or rate limits.
  • Examples: the description, or an argument, gives an example.

The rules look for English words in the descriptions, so descriptions in other languages score lower than they deserve.

Score bands

A weather sign goes with each score, so that you can scan a long list:

70 and above: clear40–69: cloudybelow 40: rain

The number is what counts; the sign only groups it.

Across the server

Some problems only show when tools are seen together:

  • Too many tools: above 25 tools, models start to tell them apart less well, and above 40 they choose less accurately.
  • Tools that are easy to mix up: pairs whose descriptions share most of their words, or whose names are alike and descriptions partly alike. A result shows the 12 most similar pairs and counts all of them; each tool's card names the tools it is easily confused with.
  • Identical descriptions: tools that share one description cannot be told apart at all.

Severity

  • Critical: a model has nothing to choose the tool by. Fix these first.
  • Major: a model is likely to choose the tool at the wrong time, or call it wrongly.
  • Minor: details that would help a model, such as the return value or constraints.

The issues alone do not fail the CLI; set --fail-under to stop a CI job on a low average (see Score with the CLI).

Issue codes

The codes are the same on the web, in the CLI and in its --json report. The sentences below are examples: a result fills in the real numbers and names.

Issues of each tool

CodeSeverityWhat a result says (example)
no_descriptionCriticalNo description.
restates_nameMajorThe description only restates the name (4 words).
too_shortMajorThe description is too short (5 words).
no_usage_contextMajorSays nothing about when to use it, or when not to.
destructive_unmarkedMajorIt may be destructive, but destructiveHint is not set.
generic_nameMajorThe name is made of generic words only.
bad_name_charsMajorThe name contains characters other than letters, digits, _, - and .
param_no_descriptionMajorMinorArguments without a description: path, recursive.
too_longMinorThe description is long (430 words). It helps selection but costs steps and tokens.
loose_string_paramMinorString arguments with no constraint or example: query.
no_requiredMinorNo `required` list, so every argument is optional.
params_not_in_descriptionMinorThe description mentions none of its 3 or more arguments.
vague_booleanMinorThe meaning of the boolean argument recursive is unclear.
no_return_infoMinorSays nothing about what it returns.
no_constraintsMinorNo constraints, side effects, auth or rate limits described, and no annotations.
long_nameMinorThe name is long (48 characters).

Issues across the server

CodeSeverityWhat a result says (example)
identical_descriptionCriticalThese tools share the same description: search_pages, search_docs.
too_many_toolsMajor52 tools. Above 40, models choose less accurately.
confusable_pairMajorget_block and retrieve_block are easy to mix up.
confusable_pair_totalMajor57 confusable pairs in all; the top 12 are shown.
many_toolsMinor31 tools. Above 25, models start to tell tools apart less well.

Versions of the scoring rules

Every result shows the version of the scoring rules it was scored with. Any change that can move a score, an item, an issue code or a severity comes with a new version, so compare scores only within one version. What static scoring cannot tell at all is in Getting started.