# Read your scores

How a tool's score out of 100 is made, what the checks across the server look for, and every issue code with its severity.

Each tool gets a score out of 100, the sum of six items. The server's average is the mean of its
tools' scores, with one decimal place. Each result also lists issues: what to fix, each with an
issue code and a severity.

## The six items

| Item | Points |
|---|---:|
| Purpose | 20 |
| When to use | 20 |
| Arguments | 25 |
| Return value | 15 |
| Constraints and side effects | 10 |
| Examples | 10 |

- **Purpose**: the description exists, says what the tool does (starting with a verb helps), is
  long enough to explain it, and does more than repeat the tool's name.
- **When to use**: the description says when to use the tool, when not to, and what to use
  instead.
- **Arguments**: every argument has a description; string arguments have a constraint or an
  example (`enum`, `format`, `pattern`, `examples`, `default`, `minLength` or `maxLength`); the
  schema lists the `required` arguments and sets `additionalProperties` to `false`.
- **Return value**: the tool has an `outputSchema`, or its description says what it returns.
- **Constraints and side effects**: the tool has `annotations` (`readOnlyHint`, `destructiveHint`,
  `idempotentHint`), and its description mentions limits, side effects, permissions, costs or
  rate limits.
- **Examples**: the description, or an argument, gives an example.

The rules look for English words in the descriptions, so descriptions in other languages score
lower than they deserve.

## Score bands

A weather sign goes with each score, so that you can scan a long list:

70 and above: clear · 40–69: cloudy · below 40: rain

The number is what counts; the sign only groups it.

## Across the server

Some problems only show when tools are seen together:

- **Too many tools**: above 25 tools, models start to tell them apart less well, and above 40
  they choose less accurately.
- **Tools that are easy to mix up**: pairs whose descriptions share most of their words, or whose
  names are alike and descriptions partly alike. A result shows the 12 most similar pairs and
  counts all of them; each tool's card names the tools it is easily confused with.
- **Identical descriptions**: tools that share one description cannot be told apart at all.

## Severity

- **Critical**: a model has nothing to choose the tool by. Fix these first.
- **Major**: a model is likely to choose the tool at the wrong time, or call it wrongly.
- **Minor**: details that would help a model, such as the return value or constraints.

The issues alone do not fail the CLI; set `--fail-under` to stop a CI job on a low average (see
[Score with the CLI](https://forecall.dev/en/docs/cli#ci)).

## Issue codes

The codes are the same on the web, in the CLI and in its `--json` report. The sentences below are
examples: a result fills in the real numbers and names.

### Issues of each tool

| Code | Severity | What a result says (example) |
|---|---|---|
| `no_description` | Critical | No description. |
| `restates_name` | Major | The description only restates the name (4 words). |
| `too_short` | Major | The description is too short (5 words). |
| `no_usage_context` | Major | Says nothing about when to use it, or when not to. |
| `destructive_unmarked` | Major | It may be destructive, but destructiveHint is not set. |
| `generic_name` | Major | The name is made of generic words only. |
| `bad_name_chars` | Major | The name contains characters other than letters, digits, _, - and . |
| `param_no_description` | Major / Minor | Arguments without a description: path, recursive. |
| `too_long` | Minor | The description is long (430 words). It helps selection but costs steps and tokens. |
| `loose_string_param` | Minor | String arguments with no constraint or example: query. |
| `no_required` | Minor | No `required` list, so every argument is optional. |
| `params_not_in_description` | Minor | The description mentions none of its 3 or more arguments. |
| `vague_boolean` | Minor | The meaning of the boolean argument recursive is unclear. |
| `no_return_info` | Minor | Says nothing about what it returns. |
| `no_constraints` | Minor | No constraints, side effects, auth or rate limits described, and no annotations. |
| `long_name` | Minor | The name is long (48 characters). |

### Issues across the server

| Code | Severity | What a result says (example) |
|---|---|---|
| `identical_description` | Critical | These tools share the same description: search_pages, search_docs. |
| `too_many_tools` | Major | 52 tools. Above 40, models choose less accurately. |
| `confusable_pair` | Major | get_block and retrieve_block are easy to mix up. |
| `confusable_pair_total` | Major | 57 confusable pairs in all; the top 12 are shown. |
| `many_tools` | Minor | 31 tools. Above 25, models start to tell tools apart less well. |

## Versions of the scoring rules

Every result shows the version of the scoring rules it was scored with. Any change that can move a
score, an item, an issue code or a severity comes with a new version, so compare scores only
within one version. What static scoring cannot tell at all is in
[Getting started](https://forecall.dev/en/docs/getting-started#limits).
